AI Evals for Product DevelopmentL1.3 · 01

Failure
discovery

Building a failure taxonomy from raw traces
Where we left offL1.3 · 02

Which failure modes ended up in your "cannot evaluate" column?

What you built in 1.2
Three layers: functional failures, adversarial vectors, coverage gaps. Then a split into can-evaluate-now and needs-better-logging.
Recall
Which failure modes could you not evaluate with v0's logs? What did you list as logging blind spots?
The scenarioL1.3 · 03

Say fifty complaints arrive in seven days, and no two describe the problem the same way.

"Gave me last quarter's data when I asked for this quarter."
"SQL failed but there was no error message."
"The narrative said 'slight decline' when revenue dropped 40%."
"Completely hallucinated a metric that doesn't exist."
Where do you even start?
Two approachesL1.3 · 04

A checklist only catches what somebody already imagined.

Checklist approach
Hallucination, SQL syntax error, latency spike, missing data. Anything that doesn't fit a bucket gets forced in or dropped.
Discovery approach
Read the traces with no categories. Write down what you see. Let the categories come from the notes.
Correct SQL with the wrong time range. A summary that minimizes a 40% drop. Wrong background information for the question. None of these fit a standard bucket.
The principleL1.3 · 05

Read first. Name second.

Top-down
Start with categories and force-fit the traces. You see what you expect and miss what you don't.
Bottom-up
Start with the traces and let categories emerge. You see what is actually there.
A 2025 study built its failure taxonomy for multi-agent systems this way: 150 traces read with no categories fixed in advance, 14 failure modes found.
The methodL1.3 · 06

Five steps from raw traces to a taxonomy you can act on.

01Read tracesOpen mind. No categories yet.
02Freeform notesOne note per trace, in your own words.
03Cluster bottom-upGroup similar notes into categories.
04Check saturationAre new traces still producing new categories?
05TriagePrompt-fix, evaluator-needed, or system-fix.
The discipline is in not skipping ahead.
SaturationL1.3 · 07

If trace 30 gives you a new category, you haven't read enough.

5
10
20
25
30
Illustration: categories found, by traces read
New category on the next 10?
Keep reading.
Nothing new on the next 10?
Move on to triage.
TriageL1.3 · 08

Every category gets one of three labels.

Prompt-fix
Your team can improve it by changing the prompt this sprint. Example: the narrative minimizes findings, so add severity language to the prompt.
Evaluator-needed
You need a metric to detect it automatically. Example: wrong time range, so build a check that compares the question's period to the SQL's date filters. Week 3.
System-fix
It needs a code or architecture change. Example: SQL fails silently, so the system has to surface errors. No prompt change fixes that.
Predict before the demoL1.3 · 09

You have 500 traces. How many distinct failure categories are in there?

Before reading any of them, write a number. Then list three failure types you expect to find, specifically.
"SQL errors" is too broad. "Wrong table joins" or "hallucinated column names" is the level we want.
The demoL1.3 · 10

The first five v0 traces.

Trace 1Asked for the revenue trend. The answer is about DAU, and calls 41,738 to 37,075 a "14.2% change."
Trace 2Asked about session duration. The SQL failed with "relation does not exist." A revenue summary came back anyway.
Trace 3Asked for the conversion funnel. The answer is about cohort retention and reports 101671% retention.
Trace 4SQL failed with no error message. The summary says DAU is "improving" from 31,981 to 48,040, "a -4.5% change."
Trace 5Asked to compare this month's revenue to last month. The answer compares mobile and desktop conversion.
The first five traces in the v0 file. Write one freeform note for each before we discuss them.
Cluster and triageL1.3 · 11

Five notes, five categories. Triage turns them into a sprint plan.

Category
Severity
Triage
Next action
Answers a different question
Critical
Evaluator-needed
Check that the answer covers the metric and period asked for.
Narrative contradicts its numbers
Critical
Prompt-fix
Instruct the narrative to state size and direction of change accurately.
Impossible value
Critical
Evaluator-needed
A range check on reported values.
SQL error with a message
Major
System-fix
Stop writing a narrative after the SQL fails.
Silent SQL failure
Major
System-fix
Surface the error. No prompt change fixes this.
PracticeL1.3 · 12

Build a failure taxonomy from the first five v0 traces.

Base version, everyone
Start from your five notes on slide 10. Cluster them into named categories, each with a one-line description and the trace numbers that belong to it. Assign severity and a triage label to each. Then write the saturation rule you'd use on the full file: when you'd stop reading, and what would keep you going.
Going further
For each category, say how you'd estimate its frequency, how hard it would be to detect automatically, and whether it needs better logging before you can count it.
Write your own categories before you look back at slide 11.
Common mistakesL1.3 · 13

Three ways the method breaks.

Mistake
What happens
Do this instead
Premature categorization
Start with buckets after five traces and force-fit everything after. The unexpected failure types disappear.
Read 30 or more traces with freeform notes before creating any category.
Insufficient reading depth
Read ten traces and declare saturation. Categories that only appear after trace 20 are missed.
Run the saturation check: ten more traces, zero new categories.
Mixing up severity and frequency
A rare but critical failure gets marked minor because it's infrequent.
Severity is impact per incident and frequency is how often, so keep them separate.
Knowledge checkL1.3 · 14

Three judgment calls. Write your answer before you read on.

01The AI writes correct SQL and reports accurate numbers, but the summary leaves out the most important finding. Prompt-fix, evaluator-needed, or system-fix? Why?
02After 25 traces you have seven categories, and the last five traces all fit existing ones. A colleague says you've reached saturation. What do you check?
03"Wrong time range" failures are rare, about 3%, but critical. "Verbose narrative" failures are common, about 15%, but minor. Which gets a metric first?
Next lessonL1.3 · 15

The taxonomy feeds three things downstream. Next, you put error bars on it.

What to log
Coverage gaps feed Week 2's logging design.
What to measure
Evaluator-needed categories feed Week 3's metric design.
What to fix now
Prompt-fix categories go into this sprint.
Lesson 1.4: failure rates as distributions instead of single numbers.
AI ANALYST LAB · aianalystlab.ai