AI Evals for Product DevelopmentL6.3 · 01

Prioritization and
iteration

Ranking the evaluation backlog with evidence, and writing down what "done" means for each item
The backlogL6.3 · 02

Eight failure modes, all backed by evidence. Which one do you fix first?

Failure mode
Frequency
Severity
Failure mode
Frequency
Severity
Retrieval failures
18%
Medium
SQL errors
12%
High
Formatting issues
15%
Low
Cost regressions
8%
Medium
Citation errors
7%
Medium
Latency spikes
5%
Low
Hallucinations
3%
High
Policy violations
0.8%
Catastrophic
From your failure taxonomy (Week 1), driver analysis (Week 4) and experiment results (Week 5).
Seven dimensionsL6.3 · 03

Score every item on seven dimensions, in two groups.

ImpactUser harmFrequencyBusiness criticalityConfidence in the evidence
VelocityFixabilityTime to learnReversibility
Each dimension is scored 1 to 5. Weights reflect what your team cares about.
User harmL6.3 · 04

User harm asks how bad it is when the failure happens.

5CatastrophicSafety or compliance violation
4HighData loss or a wrong decision
3MediumWorkflow blocked for a while
2MinorAn extra step for the user
1CosmeticFormatting inconsistency
Silent failures count. An error nobody notices can still drive a bad decision.
Frequency, criticality, confidenceL6.3 · 05

Every score has to point to an artifact you already built.

Frequency
A rate, as a share of queries. Retrieval failures: 18% of queries, from the Week 4 driver analysis.
Business criticality
Who it hits. Retrieval failures land on PM and DS users, who generate 60% of queries.
Confidence
How sure you are it's real. Taxonomy, driver analysis and retrieval metrics all agree, or one Slack message.
VelocityL6.3 · 06

Velocity: how fast you can fix it, learn from it, and undo it.

Instruction change
Retrieval architecture change
Fixability
Clear change, can happen today
Harder to predict. Two weeks in, it may not help
Time to learn
Hours, on the pre-production test set
Weeks, waiting on production metrics
Reversibility
Swap the prompt back
Migration rollback that takes days
The scoreL6.3 · 07

Multiply each score by its weight and add them up.

priority = 2 × user harm + frequency + business criticality + confidence + fixability + 1.5 × time to learn + reversibility
These are the weights for this example. Change them to match your team, and write down why.
Predict before the demoL6.3 · 08

You can fix one of these this sprint. Which one?

Work item
Frequency
User harm
Fixability
Time to learn
Reversibility
Improve retrieval quality on metric definition queries
18%
Medium
High
Fast
High
Fix policy compliance failures on queries that touch personal data
0.8%
Catastrophic
Medium
Slow
Low
Write your pick and one sentence of reasoning before we score them.
The demo: scoring the top fiveL6.3 · 09

Score each item, cite the evidence, then apply the weights.

Work item
User harm ×2
Freq
Biz crit
Confid.
Fixab.
Time to learn ×1.5
Revers.
Retrieval quality
3
5
4
5
5
5
5
Hallucination rate
5
3
5
4
3
3
4
SQL logic errors
4
4
3
4
3
4
4
Policy compliance
5
1
5
4
3
2
2
Cost guardrails
3
3
4
5
4
3
4
Blue is 4 or 5. Red is 1 or 2.
Tail risksL6.3 · 10

A rare, catastrophic failure gets a flag the score can't override.

What goes wrong
A 0.8% failure rate looks small, so the item drifts to the bottom of the backlog. Then it fires in production.
What to do
Score harm separately from frequency. Flag any item with catastrophic harm as a critical failure so it is worked on this cycle, whatever its rank.
Segment checkL6.3 · 11

Before you commit to the top item, check it by segment. Is it twice as bad anywhere?

Overall18%
PM22%
DS25%
Engineering8%
Executive15%
Retrieval failure rate by user segment, from the Week 4 driver analysis.
Acceptance criteriaL6.3 · 12

Write down what "done" means before you start.

What metric moves, and by how much?Retrieval finds the right context in the top 3 results 75% of the time, up from 68%, on metric definition queries.
Which segment has to see it?PM and DS users, who generate 60% of queries.
What evidence closes the ticket?A 50-query test set with known-correct answers shows 75% or better.
"Improve retrieval quality" is a direction. It never closes.
The whole methodL6.3 · 13

From evidence to a ranked backlog you can defend.

EvidenceFailure taxonomy, driver analysis, experiment results, findings-to-actions plan
Scoring matrixOne row per item, seven scores with citations, weights applied, critical failures flagged
Ranked backlogSorted by score, segment check on the top item, acceptance criteria on every top item
PracticeL6.3 · 14

Build a ranked iteration backlog with acceptance criteria.

Base version, everyone
Take the eight failure modes from slide 2, or 8 to 10 items from your own product. Score each on the seven dimensions with a citation per score. Apply the weights, rank, flag at least one tail risk, run the segment check on the top item, and write acceptance criteria for the top five.
Extended version, DS and engineering
By hand, rescore the five demo items with the weight on user harm doubled, then with the weight on time to learn doubled. Note which items move and which tie. Then write one paragraph on which weights your team should use, and why.
Knowledge checkL6.3 · 15

Three judgment calls. Write your answer before you read on.

01A ticket's acceptance criterion reads "Improve task completion rate." Rewrite it using the three questions.
02Item A: hallucinated numbers on revenue queries, 0.8% of queries, catastrophic harm, slow to learn. Item B: SQL success on simple lookups, 15% of queries, minor harm, fast to learn. A stakeholder says "do B, it affects way more users." What's missing, and how do you respond?
03Your top item is "fix retrieval failures on schema lookups." It fails on 35% of DS queries and 5% of PM queries. DS users generate 40% of revenue. Does the ranking change?
Next lessonL6.3 · 16

You have a ranked backlog. Who owns it?

Engineering
"We own the fixes. It's code."
Product
"We own the order. It's the roadmap."
Data science
"We own the metrics. We built the evals."
Next: the ownership model and the AI Reliability Lead.
AI ANALYST LAB · aianalystlab.ai