AI Evals for Product DevelopmentL6.1 · 01

Decision-making
under uncertainty

Ship, ramp, hold, roll back, and the evidence each one needs
Where we left offL6.1 · 02

In Week 5, what tradeoff did you find between the must-pass metrics and the rest?

Blocking metrics
Must pass before a release goes out. SQL correctness and safety checks.
Optimization metrics
Tracked for improvement. Retrieval quality and task completion. A dip doesn't stop the release.
Guardrails like latency and cost sit alongside both.
The scenarioL6.1 · 03

Your PM asks whether v2 should go out. Here's the evidence.

SQL success
71.9% to 74.6%
Task completion
57.6% to 63.0%
Retrieval precision
69.8% to 77.7%
Average latency
1,416 to 1,665 ms, limit +10%
Users averaging over 2s
1.6% to 10.6%
Cost per query
2.5 to 3.0 cents, limit +15%
Two wordsL6.1 · 04

With only "ship" and "don't ship", a mixed result has nowhere to go.

"The quality gain is too big to leave on the table."
Pushes for everyone, now.
"The latency regression is too risky."
Pushes to wait until latency is fixed.
Both people are right about the evidence. The choice only has two options.
Six decisionsL6.1 · 05

Six decision types, each with its own evidence bar.

Decision
What happens
When it fits
Ship
Full rollout to all users
High confidence, every blocking metric passes, tradeoffs acceptable
Ramp
Gradual rollout: 10%, 25%, 50%, 100%
Moderate confidence, positive direction, monitoring catches problems fast
Hold
Don't deploy yet
Evidence too thin, or a blocking metric fails and the change needs a fix first
Roll back
Reverse a change that's already live
A blocking metric fails after deploy
Scope-restrict
Deploy to some segments only
Some user groups benefit and others don't
Human in the loop
Deploy with manual review
Quality is better but edge cases need oversight
Evidence sufficiencyL6.1 · 06

Four questions tell you which decision the evidence supports.

DirectionIs the change better, worse, or mixed?
MagnitudeHow much better or worse, on each dimension?
ConfidenceHow much can you trust the numbers, given the sample size and interval width?
Risk containmentWhat could go wrong after deploy, and how fast would you see it?
Ship needs a strong answer on all four. Ramp can live with moderate confidence if containment is strong.
Tradeoff transparencyL6.1 · 07

When the signals conflict, the tradeoff goes in writing.

"We recommend [decision] because [primary improvement]. The risk is [regression]. We mitigate via [monitoring, rollback trigger, scope restriction]."
What improvedSpecific metrics and sizes
What regressedSpecific metrics and sizes
Why it's acceptableTied to what users get
What limits the riskMonitoring, trigger, scope
Predict before the demoL6.1 · 08

Which decision would you make on v2, and why?

Quality
SQL success, task completion and retrieval all up, tight intervals
Latency
+17.6% against a 10% limit, 1 user in 10 over 2 seconds
Cost
+20.2% per query against a 15% limit
Segments
Not split yet. You'll see them next
Ship, ramp, hold, roll back, scope-restrict, or human in the loop. Write one sentence with your reason before we advance.
The demo: split by segmentL6.1 · 09

Four user groups got the same change. Who got the gain, and who paid for it?

PMAbout 30% of users
DSAbout 30% of users
EngineeringAbout 25% of users
ExecutiveAbout 15% of users
For each group: the SQL gain, the latency increase and the cost increase.
The demo: run the four questionsL6.1 · 10

Walk the four questions on v2 before you name a decision.

DirectionWhich metrics moved up, which moved down?
MagnitudeHow big is the gain next to the regression, on average and at the tail?
ConfidenceIs the sample big enough that these effects are real?
Risk containmentWhat do you watch, how often, and what number reverses it?
The demo: the decision memoL6.1 · 11

The decision memo has six fields.

recommendation
The decision type, and for whom
primary_evidence
The improvements, with metrics and sizes
tradeoff_acknowledgment
What got worse, and for which segment
risk_mitigation
What you monitor, by segment, how often
rollback_trigger
The exact number that reverses the change
decision_confidence
How sure you are, in one word
PracticeL6.1 · 12

Write your own decision memo for v2.

Base version, everyone
Answer confidence and risk containment yourself, using the numbers from this lesson. Write the memo with all six fields. Then check it: does every field name a metric and a number, and is there a rollback trigger?
Extended version, DS and engineering
From the per-query cost, work out what v2 adds per day at 100,000 queries. Using the segment ranges, check whether any user group stays inside both guardrails. Write a monitoring spec for v2.1: metric, granularity, threshold, action.
Common mistakesL6.1 · 13

Three ways a ship decision goes wrong.

Mistake
What happens
Do this instead
Anchoring on one metric
Ship on the SQL gain alone. Two weeks later, users with complex questions file tickets about slow answers.
Answer all four questions for every blocking and optimization metric.
Gathering evidence forever
10,000 users and tight intervals, and someone asks for two more weeks.
Use the rubric to say when the evidence is enough to ramp.
No rollback trigger
Latency spikes on day three and the team argues about whether to revert.
Write the trigger before you deploy.
Knowledge checkL6.1 · 14

Three judgment calls. Write your answer before you read on.

01Retrieval +40%, SQL correctness +2%, latency +200ms with 18% of queries over the SLA, cost +15%. All blocking metrics pass. Would you ship, ramp or hold, and how do the four questions support it?
02v2 improves average quality by 8%, but power users, 10% of your base, see a 5% quality drop. Which idea from today applies, and what do you recommend?
03You ramped to 20% with a pre-set trigger: more than 10% of queries over 2.5 seconds. Two days in, it fires. What do you do, and why does it help that the trigger was set in advance?
Next lessonL6.1 · 15

Next: the dashboard is green and users are complaining.

Today
Six decision types, four questions, and a memo with a rollback trigger set before deploy.
6.2
What to change next when user signals and evaluation metrics disagree, and how you'll know the fix helped.
AI ANALYST LAB · aianalystlab.ai