AI Evals for Product DevelopmentL5.4 · 01

Launch readiness
and rollout gates

What has to be true before users see a change, and what you do when it starts to slip
Where we left offL5.4 · 02

What did the v2 experiment show, and which rule made the call?

Treatment effect
How much did v2 move SQL success rate, and did the interval include zero?
Guardrails
Did latency and cost stay inside their limits?
Answer from memory before you read on.
The scenarioL5.4 · 03

Say v2.1 passes its rerun. Engineering still has three questions nobody can answer.

01Is monitoring in place to catch a silent failure after launch?
02If quality drops 5% overnight, is there a runbook, and who runs it?
03Do our offline metrics predict what happens on production traffic?
The rolloutL5.4 · 04

Roll out in stages. Each stage has a gate to enter and a rule for leaving.

Shadow0% of usersv2.1 runs on live traffic. Users still see v1.
Canary1 to 5%A small group of real users sees v2.1.
Ring 110 to 25%A larger group, big enough to check segments.
Full rollout100%Everyone, with drift monitoring on.
Entry criteria say what must be true to move forward. Exit criteria say what sends you back.
Three kinds of readinessL5.4 · 05

Ready means three things, and the experiment only covers the first.

Technical validation
Offline metrics, the experiment result, shadow agreement, canary checks.
Operational readiness
Alerts configured, on-call staffed, runbook written, cost budget approved, compliance signed off.
Decision governance
Who approves each step, what evidence they need, what triggers a rollback.
ShadowL5.4 · 06

Shadow tests whether your offline numbers hold up on production traffic.

What shadow can tell you
Whether v2.1 handles real questions, real volume and real data, and whether your offline metrics agree with what happens on that traffic.
What shadow cannot tell you
How users react. Nobody sees v2.1's answers, so there are no retries, abandons or complaints yet.
PredictL5.4 · 07

On the same shadow traces, will your judge's number come out higher or lower than what actually worked?

Write one sentence: which direction will the gap go, and what causes it?
Shadow resultsL5.4 · 08

Three ways to measure the same shadow traces, from most lenient to strictest.

JudgeSQL success, judgedAn LLM judge reads the SQL and says whether it looks correct.
OracleSQL success, executedThe SQL runs and its results are compared to a known-correct answer.
PipelineTask completionThe trace made it all the way to a written narrative.
Each one checks more of what has to go right.
Mapping gates to metricsL5.4 · 09

At canary, four blocking metrics decide whether you move on.

Metric
Gate
What it protects
SQL success rate
≥ 0.74
Answer quality, the reason for the change
Average latency
≤ 1,560 ms
Users waiting on an answer
Average cost per query
≤ $0.029
The budget
Safety violation rate
≤ 1%
Users and the company
If any gate fails, you do not advance to ring 1.
PredictL5.4 · 10

Given what shadow showed, which of the four gates is most likely to be in trouble at canary?

SQL success, latency, cost or safety? Pick one and give your reason.
Canary gate checkL5.4 · 11

Check each gate twice: the point estimate, then the interval.

What the rule says
Is each observed value on the right side of its gate? If so, canary passes.
What the intervals say
Does the plausible range for any metric still reach past its gate?
That difference is what separates ramp from ship.
Rollback criteriaL5.4 · 12

Write the rollback triggers before launch, so nobody has to argue during an incident.

Automatic rollback
A blocking metric fails its gate. Traffic goes back to v1 without a meeting.
Manual investigation
A metric moves toward its gate, or a non-blocking signal jumps. Someone named looks at it within a set time.
Every trigger names a metric, a threshold, a time window and an owner.
Operational readinessL5.4 · 13

Before canary starts, check these off.

01Monitoring alerts are configured on every blocking metric, and a test alert has arrived.
02The on-call rotation is staffed with people who know this system.
03An incident runbook says how to diagnose the common failures and how to roll back.
04The cost budget for the new version is approved.
05Compliance and security sign-offs are recorded.
PracticeL5.4 · 14

Write the Rollout Decision Document for v2.1.

Base version, everyone
Take the practice canary numbers from today and mark each gate pass, warn or fail, using the intervals. Write entry and exit criteria for all four stages. Define automatic and manual rollback triggers for each blocking metric. Build an operational checklist of at least five items. Finish with a ship, ramp or hold recommendation and the evidence behind it.
Extended version, DS and engineering
Work out the gaps between the three shadow measures and what they mean for the SQL gate. Say which query types you'd suspect and how you'd check. Compare each trigger at a one-hour and a one-day window. Propose three more signals for silent failures.
Common mistakesL5.4 · 15

Four ways a rollout goes wrong.

Mistake
What happens
Do this instead
Ignore the offline-online gap
The judge number goes into the decision as if it were what users get.
Report judge, oracle and completion side by side.
No rollback triggers
Cost jumps and three people argue about what to do while it keeps climbing.
Name metric, threshold, window and owner before launch.
Treat shadow as approval
Canary is skipped, and a problem that only shows up with real users reaches everyone.
Run canary even when shadow looks clean.
Ops as an afterthought
A 2 a.m. latency spike, and on-call can't diagnose it or roll it back.
Finish the checklist before canary starts.
Knowledge checkL5.4 · 16

Two judgment calls. Write your answer before you read on.

01In shadow, the judge scores SQL success at 82% and task completion is 74%. What does the gap tell you about the offline metric? Do you move to canary, or hold and look into it first?
02Halfway through canary, average cost per query is 12% above the v1 baseline. It is still under the 2.9 cent gate. Finance has noticed. Do you move to ring 1, ramp more slowly, or roll back?
Next lessonL5.4 · 17

The full rollout, with a way back from every stage. Next, you watch it after launch.

Shadow · 0%v2.1 runs, users see v1ENTEROffline eval passedLEAVEMetrics agree with offline
Canary · 1 to 5%Small group sees v2.1ENTERShadow passed, monitoring live, runbook readyLEAVEAll blocking gates pass
Ring 1 · 10 to 25%Larger groupENTERCanary passed for 3+ daysLEAVEEvery segment checked
Full · 100%EveryoneENTERRing 1 passed for 1+ weekTHENDrift monitoring stays on
Roll back from any stage when a blocking metric fails, cost overruns its budget, or complaints spike.
AI ANALYST LAB · aianalystlab.ai