Launch readiness and rollout gates
What must be true before exposure, and what do we do if it degrades?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
v2 was rolled back in 5.3 because latency and cost went past their guardrails. This lesson is about what its fix, v2.1, would have to show before any user sees it. The shadow and canary numbers in the lesson are practice numbers for v2.1, not measurements. Even a clean rerun leaves questions open. Can you see a failure after launch? Is there a runbook if quality drops overnight? Do your offline numbers hold up on real traffic?
So you roll out in stages: shadow, where v2 runs on live traffic but users still see v1; canary, at 1 to 5 percent of users; ring 1, at 10 to 25 percent; then everyone. Each stage has entry criteria that say what must be true to move in, and exit criteria that say what sends you back.
Shadow is where you check offline-online agreement. You compare three measures of the same traces: the judge’s score for the SQL, an oracle check that runs the SQL against a known answer, and whether the trace made it all the way to a written summary. Each one is stricter than the last. At canary you run a gate check on four blocking metrics, quality, latency, cost and safety, with the latency and cost gates set from v1 plus the 5.3 guardrails, and you look at the confidence intervals as well as the point estimates. Passing a gate with room to spare is different evidence from passing it by a point.
You also write the rollback triggers before launch: automatic rollback when a blocking metric fails, manual investigation when something drifts toward its limit, each with a metric, threshold, time window and owner. And you check the operational side: alerts tested, on-call staffed, runbook written, budget approved, sign-offs recorded.
The practice is a Rollout Decision Document for v2.1, built on the practice numbers from the slides: the canary gate check, entry and exit criteria for all four stages, rollback triggers, an operational checklist, and a ship, ramp or hold recommendation with the evidence behind it.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→