Week 5: Pipelines, Experiments, and Continuous Validation · Lesson 5.4

Launch readiness and rollout gates

What must be true before exposure, and what do we do if it degrades?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome back. Last lesson the v2 experiment came back with a real gain on SQL success and two failed guardrails, latency and cost, so v2 was rolled back. The team is now working on a fix. Call it v2.1. Today is about what v2.1 has to show before any user sees it. Who sees it first? What would make you stop? Who gets paged if it breaks at night? We turn a test result into a rollout: a sequence of stages, each with a gate you have to pass before you move on, and a rollback rule you write down before you need it. Let's start with what the experiment told you.

About this lesson

v2 was rolled back in 5.3 because latency and cost went past their guardrails. This lesson is about what its fix, v2.1, would have to show before any user sees it. The shadow and canary numbers in the lesson are practice numbers for v2.1, not measurements. Even a clean rerun leaves questions open. Can you see a failure after launch? Is there a runbook if quality drops overnight? Do your offline numbers hold up on real traffic?

So you roll out in stages: shadow, where v2 runs on live traffic but users still see v1; canary, at 1 to 5 percent of users; ring 1, at 10 to 25 percent; then everyone. Each stage has entry criteria that say what must be true to move in, and exit criteria that say what sends you back.

Shadow is where you check offline-online agreement. You compare three measures of the same traces: the judge’s score for the SQL, an oracle check that runs the SQL against a known answer, and whether the trace made it all the way to a written summary. Each one is stricter than the last. At canary you run a gate check on four blocking metrics, quality, latency, cost and safety, with the latency and cost gates set from v1 plus the 5.3 guardrails, and you look at the confidence intervals as well as the point estimates. Passing a gate with room to spare is different evidence from passing it by a point.

You also write the rollback triggers before launch: automatic rollback when a blocking metric fails, manual investigation when something drifts toward its limit, each with a metric, threshold, time window and owner. And you check the operational side: alerts tested, on-call staffed, runbook written, budget approved, sign-offs recorded.

The practice is a Rollout Decision Document for v2.1, built on the practice numbers from the slides: the canary gate check, entry and exit criteria for all four stages, rollback triggers, an operational checklist, and a ship, ramp or hold recommendation with the evidence behind it.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→