Experiment design for stochastic systems
How do we run online tests when outcomes are noisy, long-tailed, and mediated through user behavior?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
A better offline score for v2 doesn’t tell you what happens when real users get it. They ask follow-ups, rephrase when they’re confused and give up when an answer looks wrong. To find out, you run an online experiment, and an AI system makes that harder than a normal A/B test. The same question can come back with a different answer on every run, so there’s more noise and you need more users to see the same effect.
The lesson walks through the design decisions you make before any data comes in: what you’re measuring, what gets randomized, the primary, secondary and guardrail metrics, the sample size from a power analysis, the stop conditions, and the rules for each outcome: ship, ramp, hold or roll back. Writing the rules first keeps you from bending them once you’ve seen a number you like.
Then you read the course’s own results. v2 raised SQL success by 2.7 points, a real gain, and it also raised latency 17.6 percent and cost 20 percent, past the 10 and 15 percent guardrails set before the test. By the rules, that’s roll back.
It also covers what happens when users share something. The AI Data Analyst has a shared cache, so control users can pick up some of v2’s better retrievals. A switchback, where everyone gets the same version in each six-hour block, is meant to remove that. On the course data it came back smaller and much noisier, with an interval that crosses zero and a carryover gap of its own, so it doesn’t change the call.
The practice is an Experiment Design One-Pager for the v1 to v2 change, built from the numbers in the lesson: check the sample size against the power analysis, match the effect and interval to a rule, work out both guardrail changes yourself, and write the decision. The extended version adds the switchback and carryover results and designs the rerun on the fix.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→