Distributional thinking for AI quality
How do we reason about AI behavior when outputs vary across inputs, contexts, and users?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Say you have 100 traces and 62 of them look acceptable. Your PM asks whether 62% is good enough to ship, and you cannot answer. You do not know whether a re-run would give 58 or 67. You do not know whether the system fails badly on 5% of cases or is mediocre on 38%. You do not know whether the variation comes from the system or from your evaluation method.
Three tools fix that. Bootstrap resampling puts a confidence interval on any metric without assuming anything about the shape of the data: resample your outcomes with replacement, recompute, repeat, and read the 2.5th and 97.5th percentiles. Segment breakdowns show what an aggregate hides. In a made-up example an aggregate pass@5 of 0.78 hides multi-table joins at 0.45, and in the v0 traces reliable@5 runs from 0.12 on trend questions to 0.50 on comparisons. Variance decomposition separates the system’s randomness from the evaluator’s noise, because the two need different fixes.
Then the decision. Full rollout when the intervals are narrow and every must-pass metric clears its threshold in every segment. Gradual ramp when the intervals are wider but the direction is positive. Hold or roll back when an interval includes the threshold, because you cannot tell pass from fail.
You compute pass@5 and reliable@5 for five trials of one query using the two formulas, bootstrap the interval, and see why five trials is not enough evidence. Then you look at the 47 questions the v0 traces hold five runs of. pass@5 reaches 1.0, reliable@5 falls to 0.36, and most of the variation is the same question passing on one run and failing on the next.
The practice is a distributional quality report: the gap between pass@5 and reliable@5 for each segment, a ship criterion written with a threshold and an interval width, and a ship, ramp or hold call per segment with the evidence behind it.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→