How much evidence is enough to decide
How many test cases do we need before the accuracy number can support this decision?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
You have 50 test queries with correct answers written ahead of time, and 42 of them come back right. That’s 84 percent. Your PM wants to roll out to everyone and asks whether 50 is enough. This lesson is how you answer that with a number instead of a guess.
Three things decide the sample size: the decision you’re making, how confident you need to be, and how noisy the measurement is. A full rollout needs a tight estimate, a limited ramp can live with a wider one, and comparing two versions needs a power analysis instead. Noise comes from the system and from the evaluator grading it.
The formula is n = z² × p(1 − p) / ε². You work through it on the 42-of-50 example against a rollout rule of plus or minus 5 points at 95 percent confidence, and see how far 50 samples falls short. Then you pull the two levers you control. Lowering the confidence level cuts the count. Tightening the margin raises it fast, because halving the margin roughly quadruples the samples.
An LLM evaluator isn’t perfectly consistent, and its kappa score tells you how much. The course uses a rough rule: divide the required sample size by kappa, so a noisy judge needs more samples to give you the same evidence.
The practice is a decision matrix: full rollout, hold and experiment, each at 90 and 95 percent confidence, with a sample size you work out from the formula and a two-sentence evidence argument in every cell. The extended version adjusts every cell for three evaluator kappa values.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→