Ground truth sources, regression suites, and synthetic data
What are we comparing against, and how do we cover cases production hasn't shown us yet?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Your AI Data Analyst scores 90% on SQL correctness. The first question to ask is what that 90% was measured against. If the expected answers were copied from an older model’s output and nobody ran them against the database, the suite can pass while users see wrong numbers.
Ground truth is the set of answers you have checked and trust. It comes from different places for different tasks. For SQL you use an execution oracle: run the generated query and a verified expected query against the warehouse and compare the result sets, so two differently written queries that return the same rows both pass. For the written summary you need people reading it against a rubric, and before you scale that you check that two annotators agree, using Cohen’s kappa. User feedback is useful for watching production and too noisy to write test cases from.
Verified cases go into a regression suite, the set every release has to pass. You start with 10 to 20 high-impact cases from your Week 1 failure taxonomy and add more through promotion. A case gets in only if it is reproducible, covers a failure mode the suite doesn’t already cover, has verified ground truth, and is severe or frequent enough to be worth maintaining. For gaps production hasn’t shown you yet, you generate synthetic queries from a set of attributes and drop the ones no real user would ask.
The practice is a Regression Suite Promotion Plan: write the oracle’s explanation for the worked Q4 failure, run three failures from your taxonomy through the promotion criteria, and list three coverage gaps. The extended version writes synthetic queries for one gap and applies the realism check.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→