Test set strategy and dataset lifecycle
How do we iterate fast without overfitting our evaluation?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Say a test set passes 78 percent of 200 curated cases, and two weeks after release production quality is at 52 percent. The cases were six months old and mostly simple lookups, while users had moved on to multi-table joins and trend comparisons. Evaluation data goes bad in three ways: it goes stale, it gets contaminated when test examples leak into development, and it loses its version history so no one can reproduce an old result.
The lesson follows a dataset through four stages. At creation you split it into train, dev and holdout, 10, 45 and 45 percent, before any judge or metric work. In development you iterate on train and dev only, and every change, label corrections included, gets a version and a changelog entry. Once the judge settles, dev becomes a locked regression suite that you rotate on a schedule. When it no longer fits production, you retire it into an archive and keep it, because past release decisions cite it.
A large part of the lesson is keeping the holdout independent. Running it after every prompt revision, copying its examples into prompts, or tuning toward what you heard it contains all leak information into development. You predict what happens to a judge’s agreement score when it finally meets the holdout, and see why the dev-set number is too high.
The demo shows how to check whether a suite still matches production, with a Kolmogorov-Smirnov test on query complexity, user role and failure category, and how to decide whether to rotate.
The practice is a Dataset Management Spec for the AI Data Analyst’s regression suite, written on paper: you work out the split sizes, write the version history, set the drift check and what triggers a rotation, and write the holdout rules and retirement criteria.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→