Evaluation pipeline architecture and environments
How do we run evaluations the same way every time and keep enough of a record to compare any two runs?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
An evaluation that someone runs once and pastes into Slack can’t be compared with anything. If SQL correctness reads 87 percent on Monday and 83 percent on Wednesday, you can’t tell whether the system got worse or the sample, the judge or the dataset changed, because none of that was saved.
This lesson sets out the structure that fixes that. Every run goes through six stages: sampling, judging, aggregation, storage, reporting and alerts. Every run also writes a record of how it was done: the dataset version, model version, judge version and configuration, code commit, sample size and sampling strategy. With that record you can look at two runs and say what differs between them. A worked example shows why: two runs on the same traces that differ only in the judge version can’t tell you whether the system improved.
You also see where the v1 AI Data Analyst actually breaks, using the first failure step each trace records, and how much evaluation to run at each stage of a rollout, from the full suite offline to sampled blocking metrics with alerts in production.
The practice is on paper. You write the run record for a baseline run, use the failure counts from the lesson to explain the top two failing steps, explain the gap between the two runs in the worked example, then design a third run that changes one thing on purpose and predict what it will show. In Lesson 5.7 you run the whole pipeline by hand.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→