Capstone lab: run the pipeline by hand
When an alert fires, can you tell from the run records whether the system got worse or the sample changed?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
This is the Week 5 lab. You run the evaluation pipeline from Lesson 5.1 yourself, twice, on real traces from the AI Data Analyst: sample, judge, aggregate, store, report and alert. You don’t need code. The two trace tables, the run record and the report table are all in the slides, and a spreadsheet or paper is enough.
Run A is a random sample of twelve traces from the full set of 1,000. Run B is a random twelve from the multi-table join questions only. You write the pass rules before you count, score each trace from the outcomes the pipeline recorded (did it reach the narrative, and if not, where did it stop), and compute completion rate, SQL success, mean latency and mean cost. Then you fill in a run record for each run and compare the numbers against three gates.
Run B fails the SQL gate. Most of the lab is about what you do next: put the two run records side by side to find what changed, look at how much one trace moves a twelve-trace number, and decide if the cause is a worse system, a weaker segment or a small unlucky sample.
The full version of the judging stage adds the oracle SQL check and the LLM judge from Week 3. Those need code and an API key, so the slide version uses recorded outcomes, and the knowledge check asks what that leaves out.
The practice is two run records, the completed report table, and one paragraph on the Run B alert. If you can get logs from an AI feature you work on, the extended version runs the same six stages on 50 of your own traces, with a second run that changes one thing.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→