Failure surfaces and annotation-based analysis
Where do AI systems break in practice, and how do we turn failures into structured evidence?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Say fifty complaints arrive in seven days and no two are described the same way. One user got last quarter’s data when they asked for this quarter. One had SQL fail with no error message. One read “slight decline” in a summary when revenue dropped 40%. Your instinct is to reach for a checklist, and a checklist only catches the failures somebody already imagined.
So you go the other way. Read the traces with no categories. Write one note per trace in your own words. Cluster the notes and let the categories come from the data. Keep reading until new traces stop producing new categories. Then triage every category with one of three labels: prompt-fix, evaluator-needed, or system-fix. The label tells you where the work goes.
The step people get wrong is saturation. Ten traces is not enough. Twenty is the minimum and thirty is better. The check is simple: read ten more, and if a new category appears you are not done.
You read the first five traces in the v0 file yourself. Answers about a different metric than the one asked. A summary that calls a rise from 31,981 to 48,040 a -4.5% change. A retention rate of 101671%. Most of them don’t fit “hallucination” or “SQL error” cleanly.
The practice is a failure taxonomy built from your notes on those five traces: named categories, each with a severity and a triage label, and the saturation rule you would use on the full file. Write your own categories before you look back at the worked answer.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→