Prioritization and iteration using evaluation evidence
What should we fix first, and how do we learn fast?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
By Week 6 you have a lot of evidence about what is broken in the AI Data Analyst: a failure taxonomy, a driver analysis, experiment results and a findings-to-actions plan. What you do not have yet is an order. Formatting issues hit 15 percent of queries and barely hurt anyone. Policy violations hit under one percent and can start a compliance escalation. Without a method, the fix order comes down to whatever is easiest or whoever argues loudest.
This lesson scores every item in the backlog on seven dimensions, one to five. Four are about impact: user harm, frequency, business criticality, and confidence in the evidence. Three are about velocity: fixability, time to learn, and reversibility. Every score has to cite the artifact it came from. The example weights count user harm twice and time to learn one and a half times, and you can change them as long as you write down why.
You work through a demo backlog, predict which of two items should go first, and then check your pick against the scores. Two checks sit on top of the ranking. Any item with catastrophic harm gets a critical-failure flag, so a rare failure cannot sink to the bottom. And the top item gets a segment check: is it at least twice as bad for any group of users?
Every item near the top also gets acceptance criteria: which metric moves and by how much, which segment has to see it, and what evidence closes the ticket.
The practice is a ranked backlog: the eight failure modes from the lesson, or 8 to 10 items from your own product, scored with citations, weighted and ranked, at least one tail risk flagged, a segment check on the top item, and acceptance criteria for the top five. The extended version reworks the demo scores by hand with the harm weight doubled and then the time to learn weight doubled, to see which items move.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→