Driver analysis: explaining variance and choosing what to change first
What is driving performance differences, and what should we change first?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
The AI Data Analyst scores 74 percent SQL correctness on its test set, just over a 70 percent blocking threshold. Your PM wants to know whether to release now or spend two more weeks on it. The overall number can’t answer that, because 74 percent looks the same whether the failures are spread evenly or packed into one kind of query.
Driver analysis splits the overall number into what each segment contributes. The impact score is the segment’s gap from the overall rate times its share of queries, so a segment that looks terrible but is 2 percent of traffic can rank below one that looks mild but is 30 percent. You sort by that score, and the top of the list is where engineering time does the most.
The findings are associations. A group of users can look worse only because they ask harder questions. So you split by two dimensions at once before blaming one factor, and you treat every finding as a hypothesis for the Week 5 experiments.
In the worked example you predict how simple and complex queries compare, then split the same test set by domain and by domain and complexity together. That second split shows an interaction effect: complexity hurts everywhere, and much more in one domain.
The practice works from the lesson’s own tables. You score the domain and each domain and complexity cell by impact, rank them, and fill an intervention hypothesis for the worst cell: the likely cause, the fix, the expected gain, and how you’ll check it worked. On the extended track you work out what the overall rate becomes if that cell gets fixed, and plan how you’d check the drivers month to month.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→