Grounding evaluation in user value
Which of our technical metrics predict what users do, and how much weight should each one get?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Your dashboard says 92% SQL correctness, 87% retrieval accuracy and a 4.2% hallucination rate. The PM asks whether users find the AI Data Analyst useful, and you can’t answer, because none of those numbers has been checked against anything a user did.
This lesson splits metrics into three layers. Leading metrics like SQL correctness are fast and computed offline on a test set. Lagging metrics like edit rate, abandon rate and satisfaction come from real users in production. Business metrics like analyst time saved move slowest of all. A leading metric is a proxy, and it only earns a say in decisions once you’ve shown it predicts a lagging one.
You check that link with a correlation study on production traces, and the strength of the correlation decides the metric’s role. Above 0.7 it can block a release. Between 0.5 and 0.7 you track it and try to improve it. Below 0.5 you use it for debugging only. You predict which of three candidate metrics tracks satisfaction most closely, see an example result, and walk through one trace where the query ran, rows were missing, and the user was still satisfied. You also see metrics that are easy to log and predict nothing, and what happens when a team optimizes one of them.
The practice is a correlation table and a Leading-to-Lagging Metric Map. You put the six example metrics in one table with their correlations, intervals and tiers, decide which one could block a release, and fill in the map for one metric from your Lesson 1.3 failure taxonomy.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→