Metric strategy: blocking metrics vs optimization metrics
Which metrics should be able to hold a release on their own, and why?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
By Week 4 the AI Data Analyst has twelve evaluation metrics, and your PM still can’t get an answer to whether v2 should go out. Twelve good numbers don’t make a decision on their own, because the team hasn’t said which ones matter for a release.
This lesson gives every metric one of two jobs. A blocking metric has a threshold, and if it misses, the release waits. An optimization metric has a target, and missing it starts work without holding the release. To sort them you ask one question of each metric: if this one fails and everything else passes, do we release?
The metrics also get a shape. User trust sits at the top, tracked through a proxy. Four drivers sit under it: answer correctness, response latency, cost per query and multi-turn coherence. The metrics you compute from traces sit under those. When a driver drops, the tree tells you where to look first.
Every blocking threshold needs a written reason, meaning a sentence on what breaks below that number. The reasons usually come from one of four places: the output has to work at all, a cost or revenue limit, what users already have today, or the point where people give up. You predict how many of the twelve metrics should be blocking, then compare your number with a worked classification.
The practice is a metric system spec. You classify four metrics that are hard to call, write release criteria for a v2 retrieval change, and, on the extended track, put numbers on your threshold reasons and test one tradeoff: recall up 5 points, latency up 20 percent.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→