Metric specifications, thresholds, baselines, and release criteria
What does good enough to release mean in numbers, written down before we see the results?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
v2 of the AI Data Analyst improves SQL success and retrieval, loses a little narrative quality, and runs slower. Engineering thinks it’s great, design sees a regression, and data science is worried about latency. Everyone is reading the same numbers, and no one agreed in advance what they mean or where the lines are.
This lesson writes that agreement down. A metric spec pins each metric to one definition in eight parts: the calculation, the unit of analysis, the population, the segments, sampling, aggregation, ownership and versioning. Release criteria then put each metric in one of three classes. Blocking metrics must pass. Guardrails are blocking metrics that protect against harm, like cost and latency. Optimization metrics are tracked and don’t gate the release.
Thresholds follow four patterns: an absolute floor, a limit on how far a metric can drop from the last version, a check that an improvement isn’t noise, and a floor each user role has to clear. All of them get set from the v1 baseline and the business needs before anyone looks at v2’s numbers. Drawing the line after you’ve seen the result guarantees a pass.
In the worked example, built on the lesson’s scenario baselines, you predict where to put the SQL success threshold given v1’s baseline and its interval, then see how segment baselines and a stricter correctness check shape the final thresholds.
The practice writes the spec for narrative quality, classifies your metric inventory, writes the release gate as a set of conditions, runs the v2 scenario through it, and documents one conflict between two metrics. On the extended track you read the segment intervals to decide which segments need their own threshold, and check that v1 passes its own gate.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→