Metric design patterns for AI features
Which measurement pattern fits each part of this feature, and at what level do we measure it?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
The AI Data Analyst has a score for retrieval, one for SQL and one for the narrative, and you still can’t tell your PM whether it’s ready. Each metric was built by whoever owned that component, and no one decided which one can hold a release or what level to measure at.
This lesson gives you a starting point for each metric instead of a blank page. A measurement archetype is a reusable plan for one type of AI feature. It sets the unit of measurement, the kinds of quality that matter, the blocking and optimization metrics, and how scores roll up. Six archetypes cover most features: drafting, summarization, extraction, RAG, agents and decision support.
The AI Data Analyst is a hybrid, so each component gets its own archetype. Retrieval follows RAG and is measured on its own. SQL generation is checked by running the query and comparing the result with a known answer. The narrative follows summarization. You predict which unit of measurement fits the narrative, see why one end-to-end score can’t tell you which component to fix, and see why SQL correctness blocks the release while narrative conciseness gets tracked.
The practice fills in the archetype template for the AI Data Analyst. You pick the part of the question whose extraction errors would hurt most, justify a precision bar for refusal detection, and work out a ten-question session’s pass rate two ways, worst case and average, from the lesson’s 89 percent SQL correctness. The two land far apart.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→