Translating evaluation signals to product actions
Given what we observed, what should we change next, and how will we know it helped?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
v2 was held in 6.1. Say the fix, v2.1, goes live. Two days later, in this worked scenario, every tracked metric is green and support tickets have tripled. Users say it takes forever and keeps asking them to clarify obvious questions. When user behavior and your evaluation metrics disagree like this, believe the users first, then find out what your metrics are not measuring.
You track four behaviors next to your automated scores: edit rate, reformulation rate, abandonment rate and escalation rate. A simple grid of behavior against metrics shows which case you are in. High behavioral signals with passing metrics means the metrics are incomplete, so you fix the measurement before you change the system.
In the demo you plot how much users edited each answer against the judge’s score. A cluster of answers the judge passed and users rewrote leads to one trace: a correct answer buried in 87 words of jargon. The judge scores faithfulness and completeness and has no dimension for readability.
From there you pick what to change. There are six places to change an AI system: prompts, retrieval, model config, UX constraints, guardrails and data quality. Each candidate fix gets a hypothesis (the change, the metric, how much it should move, why, and how you will check), a priority score of impact times confidence divided by effort, and staged validation: test set, shadow mode, a 10% experiment and full rollout, with a rollback if any stage fails.
The practice is a findings-to-actions plan for the v2.1 scenario, using the numbers and complaints in the lesson: explain why the dashboard missed the slow answers and the needless clarifying questions, then propose at least three ranked interventions with hypotheses and validation plans, plus at least one change to your metric suite.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→