What AI evaluation is and why it requires a different approach
Why don't the evaluation methods I already know work for AI-powered product features?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Your team launches an AI Data Analyst. It takes a question in plain English, writes SQL, runs it, and hands back a chart and a short summary. It passes code review. It works in the demo. Then say that in the first week about 15% of users report that the same question gives different answers, and engineering finds nothing wrong.
Nothing is broken. A language model samples every answer, so the same input can come back different with every component doing what it was designed to do. The testing habits you already have, run it once, assert the exact output, release, do not carry over.
You look at one query run five times, a made-up example, and then at the 47 questions the course’s v0 traces hold five runs of. Then you put two numbers on it. pass@k asks whether at least one of k runs succeeded. That tells you the system is capable. reliable@k asks whether all k runs succeeded. That tells you what your users experience. In the v0 traces pass@5 is 1.0 and reliable@5 is 0.36. The gap between the two is the number a product manager needs to see, and its size tells you whether you have a consistency problem or a capability problem.
You also see why temperature zero reduces variance without removing it, why a benchmark score screens a model but does not make the ship decision, and why “evaluation” means three different things to a PM, a data scientist and an engineer. The course treats it as evidence for four decisions: ship, ramp, hold, or roll back.
The practice is a Non-Determinism Report you do from the slides: pass@5 and reliable@5 for the five-run demo query and for the 47 repeated v0 questions, where each result lands on the gap table, and two sentences on what the gap means for the ship decision.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→