Product evaluation framework for AI systems
What is the full evaluation surface for an AI feature, from inputs to user outcomes?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Knowing that a system fails one time in five does not tell you where. Was the SQL wrong? Did retrieval pull the wrong context? Did the summary make a number up? With the thin logs a v0 system writes, all three look the same from the outside.
The AI Data Analyst has six stages: query parsing, context retrieval, SQL generation, SQL execution, chart rendering and narrative synthesis. Each one fails in its own way, and you cannot test SQL correctness the way you test narrative accuracy. So you build a map organized by stage.
The map has three layers. The functional surface is what can go wrong when a regular user asks a regular question. The adversarial surface is what someone could do on purpose: prompt injection, jailbreaks, policy violations, private data leaks, resource exhaustion. The coverage gaps are what you cannot see yet: failures you suspect but have not observed, rare scenarios not in your sample, and behavior your logging cannot record at all.
The method is to predict failure modes from the architecture first, by looking at each stage’s inputs, outputs and dependencies, and then read real traces to confirm the predictions and find what you missed. You work through SQL generation as the example, then look at what the 500 v0 traces can show. The visible failures are in SQL and the narrative. Retrieval never appears, because v0 logs nothing about it.
The last step splits the map into what you can evaluate with today’s logging and what needs better instrumentation first. That second list becomes your Week 2 priorities.
The practice is your own evaluation surface map: failure modes for SQL execution and chart rendering predicted from the architecture, each attack type matched to the stages it hits, at least three coverage gaps, and the evidence split.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→