Decision-making under uncertainty
Given conflicting evidence, what decision is justified?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
By Week 6 you can tell whether a change made the AI Data Analyst better or worse. That doesn’t make the decision for you. In the course’s experiment, v2 raised SQL success by 2.7 points, task completion by 5.4 and retrieval precision by 7.9. It also raised average latency 17.6 percent and cost per query 20 percent, past both guardrails, and the share of users averaging over the 2-second SLA went from 1.6 to 10.6 percent. Your PM wants to know if it goes out anyway.
With only “ship” and “don’t ship” to choose from, a result like that turns into an argument. This lesson gives you six decision types instead: ship, ramp, hold, roll back, scope-restrict, and deploy with a human in the loop. Each one needs a different strength of evidence.
To pick between them you ask four questions. Which way did each metric move? By how much? How much can you trust the numbers? And if it goes wrong after deploy, how quickly would you find out? Then you write the tradeoff down in one sentence: what you recommend, what improved, what got worse, and what limits the risk.
You work through v2 by user group. Every group gained, executives least, and every group paid the same latency and cost, so there’s no group to ship to and scope-restrict is out. The slowness lands on users who ask complex questions, 31 percent of whom average over 2 seconds. The evidence points to hold: v2 stays off while the team fixes latency and cost, starting with complex questions.
The practice is a decision memo for v2 with six fields: recommendation, primary evidence, tradeoff acknowledgment, risk mitigation, rollback trigger and decision confidence. You answer confidence and risk containment from the numbers in the lesson and check that every field names a metric and a number. The extended version works out the daily cost at 100,000 queries, checks whether any user group stays inside both guardrails, and writes a monitoring spec for v2.1.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→