Monitoring for drift and regressions
How do we detect silent failures in production without evaluating everything?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Say v2.1, the fix from 5.4, clears its rollout and is live. The evidence behind that decision starts to age. Users ask different questions, someone renames a column in the warehouse, the model behind your judge gets updated. Say the system handles 50,000 questions a day and you can afford to evaluate about 100, so the lesson is about which 100 to look at and what else to watch.
You weight the daily sample toward traces with warning signs: negative feedback, low confidence, question types that have been hard before. The weighted sample passes at a lower rate than a random one, which is what you want, and you report both numbers with what each one answers.
Then four checks, one per kind of drift. A stability index on the mix of question types for input drift. A distribution test on features of the generated SQL for output drift. A sentinel set of questions with known answers, run every week, for concept drift. And agreement between the judge and human labels on a fixed set, for judge drift. The sentinel example shows why a slow decline needs a check against the baseline: a week-over-week alert misses it because each step is small.
The practice is a monitoring plan for v2.1: sampling budget and weights, a detection method for each drift type, tiered thresholds, a table of who acts on which alert and how fast, and the response workflow.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→