Week 5: Pipelines, Experiments, and Continuous Validation · Lesson 5.5

Monitoring for drift and regressions

How do we detect silent failures in production without evaluating everything?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome back. Last lesson ended with the rollout plan for v2.1, the fix for v2's latency and cost: shadow, canary, ring 1, full rollout, with a way back to v1 from every stage. Say v2.1 cleared every stage and it's live. The numbers in today's lesson are a worked scenario for that. Now a different question starts. The decision to roll out was right on the evidence you had that week. Is it still right three weeks later? Users change what they ask. Someone renames a column in the warehouse. The provider quietly updates the model behind your judge. Today is about watching a live system for those changes, which we call drift, when you can only afford to evaluate a small slice of what comes in. Let's start with the question you'd have to answer.

About this lesson

Say v2.1, the fix from 5.4, clears its rollout and is live. The evidence behind that decision starts to age. Users ask different questions, someone renames a column in the warehouse, the model behind your judge gets updated. Say the system handles 50,000 questions a day and you can afford to evaluate about 100, so the lesson is about which 100 to look at and what else to watch.

You weight the daily sample toward traces with warning signs: negative feedback, low confidence, question types that have been hard before. The weighted sample passes at a lower rate than a random one, which is what you want, and you report both numbers with what each one answers.

Then four checks, one per kind of drift. A stability index on the mix of question types for input drift. A distribution test on features of the generated SQL for output drift. A sentinel set of questions with known answers, run every week, for concept drift. And agreement between the judge and human labels on a fixed set, for judge drift. The sentinel example shows why a slow decline needs a check against the baseline: a week-over-week alert misses it because each step is small.

The practice is a monitoring plan for v2.1: sampling budget and weights, a detection method for each drift type, tiered thresholds, a table of who acts on which alert and how fast, and the response workflow.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→