AI Evals for Product DevelopmentL5.5 · 01

Monitoring for drift
and regressions

How to catch silent failures in production without evaluating every query
Where we left offL5.5 · 02

Three weeks after rollout, how would you know the decision still holds?

What you had at launch
An experiment, a passed gate check, a rollout plan with rollback triggers.
What you need now
Evidence, every week, that the conditions behind that decision haven't changed.
Write one sentence: what would you look at?
The scenarioL5.5 · 03

The metrics look fine, users are complaining, and you can't evaluate everything.

50,000queries a day through the AI Data Analyst
$2,000a day to run the full judge suite on all of them, plus 800 ms per request
100traces a day you can actually afford to evaluate
Which 100 you pick decides what you're able to see.
Four kinds of driftL5.5 · 04

Four different things can move after launch, and each needs its own check.

Input driftThe mix of questions users ask changes.PSI on question types
Output driftThe system answers the same kinds of questions differently.KS test on SQL features
Concept driftThe right answer to the same question changes.Sentinel set with known answers
Judge driftYour judge starts scoring differently.Agreement on a fixed set over time
Sudden and gradualL5.5 · 05

A threshold catches a sudden drop. A slow slide needs a comparison to the baseline.

Sudden drift
SQL success rate falls from 0.82 to 0.68 overnight. Blocking threshold is 0.75.
Gradual drift
SQL success rate slides from 0.82 to 0.78 over three weeks. No single day crosses 0.75.
Which check catches which, and which one never fires?
Signal-based filteringL5.5 · 06

Spend the 100 traces where the warning signs are.

Signal on the traceWeightWhy it's informative
User gave negative feedback3xA thumbs-down or a reported issue
Low confidence from the system2xThe system was unsure of its own answer
Failure-prone segment2xQuestion types that have been hard before, like multi-turn
Everything else1xStill sampled, so you keep a view of normal traffic
PredictL5.5 · 07

A random sample of 100 passes at 85%. Will the weighted sample pass higher, lower or about the same?

Pick one and write the reason. Then: which number describes the system's quality?
The answerL5.5 · 08

The weighted sample passes lower. That tells you the weights are doing their job.

Random sample
An estimate of quality across all traffic. Use it for "how good is the system?"
Weighted sample
A view of the risky traffic. Use it to find failures and to watch the trend.
Report both, and say which question each one answers.
Input driftL5.5 · 09

Did the mix of questions change enough to call it drift?

Question type
Days 1 to 14
Days 22 to 28
Lookup
25%
22%
Trend
30%
24%
Comparison
25%
35%
Aggregation
20%
19%
PSI < 0.1 stable
0.1 to 0.25 minor drift, keep watching
> 0.25 major drift, investigate
Output driftL5.5 · 10

Is v2.1 writing different SQL than it did in the baseline weeks?

What you measure
The number of JOIN, WHERE, GROUP BY and ORDER BY clauses in each generated query, as a rough measure of how complex it is.
How you compare
A two-sample KS test: baseline days against the latest week. Alpha set at 0.01.
A significant result says the shape changed. It doesn't say whether the change is bad.
Concept driftL5.5 · 11

Run 50 questions with known answers every week. Which alert fires first?

Week-over-week alert
Fires when correctness drops 10 points or more from the previous week.
Baseline alert
Fires when correctness is 10 points or more below the first week after launch.
The sentinel set: 50 canonical questions, with correct results you've already verified.
Judge driftL5.5 · 12

Check the judge before you trust anything the judge told you.

What you check
Each week, run the judge on a fixed set with human labels and compute Cohen's kappa. Recalibrate if it falls below 0.7.
What you pin
A dated model version for the judge, so a provider update can't change your scoring without you knowing.
The monitoring planL5.5 · 13

An example plan: who acts on each alert, how fast, and what they do.

Condition
Action
Owner, example
Respond within, example
One drift type fires, blocking metrics fine
Investigate the segment
On-call analyst
1 business day
Concept drift on the sentinel set
Find the cause, fix context or roll back the change
Engineering lead
4 hours
Several drift types in the same window
Escalate, consider ramping down
Engineering manager
1 to 4 hours
A blocking metric fails its gate
Roll back to v1
Automatic, then on-call
Immediately
The owners and times are made up to show the shape. You fill in your own with your team.
PracticeL5.5 · 14

Write the monitoring plan for v2.1.

Base version, everyone
Sampling: budget and signal weights. Detection: a method for each of the four drift types. Thresholds, with log, alert and escalate tiers. The decision table: condition, action, owner, response time. The response workflow from alert to post-incident review.
Extended version, DS and engineering
Work out PSI by hand on the table from slide 9. Design a semantic drift check on the questions. Write how you'd test whether your signal weights really point at failures.
Common mistakesL5.5 · 15

Three ways monitoring stops telling you the truth.

Mistake
What happens
Do this instead
Signal weights set too high
Negative feedback at 10x fills the sample with the worst cases. The weighted pass rate drops far below the random one and the team chases a problem that isn't there.
Cap weights around 5x. Keep a random sample running next to the weighted one.
Thresholds too low
Normal weekly variation trips the alert every Monday. By the time real drift arrives, nobody reads the channel.
Use tiers: log above 0.1, alert above 0.2, escalate above 0.25.
Unpinned judge
The provider updates the judge model, it scores more leniently, and the pass rate goes up while real quality falls.
Pin a dated version. Recalibrate before any judge change.
Knowledge checkL5.5 · 16

Three judgment calls. Write your answer before you read on.

01Say weekly quality scores slip about one point a week for twelve weeks. Your alert fires on a drop of 5% or more week over week. It never fires. What went wrong, and what check would have caught it?
02The KS test on SQL clause counts gives p = 0.003 with alpha at 0.01. Does that mean you should act? What would you check next?
03The weighted sample passes at 78% and the random sample at 85%. A colleague says true quality is 78% because the weighted sample is more thorough. What's wrong with that?
Next lessonL5.5 · 17

Four checks, one plan. Next, they run continuously instead of once a day.

Input driftPSI on question typesPlus a look at whichever category moved
Output driftKS test on SQL featuresA changed shape is a reason to check correctness
Concept driftSentinel set, against the baselineCatches the slow decline a weekly step alert misses
Judge driftKappa on a fixed labeled setTells you whether to trust the other three
AI ANALYST LAB · aianalystlab.ai