AI Evals for Product DevelopmentL5.6 · 01

Evaluation automation,
end to end

Connecting the offline pipeline, the rollout gates and production monitoring into one system
Where we left offL5.6 · 02

Gradual drift and sudden drift need different detectors. Which catches which?

Sudden drift
SQL success rate falls from 0.82 to 0.68 overnight and crosses the blocking threshold.
Gradual drift
SQL success rate slides from 0.82 to 0.78 over three weeks. No single day crosses the threshold.
Answer from memory before you advance.
The problemL5.6 · 03

A daily batch job finds a Friday afternoon regression on Saturday morning.

Friday 2 PMModel update goes outSQL success rate drops to 0.68
Saturday 6 AMBatch job runsThe report is generated
Saturday 9 AMSomeone reads itThe PM opens the email and starts looking
All that timeUsers see bad answersThe team doesn't know yet
The logic from 5.5 is correct. The schedule is the problem.
Three limitsL5.6 · 04

Batch monitoring has three operational limits.

Reporting delay
You see today's problem tomorrow.
Aggregate-only views
When a number drops, finding the failing requests means writing new code.
No instant alerting
A threshold breach waits for the next report. No one is paged.
The whole systemL5.6 · 05

One loop: evaluate offline, gate each rollout step, monitor production, feed what you learn back.

From 5.1Offline pipelineSample, score and store every run with its metadata
From 5.4Gate checkBlocking metrics checked before each rollout step
From 5.5Production monitoringThresholds for sudden drops, a detector for slow drift
From 5.2Back into the test setNew failure patterns become new test cases
Every piece reports into one decision: ramp, hold or roll back.
Step 1: ingestionL5.6 · 06

Send every trace to one place the whole team can query.

Production loggerOne trace per request: question, SQL, result, latency, errors
Platform SDKA few lines added to the logger. Sends in 60-second batches
Trace storageSearchable by time, segment and metric
Sampling rate100% during a ramp, 1 to 5% once stable
Step 2: dashboardsL5.6 · 07

Put your own metric specs on the dashboard.

What the platform shows first
Token counts, model latency, error rate. Generic numbers every customer sees.
What you configure
SQL success rate with its blocking threshold. P95 latency with its guardrail. Both split by user segment.
The platform displays metrics. Your metric specs define them.
Step 3: alertsL5.6 · 08

Blocking alerts page someone now. Awareness alerts wait for business hours.

Blocking alert
SQL success rate below 0.75 for 3 hours, or P95 latency above 2,500 ms for 1 hour. Page on-call through PagerDuty, Slack or SMS.
Awareness alert
Cost per query up 10% over 24 hours, or a new query pattern appears. Email or a Slack channel, reviewed during the day.
PredictL5.6 · 09

Where should the alert line go?

SQL success rate is stable at 0.82. On normal days it moves up or down by about 0.03. Do you alert on any day below 0.82, or somewhere lower? If lower, where, and why?
Treat 0.03 as one standard deviation of the daily value. Write your answer before you advance.
The revealL5.6 · 10

Set the line from normal variation, so an alert means the drop is probably real.

Alert on any drop below the baseline
Normal days fall below the average often. The alert fires on noise and the team learns to ignore it.
Alert below a statistical threshold
Fire only when a day this low would be unusual under normal variation. Choose 95% or 99% confidence.
Step 4: drift detectionL5.6 · 11

A drift detector runs all the time and flags a shift that never crosses the threshold.

BaselineThe stable week after rollout, when SQL success averaged 0.82
Current windowThe last 7 days, compared to the baseline every day
Drift scoreHow different the two weeks' patterns look. Alert above 0.05
The threshold alert watches the level. The drift detector watches the shape.
The ramp decisionL5.6 · 12

"Can we ramp from 10% to 50%?" The dashboard answers it in one view.

SQL success rate7-day value against the 0.75 blocking threshold
P95 latency7-day value against the 2,500 ms guardrail
Drift scoreCurrent week against baseline, 0.05 alert line
By segmentExecutives and PMs side by side
Same evidence as the daily report. Faster to read, and anyone can open it.
PracticeL5.6 · 13

Wire up monitoring for one AI feature.

Platform version
You need a free account on a monitoring platform and some logged traces with a timestamp, a success flag and latency. Load 100 traces, chart SQL success and P95 latency by segment, set one blocking and one awareness alert, send a test alert, add a drift detector with a 7-day window.
Design-only version
No account or traces needed. Write the spec: sampling rate during and after a ramp, the dashboard metrics and thresholds, each alert with its duration and channel, and the drift detector's baseline, window and alert line.
Common mistakesL5.6 · 14

Three ways automated monitoring goes wrong.

Mistake
What happens
Do this instead
Thresholds too tight
Dozens of alerts in the first week. The team mutes the channel and a real regression gets lost.
Set lines from normal variation and require the metric to stay bad for a few hours.
Retiring the scripts
The platform goes down during a deployment and there is no monitoring at all.
Keep the 5.5 scripts running as the fallback.
Uniform 1% sampling
A small, high-stakes group produces almost no sampled traces, so its failures stay invisible.
Sample high-risk groups at 100% and the rest at 1%.
Knowledge checkL5.6 · 15

Three judgment calls. Write your answer before you read on.

01SQL success rate fell from 0.82 to 0.79 over 3 hours. The blocking alert is set at 0.75 and it didn't fire. Is the alert broken?
02The drift detector fires on day 15. You find that a new question type, year-over-year comparisons, became common that week. Real drift or a false alarm?
03You must catch regressions within 4 hours of a deployment. One tool has stronger drift detection, the other has faster alert delivery. Which feature matters more here?
Next lessonL5.6 · 16

Next, you run the pipeline by hand, from sampling to the alert.

Offline pipeline
Sample, score, store each run with metadata
Gate check
Blocking metrics before each rollout step
Production monitoring
Traces flowing in, dashboards on your metrics, two kinds of alerts, a drift detector
The decision
Ramp, hold or roll back, from evidence that is current
AI ANALYST LAB · aianalystlab.ai