AI Evals for Product DevelopmentL6.5 · 01

Evaluation cadence
and governance

Who looks at the evaluation system, how often, and what happens to the work that gets pushed back
Where we left offL6.5 · 02

You made a call on v2. Who watches the metrics once a release goes out?

What you decided
A recommendation on the v2 change, backed by experiment results and a named owner for the decision.
What the decision doesn't cover
Who checks the metrics next week, next month and next quarter, and who acts when one of them moves.
Write down a role, and how often that person would look.
The scenarioL6.5 · 03

Six months after launch, the SQL judge had drifted and nobody had checked.

LaunchSay v2.1 goes live. The SQL correctness judge agrees well with human reviewers: Kappa 0.72.
Month 2The judge starts drifting. No review looks at it.
Month 4The judge is approving SQL that reviewers would reject. Still nobody looks.
Month 6A production incident, with Kappa down to 0.48, forces an emergency fix.
"The monthly review got cancelled twice and never rescheduled."
Three tiersL6.5 · 04

Different problems show up at different speeds, so review at three speeds.

Weekly check52× a year · 45 minAcute problems: something is broken right now. An error spike, a latency jump, a failed must-pass metric.
Monthly review12× a year · 90 minSlow degradation: a judge drifting, a user segment getting worse, test coverage thinning out.
Quarterly audit4× a year · half dayWhether you are still measuring the right things for what users need.
Tier 1L6.5 · 05

The weekly check flags problems and hands each one to someone.

Automated metric review
15 min
Latency, cost, error rates, and the must-pass metrics that block a deploy
Dashboard review
20 min
Is anything outside its acceptable range, or creeping toward it?
Incident triage
10 min
New failure modes and user reports that need investigating
OwnerAI Reliability Lead, with the on-call engineer
OutputEach issue flagged for deeper analysis or for immediate rollback
Tier 2L6.5 · 06

The monthly review looks at how things moved over 30 days.

Metric trend analysis
30 min
Improving, stable or degrading over the last 30 days?
Judge calibration check
30 min
Sample 20 real user queries. Do the judge's scores still match a human reviewer's?
Segment health review
20 min
Which user groups or query types are getting worse?
Evaluation debt review
10 min
What work got deferred, and how risky is it now?
OwnerPM, data scientist, ML engineer and AI Reliability Lead
OutputUpdated thresholds, scheduled recalibration, or a risk escalated to leadership
Tier 3L6.5 · 07

The quarterly audit asks whether you are measuring the right things.

Regression suite refresh
90 min
Remove test cases that no longer look like real use. Add failure modes found in production.
Full judge recalibration
90 min
Run the judge on 200+ held-out examples with known answers. How often does it catch bad SQL, and how often does it pass good SQL?
Experiment backlog
60 min
Which improvements are worth testing next?
Evaluation system health
60 min
Do the metrics still line up with what users need?
OwnerThe extended team, with the VP of Engineering and Head of Product
OutputMetric changes, tooling updates, roadmap adjustments
OwnershipL6.5 · 08

Each review gets one Accountable owner, and only one.

Activity
AI Rel. Lead
DS
PM
ML Eng
VP Eng
Weekly quality check
A
R
I
I
I
Monthly judge calibration
R
A
I
R
I
Monthly segment health
C
R
A
C
I
Monthly debt review
C
R
A
C
I
Quarterly system health
R
R
R
R
A
R does the work and A owns the outcome. C gives input, and I gets updates.
Evaluation debtL6.5 · 09

Deferred evaluation work goes in a register, with a risk level and a date.

Deferred work
Risk
Deferred
Target paydown
Judge recalibration, SQL correctness
Medium
2025-01-15
2025-02-28
Instrumentation for a new failure mode: multi-turn drift
High
2025-01-08
2025-02-15
Test queries still written for the Q4 2024 schema
Medium
2025-01-22
2025-02-28
Reviewed every month: pay it down, re-date it, or accept the risk. No more than 10 open entries.
PredictL6.5 · 10

A five-person team, tight deadlines. Which tier gets skipped first?

Weekly check
45 minutes, every week
Monthly review
90 minutes, once a month
Quarterly audit
Half a day, with leadership
And if that tier is skipped for two months, what is the first symptom the team would see? Write both down.
The replayL6.5 · 11

Replay the drift with a monthly review in place.

LaunchJudge agrees well with human reviewers.
Month 1Monthly check on 20 real queries. Slight drift, still in range.
Month 2Monthly check falls below the team's warning line. Debt entry: full recalibration needed.
Month 3Quarterly audit recalibrates the judge on 200+ examples and fixes the rubric.
What actually happened: the monthly reviews were cancelled and the quarterly audit slipped.
The costL6.5 · 12

Can a five-person team afford this?

Tier
Time each time
How often
Weekly check
45 min
4 times a month
Monthly review
90 min
Once a month
Quarterly audit
300 min
Once every 3 months
Work out the monthly total, then the load on the person who attends all three.
Cadence and the ship decisionL6.5 · 13

How fast you can detect a regression changes how much uncertainty you can ship with.

Weekly check and a clear escalation path
A regression shows up at the next weekly review and can be rolled back quickly. You can ship with a confidence interval that still overlaps the baseline.
Next review in three weeks, or never
A regression reaches users for weeks before anyone sees it. Hold until the evidence is much stronger.
Common mistakesL6.5 · 14

Four ways a cadence breaks down.

Mistake
What happens
Do this instead
Skipping the monthly review
Slow drift stays invisible until it turns into an incident.
Keep it. If time is tight, cut it to 60 minutes.
Meetings with no one who acts
The same regression gets discussed every week and nothing changes.
Every flagged issue gets an owner and a deadline, or it escalates.
A debt register nobody pays down
It grows past anyone's interest and gets abandoned.
Cap it at 10. Pay down, re-date or accept.
A cadence for a bigger team
It's thorough on paper and nobody can keep it up.
Fit it to the team you have. Add reviews as the team grows.
PracticeL6.5 · 15

Build your Evaluation Cadence Package.

Base version, everyone
Fill the cadence calendar: 11 activities across the three tiers, each with a duration, owner and output. Fill a RACI for six activities across five roles, one A per row. Write three debt register entries, at least one High.
Extended version, DS and engineering
Write a three-level escalation playbook: tactical, cross-functional, strategic. Then run a skip simulation: what builds up if Tier 2 is skipped for two months, how would you find out, and how long would the fix take?
Knowledge checkL6.5 · 16

Three judgment calls. Write your answer before you read on.

01A team runs a weekly 30-minute check and a quarterly 4-hour audit, with nothing in between. At the audit, their SQL judge's Kappa is 0.48, down from 0.72 at launch. Which Tier 2 activity would have caught it earlier, and how often should it run?
02The data scientist is Accountable for judge calibration. A calibration check fails while they're on vacation. Who decides whether to pause deploys, and how would you change the RACI so this can't stall?
03Your debt register has 8 open entries: 3 Low, 4 Medium, 1 High. Leadership asks what risk you're carrying. What do you tell them, and which debt do you pay down first?
Next lessonL6.5 · 17

Next: reporting what the evaluation found so people act on it.

This lesson
Weekly, monthly and quarterly reviews, one Accountable owner for each, and a debt register for the work that gets pushed back.
Lesson 6.6
Claims that match the strength of the evidence, reporting segments and tails as well as averages, and the same results written for an executive, a PM and an engineering team.
AI ANALYST LAB · aianalystlab.ai