AI Evals for Product DevelopmentL6.4 · 01

Ownership
model

The AI Reliability Lead, decision rights, escalation triggers and evaluation debt
Where we left offL6.4 · 02

Who on your team signs off on the ship decision, and what evidence would they need?

Who signs off
The PM owns the decision. The ML engineer owns model quality. The product engineer owns deployment.
What they need to trust
Metric specs, experiment results, the judge calibration report, the monitoring plan.
Who built each piece of evidence, and who checked it since?
The problemL6.4 · 03

An evaluation system with no owners decays, and nothing announces it.

At launch
The judges are calibrated, the regression suite covers every query type, someone watches the dashboard, and the thresholds are written down.
Six months later
The judges have not been checked since launch, the suite misses new query types, the dashboard has no watcher, and the thresholds sit in a Slack thread.
PredictL6.4 · 04

Your SQL judge was calibrated six months ago. What happened to its false positive rate?

At launch
The system runs on GPT-4 and writes correct SQL 60% of the time. The judge is calibrated on that output: 89% true positive rate, 6% false positive rate.
Two months ago
The system moves to GPT-4o and writes correct SQL 78% of the time. Nobody has checked the judge since launch.
False positive rate: how often the judge flags correct SQL as wrong. Write your prediction before we advance.
The revealL6.4 · 05

Nobody can say, and the judge's flags got less trustworthy either way.

If the rate held steady
More of the SQL is correct, so a larger share of the judge's flags land on correct SQL.
If the rate moved
GPT-4o writes SQL differently from the output the judge was calibrated on. Only recalibration tells you how far it moved.
The six-month auditL6.4 · 06

What the audit found on the AI Data Analyst.

Component
What exists
What's wrong
Four LLM judges
SQL correctness, retrieval relevance, narrative faithfulness, hallucination detection
Last calibrated at launch
Regression suite
80 examples from launch
Three query types added since, none covered
Monitoring dashboard
SQL success rate, retrieval latency, query volume
Nobody assigned to watch it
Metric thresholds
A Slack thread from four months ago
Nobody is sure what they are
The two code-based metrics, execution success and latency, were fine.
Three things to ownL6.4 · 07

Someone owns the data, someone owns the metric, and someone owns the decision.

RowThe dataThe pipelines and traces that produce every record you evaluate. Is the data arriving, and is it complete?
ColumnThe metricWhat each metric means and whether it's measured correctly. Judges, thresholds, the regression suite.
DecisionShip and roll backWho can say ship, ramp, hold or roll back, and who can stop a release that's already out.
The AI Reliability LeadL6.4 · 08

One person owns the health of the evaluation system and the risk posture of each release.

What the lead watches
Every activity has an owner. Debt items move by their dates. Escalation triggers fire and reach the right person.
What the lead brings to the ship review
How much the evidence can be trusted today, and which known gaps sit under the decision.
A role someone wears. On a small team it's usually one of the people already in the matrix.
RACIL6.4 · 09

Every evaluation activity needs one Accountable owner, and only one.

R · Responsible
Does the work.
A · Accountable
Owns the outcome and makes the final call.
C · Consulted
Gives input before the decision.
I · Informed
Hears about it after.
Zero A's and nobody owns it. Two A's and nobody can settle a disagreement.
The desired stateL6.4 · 10

The AI Data Analyst team, with one A per row.

Activity
PM
ML Eng
Domain Expert
Product Eng
QA
Instrumentation design
C
R
C
A
I
Judge calibration
I
A
C
I
R
Threshold setting and review
A
C
R
I
I
Regression suite maintenance
C
R
I
A
R
Production monitoring
I
R
I
A
I
Ship, hold and roll back
A
C
C
I
I
Evaluation debtL6.4 · 11

The debt register turns "we should recalibrate sometime" into an item with an owner and a date.

Item
DEBT-001: four LLM judges uncalibrated since launch, six months ago
Risk
High
Impact if unaddressed
False positive rate unknown. Release decisions rest on evidence nobody has checked.
Owner
ML Engineer
Target date
March 15
Status and notes
Open. Schedule a two-day calibration sprint.
Escalation triggersL6.4 · 12

Decide ahead of time what escalates, to whom, and how fast.

Trigger
Severity
From
To
Response
Safety metric fails in production
Critical
On-call engineer
PM and ML Eng
Within 1 hour
Primary metric degrades more than 5% within 7 days
High
ML Engineer
PM
Within 24 hours
Judge true or false positive rate shifts more than 10%
High
ML Engineer
PM
Within 48 hours
Experiment shows a segment regression
Medium
Data Scientist
PM
Before the decision
PracticeL6.4 · 13

Design the ownership model for your AI Data Analyst.

Base version, everyone
Start from the RACI on slide 10 and add at least three activities, one A per row. Use the audit on slide 6 and your own work from Weeks 2 to 5 to list at least five debt items. Put them in a register with risk, owner and target date. Define at least three review types and three escalation triggers.
Extended version, DS and engineering
Estimate remediation cost in engineering weeks. Model what fails if the top three debt items stay open. Design a dashboard for the health of the evaluation system itself.
Common mistakesL6.4 · 14

Three ways an ownership model fails.

Mistake
What happens
Do this instead
Everyone is R, nobody is A
Each person assumes someone else has the judge drift. Six months later nobody can say whose job it was.
One A per row, and that person agrees it's theirs.
A register nobody reviews
Twelve medium-risk items from six months ago, no dates, no status updates.
Review it on a fixed schedule, and escalate high-risk items past their date.
No escalation path for alerts
SQL success rate slides for three days and whoever notices it decides what to do.
Triggers with a severity, an owner and a response time.
Knowledge checkL6.4 · 15

Three judgment calls. Write your answer before you read on.

01Your RACI lists the ML engineer as both R and A for judge calibration, and the domain expert as A for threshold setting. What's wrong, and how do you fix it?
02The SQL judge hasn't been recalibrated since launch, six months ago. You mark it low risk because it was accurate at launch and the SQL metrics look stable. A teammate says high risk. Who's right?
03SQL success rate dropped from 74% to 68% over five days. Your trigger says escalate to the PM if the primary metric degrades more than 5% within 7 days. Escalate now, wait two days, or investigate first?
Next lessonL6.4 · 16

The ownership model has three layers. Next, how often each review runs.

RACI matrixWho does what, with one Accountable owner per activity.
Debt registerWhat needs fixing, with a risk level, an owner and a target date.
Escalation triggersWhat gets escalated, to whom and how fast, with the AI Reliability Lead watching the whole system.
AI ANALYST LAB · aianalystlab.ai