AI Evals for Product DevelopmentL5.2 · 01

Test set strategy
and dataset lifecycle

How to iterate fast without fitting your evaluation to the data you measure on
Where we left offL5.2 · 02

Every run record in 5.1 stores a dataset version. What does that field actually protect you from?

What you built in 5.1
Run records with run_id, dataset_version, judge_version, sample size and sampling strategy.
Recall
What happens to a comparison between two runs if the data behind the same version name changed in between?
The scenarioL5.2 · 03

78% on the test set. 52% in production two weeks later.

Offline
200 curated test cases, collected six months earlier. Mostly simple lookups: "What was last month's revenue?"
Production
Users now ask multi-table joins and trend comparisons: "Compare conversion rates across three product lines by region."
The offline number measured a mix of questions that users had stopped asking.
Three ways it breaksL5.2 · 04

Evaluation data goes bad in three ways.

Staleness
The data ages out as user behavior, system capabilities or product requirements change.
Contamination
Test examples leak into development: into training data, into prompts, into judge calibration you have repeated too many times.
Versioning gaps
The dataset changed after an experiment and there's no record of which version the experiment used.
The lifecycleL5.2 · 05

A dataset moves through four stages, with a gate between each.

CreationSplit into train, dev and holdout before any metric work
DevelopmentIterate on train and dev only. Version every change.
Regression suiteDev graduates to a locked suite. Rotate on a schedule.
RetirementArchive with full history. Never delete.
You don't skip a stage, and you don't go backward.
Stage 1: creationL5.2 · 06

Split first, before you write a single judge prompt.

Train10%Few-shot examples for the judge. Small on purpose.
Dev45%Where you iterate. Refine judge prompts, compute agreement with human labels, calibrate thresholds.
Holdout45%, lockedRead-only until the final pre-release validation. One use per release cycle.
From a 200-trace sample of traces_v1_full: 20 train, 90 dev, 90 holdout, stratified by failure category.
LeakageL5.2 · 07

A holdout stops being a holdout the moment it shapes a decision.

Checking progress on it
Running the holdout after each prompt revision "just to see"
Copying examples
A holdout query pasted into a judge prompt or a few-shot example
Using its statistics
"The holdout has lots of multi-join queries, let's tune for those"
Reusing it forever
The same set used across release after release
PredictL5.2 · 08

Will the holdout agreement score be higher, lower, or the same as 0.82?

Your judge's agreement with human labels (Cohen's kappa) on a 150-query dev set rose from 0.65 to 0.82 over 12 prompt revisions in three months. You now run it on the holdout.
Write higher, lower or the same, with one sentence of reasoning, before we advance.
The generalization gapL5.2 · 09

Every revision you kept was chosen because it scored well on the dev set.

Dev set
Measures how well you tuned the judge to those 150 queries.
Holdout
Measures how the judge does on queries it never influenced.
The difference between the two is the generalization gap.
Stage 2: developmentL5.2 · 10

Every change to the dataset gets a version and a changelog entry, the way code does.

v1.020 / 90 / 90Initial split from a 200-trace sample of traces_v1_full, with a content hash
v1.140 / 90 / 90Added 20 multi-table join queries to train. Never to holdout.
nextlabel fixA corrected label is a new version too, with who changed it and why
Each manifest records name, version, date, split sizes, source dataset and what changed.
Stage 3: regression suiteL5.2 · 11

Once the judge settles, dev becomes a locked regression suite, and rotation starts.

On a schedule
Every quarter, retire the oldest 20% of the suite, or the examples that no longer match production. Replace them with a stratified sample from the latest production month.
On a signal
Rotate right away after a major product change, or when a drift check shows the suite no longer matches production.
The demo: is the suite still current?L5.2 · 12

Compare the suite to recent production on three dimensions. Which ones drifted?

Query complexityShare of simple lookups, multi-table joins, trend analyses and comparisons
User roleShare of PMs, data scientists, engineers and executives
Failure categoryThe mix of failure types the system produces
A Kolmogorov-Smirnov (KS) test on each dimension. p below 0.05 means the suite has drifted on that dimension.
Stage 4: retirementL5.2 · 13

Retire a dataset into an archive. Never delete it.

When to retire
Drift past the threshold for two quarters in a row, or a product change that makes the query patterns obsolete.
What the archive keeps
Version, creation date, last use date and reason for retirement. Read-only.
A Q2 release decision that cites suite v2.3 still has to be checkable in Q4.
PracticeL5.2 · 14

Write the Dataset Management Spec for the AI Data Analyst's regression suite.

Work out the split sizes and the changelog on paper, then fill in all six parts.
Lifecycle stageWhich of the four stages the dataset is in, and its current version
Split configurationRatios, sizes, and what the split is stratified by
Rotation policyCadence, what gets retired, what replaces it
Contamination checksWhich dimensions, how often, and what p-value triggers rotation
Holdout protectionWho can read it, when, and how many times
Retirement criteriaWhat drift or product change ends the dataset's life
Common mistakesL5.2 · 15

Four ways teams lose trust in their own evaluation data.

Mistake
What happens
Do this instead
Tuning on the full dataset
Kappa looks like 0.85 on the data you tuned on and drops to 0.68 on fresh data.
Split first. Report on data the judge never influenced.
Running the holdout every iteration
After ten peeks, the holdout is a second dev set.
One use per release cycle, then rotate it into dev.
Not versioning
A 5% improvement can't be reproduced after 30 examples were added and 10 labels changed.
Version every change. Experiments cite a version.
Never rotating
The suite measures last year's questions.
Check drift quarterly. Rotate when p is below 0.05.
Knowledge checkL5.2 · 16

Three judgment calls. Write your answer before you read on.

01Your suite is nine months old. KS p-values against production: 0.03 on query complexity, 0.42 on user role, 0.01 on failure category. Which dimensions drifted, and what do you do before using the suite for a release decision?
02You split 80% dev, 20% holdout. After five months of judge iteration, dev kappa is 0.85 and holdout kappa is 0.71. A colleague says the judge is bad. What is the likelier explanation?
03Your suite has 200 examples and production produces about 5,000 new traces a quarter. With quarterly rotation, how many do you retire, and how do you choose them?
Next lessonL5.2 · 17

Five rules keep the data honest. Next, you compare two versions of a system that answers differently every run.

01Split firstBefore any judge or metric work
02Version every changeIncluding label corrections. Experiments cite a version.
03Protect the holdoutRead-only, one use per release cycle
04Rotate on a cadenceQuarterly, or right after a major product change
05Check for driftCompare the suite to production and rotate when it drifts
AI ANALYST LAB · aianalystlab.ai