AI Evals for Product DevelopmentL5.1 · 01

Evaluation pipeline
architecture

Stages, run metadata, and how much evaluation to run at each deployment stage
Where we left offL5.1 · 02

Which of your metrics were blocking, and what threshold did you put on each?

Blocking metrics
The ones where a miss stops the release, whatever else improved.
Thresholds
The line for each one, and where that number came from.
Answer from memory before you move on.
The problemL5.1 · 03

Two numbers in Slack. Did quality drop?

Monday, after a normal runSQL correctness: 87%
Wednesday, after a prompt changeSQL correctness: 83%
Which dataset?
Which judge and settings?
Which model version?
How to rerun it?
The shiftL5.1 · 04

An evaluation pipeline keeps every run, with enough context to compare it.

One-off script
Someone runs it when they remember. The number goes to Slack. There's no history, so two runs can't be compared.
Pipeline
Defined stages. Every run is stored with its metadata. Anyone can query the history and see what changed between runs.
The stagesL5.1 · 05

Six stages, each with a defined input and output.

SamplingChoose which traces to evaluate
JudgingScore each one: oracle checks, LLM judges
AggregationRoll up to pass rates, means, P95s, segments
StorageWrite results and run metadata to a table
ReportingMetric cards and comparisons from stored runs
AlertsNotify when a metric crosses a threshold
Run metadataL5.1 · 06

Every run records how it was done.

run_id
run_a_7f3c2e10
dataset_version
traces_v1_full 1.0.0
model_version
v1
judge_version
judge_v2.1
judge_config_hash
a3f5b8c2
timestamp
When the run started
commit_sha
7d3a91f
sample_size
50
sampling_strategy
first_50
run_type
offline
Change any one of these and you have a different evaluation.
PredictL5.1 · 07

Run B scored 4 points higher. Is it better?

Field
Run A
Run B
dataset
traces_v1_full, 200 traces
traces_v1_full, same 200 traces
judge_version
judge_v2.1
judge_v2.3
pass rate
78%
82%
Write down yes, no, or can't tell, and which field decided it for you.
The answerL5.1 · 08

When the judge changes, the comparison measures the judge too.

What changed
Only the judge version. Same traces, same dataset, same system.
How to settle it
Score both runs with the same judge, or check the new judge against human labels first.
Environment ladderL5.1 · 09

Match how much you evaluate to where the system is running.

Offline staging
Full regression suite. All metrics, all segments, all traces. This sets the baseline.
Shadow
Real production queries, no user sees the output. Sampled, blocking metrics only.
Beta
A limited user group. Targeted segments, and the optimization metrics are tracked.
Online experiment
Randomized test. Statistical comparison on the decision metrics.
Production
Sampled monitoring. Alerts on blocking metrics, within a cost budget.
One run, worked throughL5.1 · 10

A single offline run, from sample to stored record.

SamplingThe first 50 traces from traces_v1_full, the v1 baseline set from Week 2
JudgingOracle check on the SQL result, plus the LLM judge on the narrative
AggregationComposite pass rate: the share of traces where SQL and narrative both pass
StorageOne row in eval_runs with the scores and all ten metadata fields
ReportingQuery the row back and confirm every field is filled in
Where it breaksL5.1 · 11

Every v1 trace records the first step where it stopped. Which step fails most?

query_received
intent_classified
context_retrieved
sql_generated
sql_executed
chart_rendered
narrative_written
1,000 v1 traces. Guess the step with the most failures before you count.
Compare two runsL5.1 · 12

Run B changes one thing on purpose. What explains the difference?

Field
Run A
Run B
dataset_version
traces_v1_full 1.0.0
traces_v1_full 1.0.0
sampling_strategy
first_50
last_50
sample_size
50
50
judged by
Same checks
Same checks
completed all steps
?
?
PracticeL5.1 · 13

Design the run record, find where it breaks, and compare two runs.

Base version, everyone
Write the eval_runs record for Run A. Take the first_failure_stage counts from the lesson, work out each step's share of the failures, and explain the top two. Write one paragraph on Run A against Run B. Then design a Run C with one deliberate change and predict what it will show.
Extended version, DS and engineering
Write the lineage record linking Run B to Run A. Describe the query you would use to list every run in time order. List the fields you would add to the schema, and why.
In Lesson 5.7 you run the whole pipeline by hand. Today is the design and the comparison.
Common mistakesL5.1 · 14

Five ways evaluation infrastructure breaks.

Mistake
What happens
Do this instead
Results live in Slack
"Has SQL correctness improved?" No one can say, because nothing was stored.
Every run writes a row to eval_runs.
Incomplete metadata
87 to 83 could be a regression or a new judge. You can't tell which.
Record dataset, judge, sample and code version on every run.
No lineage
How the evaluation changed gets rebuilt from memory.
Record which run each new run refines.
Same depth everywhere
The full suite on every production request adds latency and cost.
Use the environment ladder.
Stages that aren't connected
Aggregation runs on last week's judge scores after the judge step fails.
Explicit stages with defined inputs and outputs.
Knowledge checkL5.1 · 15

Three judgment calls. Write your answer before you read on.

01Run A: 87% with judge_v2.1. Run B: 91% with judge_v2.3. Your PM asks whether this 4 point gain means v2 is ready to ship. What do you check before you answer?
02Of 210 failed v1 traces, 148 stopped at SQL execution and 36 at context retrieval. A teammate wants to spend the next sprint on retrieval. What do you tell them?
03A colleague proposes running the full suite, 500 traces, 8 metrics and 3 LLM judges, on every production request. Why won't that work, and what would you do instead?
Next lessonL5.1 · 16

A pipeline is stages, a record of every run, and depth that fits the environment. Next: the data it runs on.

StagesSampling, judging, aggregation, storage, reporting, alerts.
Run metadataDataset, model, judge, code version, sample. Enough to compare any two runs.
Environment ladderFull suite offline, sampled blocking metrics in shadow and production.
AI ANALYST LAB · aianalystlab.ai