AI Evals for Product DevelopmentL2.4 · 01

The regression
safety net

Catch regressions before a deploy, when the same input can give you different outputs
Where we left offL2.4 · 02

Which v1 fields would a regression check actually read?

generated_sqlsql_successsql_errorretrieval_querynum_docs_retrievedmodel_versionprompt_template_version
With only the question and the final answer logged, what could you test?
The scenarioL2.4 · 03

A small prompt change passes code review and breaks multi-table joins.

The changePrompt update makes the SQL easier to read
Code reviewApproved, looks like an improvement
DeployMerged and released
Twenty minutes laterUsers report join questions failing
The causeOne line about foreign keys dropped from the prompt
The test case for this failure already existed. Nothing ran it before the deploy.
Why this is harder for AIL2.4 · 04

Normal regression tests assume the same input gives the same output.

Traditional software
Run the test, check for an exact match. If it passes once with the same code, it passes again.
AI systems
The same input can come back different. A suite can pass on one run and fail on the next with no code change.
The job is to tell a real regression from run-to-run variation.
What goes in the suiteL2.4 · 05

Build the suite from your failure taxonomy.

TaxonomyThe categories from Lesson 1.3
FilterCategories that need ongoing measurement
SampleOne real trace per category
ExpandEdge cases, adversarial inputs, known correct answers
SuiteA curated set that runs on every change
You're choosing cases that cover every category you care about, rather than testing everything.
Metric classificationL2.4 · 06

Not every failed metric should stop a deploy.

Tier
Pass rule
Gate behavior
Examples
Blocking
Every case passes
Any failure blocks the deploy
SQL correctness, tool call success, policy compliance
Optimization
No regression beyond a tolerance band
Blocks on a real regression, warns below target
Narrative quality, latency, retrieval precision
Informational
None
Logged, never blocks
Retrieval diversity, token usage
Multi-trial protocolL2.4 · 07

Run each noisy case more than once, and track two numbers.

pass@3: can it do it?
The case passes at least once in three trials.
reliable@3: does it do it every time?
The case passes on all three trials.
For continuous metrics like latency, compare to a band around the baseline instead of an exact value.
The gateL2.4 · 08

The suite runs on every pull request and decides whether it can merge.

Pull requestPrompt, model or code change
CI runs the suiteGitHub Actions or whatever your team uses
Blocking checksKnown-answer comparisons, pass or fail
Optimization checksMulti-trial, compared to the band
ResultBlocked with the failures linked, or cleared to merge
A dashboard tells you what already happened. A gate stops it before users see it.
Suite evolutionL2.4 · 09

Every failure you find in production becomes a new case.

Add
A failure the suite missed is a gap. Add a case for it and version the suite.
Keep
A case that always passes is still guarding a fix.
Retire
Only duplicates, or cases whose scenario no longer exists.
Predict before the demoL2.4 · 10

You rerun the suite with no code change. What happens to each group?

10 SQL oracle cases
Run the generated SQL and the correct SQL, compare results. One run each.
10 narrative judge cases
An LLM judge scores the summary. Three trials each, tracked as pass@3 and reliable@3.
Write your prediction for each group, and why, before we advance.
The demo: two runs, no changeL2.4 · 11

The oracle checks held. The judge checks moved.

SQL oracle cases
Same cases passed on both runs. The same case failed on both runs.
Narrative judge cases
Both pass@3 and reliable@3 dropped on the second run, with nothing changed.
Example numbers in the notes, made up to show the pattern. Your own reruns give you the real ones.
What it means for the gatesL2.4 · 12

Gate hard on stable checks. Put the noisy ones behind a band you measured.

Oracle checks
Stable across reruns. Safe to make blocking.
Judge checks
Optimization tier, multiple trials, and a tolerance band wider than the movement you saw with no change.
Set the band from your own no-change reruns.
PracticeL2.4 · 13

Build a 20-case regression suite with gating rules.

Base version, everyone
Write 20 cases from your 1.3 taxonomy: at least 5 categories and at least 5 known-answer queries. Classify each metric. Use the two no-change runs on slide 11 to set a tolerance band for the judge cases. Write the gating rules in plain English, test them on 5 pull requests you describe, and write the add-and-retire policy.
Going further
Write the gating rules as pseudocode. Use the formula from 2.2 to estimate how many cases you'd need before a 5-point regression stands out from rerun noise.
The artifactL2.4 · 14

A threshold table that says what blocks and what warns.

Metric
Tier
Threshold
Gate
SQL correctness
Blocking
100% of suite cases
Blocks on any failure
Policy compliance
Blocking
100% of suite cases
Blocks on any failure
Latency p95
Optimization
Within 110% of baseline
Blocks above the band
Narrative quality
Optimization
Target 80% pass@3
Blocks on regression, warns below target
Token usage
Informational
None
Logged only
Common mistakesL2.4 · 15

Three ways a regression gate goes wrong.

Mistake
What happens
Do this instead
Over-gating
A single-trial judge score blocks deploys. The gate fails on noise and people learn to ignore it.
Put noisy metrics in the optimization tier with multiple trials and a measured band.
Under-gating
Every metric is informational to avoid flaky runs. Nothing blocks, so regressions go out anyway.
Make at least one stable metric blocking.
Tolerance mismatch
The band is tighter than the natural variation, or far looser than it.
Rerun the suite 5 times with no change and set the band just above what moves.
Knowledge checkL2.4 · 16

A judge scores one case Pass, Fail, Pass. Blocking or optimization?

Trial 1: Pass
Trial 2: Fail
Trial 3: Pass
Narrative quality on a complex query, in a 30-case suite, scored by majority vote. The rule: "block the deploy if any blocking metric fails." Write your answer and your reason.
Next lessonL2.4 · 17

The suite works for this pipeline. Next, how instrumentation changes with the type of system.

LLM apps
Prompt in, response out.
RAG systems
Retrieval, then generation.
Agents
Choosing and calling tools.
AI ANALYST LAB · aianalystlab.ai