AI Evals for Product DevelopmentL3.2 · 01

Ground truth and
regression suites

What you compare against, and the set of cases every release has to pass
Where we left offL3.2 · 02

In Lesson 1.3 you built a failure taxonomy from raw traces.

Which category had the highest user impact, and why would you cover it first?
The scenarioL3.2 · 03

SQL correctness is 90%. Compared to what?

Previous modelIts outputs are used as the expected answers.
Copied into the suiteNobody runs them against the database.
Tests passThe new system agrees with the old answers.
ProductionUsers see wrong numbers in their dashboards.
A metric can only be as right as the answers it is scored against.
Three sourcesL3.2 · 04

Match the ground truth source to the task.

Source
Cost
Delay
Use it for
Execution oracle
Near zero
Seconds
SQL correctness, tool outputs, anything you can run and check
Expert annotation
2 to 3 min per trace
Hours to days
Narrative quality and other calls that need human judgment
User feedback
None
Days to weeks
Production validation, directional only
Use an oracle wherever you can. Spend annotation time where only a person can judge.
Execution oraclesL3.2 · 05

The oracle runs both queries and compares what comes back.

Generated SQL
Written by the AI Data Analyst. It might use a subquery, a different join order, or a different filter than the reference.
Expected SQL
A verified query for the same question. The oracle runs both against the warehouse and compares the result sets.
The oracle passes the trace when the result sets match, within a stated tolerance for rounding.
Where annotation goesL3.2 · 06

Don't spend annotation time on anything an oracle can check.

Use the oracle
SQL correctness, tool outputs, retrieval results you can check against a known list.
Use annotation
Faithfulness: does the narrative match the data? Completeness: does it cover the key findings?
CalibrationL3.2 · 07

Before you scale annotation, check that your annotators agree.

Cohen's kappa
How often two raters agree, after removing the agreement you'd expect by chance. 0 is chance, 1 is perfect.
The rule for this course
Above 0.6, keep annotating. Below 0.4, fix the rubric before you label anything else.
Regression suiteL3.2 · 08

A regression suite is the set of verified cases every release has to pass.

New versionA prompt, model or code change.
Run the suiteEvery case, scored against verified ground truth.
All pass?The gate before release.
Ship or triageA failure is a real regression or an outdated case.
Start with 10 to 20 high-impact cases from your failure taxonomy. Grow it by promotion. Retire cases that stop matching the product.
PromotionL3.2 · 09

A case gets into the suite only if it meets all four criteria.

Reproducible. You can re-run it with your current instrumentation.
Distinct. It covers a failure mode the suite doesn't already cover.
Verified ground truth. From the right source for the task.
Justifies the cost. Severity or frequency is worth maintaining it.
Every promoted case gets a one-line justification.
Predict before the demoL3.2 · 10

Three candidate traces. Which ones do you promote?

Trace
What happened
What the suite already has
A
A SQL logic error that produces wrong revenue numbers. Checked with the oracle.
No cases of this kind of logic error.
B
A SQL syntax error: a missing semicolon.
Two cases of the same syntax error.
C
The user typed gibberish and the system returned an error.
Nothing for this.
Run each one through the four criteria and write down promote or reject, with the criterion that decided it.
The oracle demoL3.2 · 11

"What was Q4 revenue?" Does the generated SQL pass?

GeneratedSELECT SUM(revenue) FROM orders WHERE quarter = 'Q4'
ExpectedSELECT SUM(revenue) FROM orders WHERE date >= '2024-10-01' AND date < '2025-01-01'
The expected query is the verified one. Write down pass or fail before we go on.
Synthetic dataL3.2 · 12

Fill the gaps production hasn't shown you yet.

01Find the gapA segment of your evaluation surface map (L1.2) with no cases.
02Define attributesFor example: ambiguous entity, trend question, high complexity.
03Generate queriesAn LLM writes queries constrained to those attributes.
04Filter for realismWould a real user ask this? Drop the ones they wouldn't.
Then verify ground truth for each survivor and send it through the same four criteria.
PracticeL3.2 · 13

Build a Regression Suite Promotion Plan.

Base version, everyone
Write the oracle's explanation for the slide 11 failure, and one generated query for the same question that would pass. Run three failures from your 1.3 taxonomy through the four criteria. List 3 coverage gaps from your surface map. Fill in the plan.
Going further
Pick one gap. Write attribute combinations for it, then five queries from them. Apply the realism check, and name the ground truth source for each query you keep.
Common mistakesL3.2 · 14

Three ways a suite stops being useful.

Mistake
What happens
Do this instead
Promoting duplicates
The suite fills with variations of one easy error. It gets slow, people stop running it, and the dangerous failures have no cases.
Apply the distinct and cost criteria to every case.
Skipping the realism filter
Synthetic queries nobody would ask set a baseline that production never matches.
Filter every synthetic query before it's promoted.
Freezing the suite
The product changes and the suite keeps testing the old one.
Version it, promote on a cycle, retire stale cases.
Knowledge checkL3.2 · 15

Three judgment calls. Write your answer before you read on.

01The AI Data Analyst writes SQL that runs but aggregates over the wrong date range. What ground truth source do you use for this failure, and why?
02A suite has 200 cases and takes 15 minutes. A PM asks why it can't run on every commit. What's the problem, and what do you recommend?
03You promoted 10 synthetic queries without a realism filter. Three months later the LLM judge pass rate on the suite is 95% and production user satisfaction is 70%. What went wrong, and how do you fix it?
Next lessonL3.2 · 16

You can now build a suite on verified answers. Next, what to compute from it.

What you can do now
Pick the ground truth source for a task, check annotator agreement, promote cases with four criteria, and fill gaps with filtered synthetic queries.
Lesson 3.3: evaluation signals
What signals you can compute when ground truth is imperfect, and what each one is for: gates, diagnostics and drivers.
AI ANALYST LAB · aianalystlab.ai