AI Evals for Product DevelopmentL3.1 · 01

Grounding evaluation
in user value

Which metrics predict what users experience, and how much weight each one gets
From Week 2L3.1 · 02

Which of the v1 signals let you measure something v0 could not?

v0 instrumentationquery_idsql_texttimestamp
v1 instrumentationedit_occurredsession_abandonedsatisfaction_ratingfollow_up_count
Think about what each field tells you about the person on the other side.
The scenarioL3.1 · 03

The dashboard is green. The PM asks whether users find it useful, and you can't answer.

SQL correctness92%
Retrieval accuracy87%
Hallucination rate4.2%
These numbers describe the system. None of them has been checked against what users did.
Why it happensL3.1 · 04

In a June 2025 Deloitte survey, a third of generative AI users said they had run into incorrect or misleading answers.

In testing
Curated test sets, clean data and questions chosen because they have answers. Accuracy scores come out high.
In production
Ambiguous questions, unfamiliar jargon and topics combined in ways the test set never covered.
Source: Deloitte, 2025 Connected Consumer survey of about 3,500 US consumers.
Two layersL3.1 · 05

Leading metrics are fast proxies. Lagging metrics are what users actually did.

Leading metrics
Fast, offline, computed on a test set in minutes. SQL correctness, retrieval accuracy, hallucination rate.
Lagging metrics
Slower, need real users in production. Edit rate, abandon rate, follow-up rate, satisfaction.
Before a leading metric drives a decision, check that it predicts a lagging one.
Three layersL3.1 · 06

Business metrics sit on top. They are the slowest to move and the closest to value.

Leadingfast · offlineSQL correctness, retrieval accuracy, hallucination rate
Laggingslow · productionEdit rate, abandon rate, satisfaction
Businessslowest · aggregateAnalyst time saved, decision confidence, support tickets
Each arrow is an assumption until someone checks it.
Checking the linkL3.1 · 07

Four ways to check whether a leading metric predicts a lagging one.

01Correlation study. Compute how strongly the technical metric and the user signal move together across traces.
02Offline and online agreement. Check that versions scoring higher on the test set also do better with real users.
03User signal tracking. Log edits, abandons and follow-ups next to the evaluation scores on every trace.
04Periodic re-validation. Recompute the link every quarter as the system and its users change.
Decision authorityL3.1 · 08

How strongly a metric tracks a user outcome decides what it is allowed to do.

Correlation
r range
Role
What it can do
Strong
r > 0.7
Blocking metric
Must pass before a release goes out.
Moderate
0.5 to 0.7
Optimization metric
Track it and improve it. Don't gate on it.
Weak
r < 0.5
Feedback only
Useful for debugging. Keep it off the release gate.
A blocking metric needs r above 0.7 with at least one lagging metric.
Proxy theaterL3.1 · 09

Retrieval latency went from 150ms to 80ms. User satisfaction did not move.

What the team saw
A precise measurement, a real improvement, a green dashboard.
What users saw
The same wrong data, arriving a little faster.
Before you optimize a metric, ask: if this improves by 20%, what user behavior changes?
PredictL3.1 · 10

Three candidate metrics. Which one tracks user satisfaction most closely, and which least?

ASQL execution success. Did the query run without errors?
BResult set completeness. Did the query return all the relevant rows?
CNarrative actionability. Does the summary help the user make a decision?
Write down your ranking before we advance.
The resultL3.1 · 11

In the example, the metric closest to what the user reads tracks satisfaction best.

What users see
The summary. Whether it helps them decide what to do next.
What users don't see
Whether the SQL errored and retried, or how many rows came back.
Put each metric in a tier: blocking, optimization or feedback only.
Easy to logL3.1 · 12

Some metrics are right there in the logs and predict nothing about satisfaction.

Retrieval latency
Still useful for infrastructure planning.
Narrative length
Still useful for cost estimates.
Total tokens
Still useful for budget planning.
A confidence interval that crosses zero means you can't tell the correlation apart from noise.
One traceL3.1 · 13

Trace t_0042: the query ran, some rows were missing, and the user rated it 4 out of 5.

SQL execution success1
Result set completeness0.70
Narrative actionability0.90
User satisfaction4 / 5
Judge this trace by each metric in turn. Which one matches what the user experienced?
PracticeL3.1 · 14

Build a correlation table and a Leading-to-Lagging Metric Map.

Base version, everyone
Put the six metrics from slides 11 and 12 in one table with r, the interval and a tier. Say which one you'd let block a release, and which you'd keep for debugging only. Fill the metric map for one metric from your Lesson 1.3 failure taxonomy.
Going further
Write the validation plan for that metric's link: the user signal you'd check it against, such as edits or abandoned sessions, how many traces you'd need, and what the interval has to show before the metric can block.
Common mistakesL3.1 · 15

Five ways grounding goes wrong.

Mistake
What happens
Do this instead
Assume a metric is valid
Release gates built on SQL correctness. After launch, abandon rate climbs from 12% to 25%.
Correlate it with a lagging metric before it gates anything.
Optimize a proxy
Latency drops from 150ms to 80ms. Satisfaction stays flat.
Name the user behavior that should change first.
Test set drifts from production
95% quality on the test set, 70% satisfaction from real users.
Compare offline and online scores. Rebalance the test set.
Let a proxy go stale
A gate validated in Q1 is still gating in Q3 after users' priorities moved.
Recompute each correlation quarterly.
Mix up the layers
Waiting two weeks on satisfaction surveys to judge each change.
Iterate on leading metrics. Validate on lagging ones.
Knowledge checkL3.1 · 16

Three judgment calls. Write your answer before you read on.

01SQL correctness 92%, retrieval 88%, hallucination 3.8%. The PM asks whether to ship. What question has to be answered first?
02Your best offline metric correlates with satisfaction at r = 0.68 on 150 traces, 95% CI 0.58 to 0.76. Your gate needs 0.7. Do you use it to block releases?
03After launch, SQL correctness and hallucination rate hold steady, but abandon rate rises from 12% to 19%. What does that tell you, and what do you do?
Next lessonL3.1 · 17

You know which metrics predict user value. Next, you build ground truth for those metrics.

What you can do now
Sort metrics into leading, lagging and business. Correlate a leading metric with a user signal, and give it the authority its correlation supports.
Lesson 3.2: ground truth and regression suites
Gold-standard datasets, regression suites that catch real failures, and synthetic data for edge cases production hasn't shown you yet.
AI ANALYST LAB · aianalystlab.ai