AI Evals for Product DevelopmentL4.1 · 01

Metric strategy

Which metrics block a release, and which ones you keep improving
Where we left offL4.1 · 02

Which of your judges would you trust to block a release?

Release gate
If this number misses its threshold, the release waits. You need to trust the measurement.
Continuous improvement
Tracked over time and reviewed on a cadence. Useful even when the measurement is noisy.
Take your Judge Report Card from 3.6. Sort your judges into the two columns.
The scenarioL4.1 · 03

Twelve metrics, and your PM asks one question: "Should we ship v2?"

Retrieval recall@5
SQL correctness
Narrative faithfulness
Cost per query
Judge agreement
Latency p95
Multi-turn coherence
Policy compliance
and four more
"Release it if all twelve look good" is not an answer anyone can act on.
Two jobsL4.1 · 04

Every metric gets one of two jobs.

Blocking metric
Has a threshold. If it misses, the release does not go out, however good everything else looks.
Optimization metric
Has a target. Tracked and reviewed on a cadence. Missing it starts work, and the release still goes out.
The job decides the threshold, how often you watch it, who owns it, and how good the measurement has to be.
The decision testL4.1 · 05

Ask one question of each metric.

"If this metric fails and everything else passes, do we release?"
No, we holdIt's a blocking metric. It needs a threshold and a written reason for it.
Yes, we note it and keep improvingIt's an optimization metric. It needs a target and an owner.
The metric treeL4.1 · 06

Organize the metrics as a tree, from the outcome you care about down to what you can measure in a trace.

North starOne measure of the value users get. You usually can't compute it from traces.
DriversA handful of mid-level metrics that move the north star.
Trace-level metricsOutput quality you compute from traces. Each one feeds a driver.
The metric treeL4.1 · 07

The AI Data Analyst's tree.

User trustProxy: share of queries where the user acts on the answer within one hour
Answer correctnessRetrieval recall@5, SQL correctness, narrative faithfulness
Response latencySQL execution latency, retrieval latency
Cost per queryModel and infrastructure cost for one query
Multi-turn coherenceKeeps context across a conversation
If answer correctness drops, the tree tells you which three metrics to check first.
Threshold rationalesL4.1 · 08

Every blocking threshold needs a written reason.

Without a reason
SQL correctness sits at 91%. A change drops it to 89%. Two other metrics improved. No one wrote down why the line is 90%.
With a reason
The threshold says what breaks below it. The hold or release call follows from that.
Threshold rationalesL4.1 · 09

Four places a threshold's reason comes from.

ExecutionThe output has to work at all.SQL has to parse, so syntax validity sits close to 100%.
Business impactA cost or revenue limit.Above $0.15 a query the feature loses money.
Competitive benchmarkMatch what people use today.Answers at least as good as the manual dashboards.
User toleranceWhere behavior tips over.Past about 5 seconds, users give up.
Predict before the demoL4.1 · 10

How many of the twelve metrics would you classify as blocking?

Apply the decision test to each one. Write down a number out of twelve.
For each metric you count, you should be willing to hold a release over that metric alone.
The demo: the classificationL4.1 · 11

The worked classification: what blocks this release, and why.

Blocking metric
Threshold
Reason
SQL syntax validity
>98%
Execution. If the SQL doesn't parse, the user gets nothing back.
Cost per query
<$0.15
Business impact. $0.10 infrastructure plus a $0.05 margin at 10,000 queries a day.
Policy compliance
>97%
Regulatory. A violation can have legal consequences.
Every other metric on the list is tracked as optimization.
PracticeL4.1 · 12

Classify four hard metrics and write release criteria for v2.

Base version, everyone
Classify retrieval recall@5, latency p95, judge agreement and multi-turn coherence with the decision test, one sentence of reason each. Then write release criteria for the v1 to v2 retrieval change: what must improve, what must hold steady, and which tradeoffs you'd accept.
Extended version, DS and engineering
Put a number on each threshold reason, the way slide 9 did for cost. Map which metrics depend on which. Then test one tradeoff against your criteria: recall@5 up 5 points, latency p95 up 20%.
The output is a metric system spec: the classification with reasons, the tree, and the v2 release criteria.
Common mistakesL4.1 · 13

Five ways a metric strategy breaks.

Mistake
What happens
Do this instead
Everything is blocking
Each extra gate is another chance to hold a good release on noise.
Keep blocking to the few metrics you'd hold for alone.
No threshold reason
The metric sits near the line and the team can't defend the line.
Write what breaks below the number.
A flat list with no tree
Twelve numbers, no story, arguments about each one.
Connect every metric to a driver and the north star.
Blocking metrics mixed up with runtime checks
The wrong thing gets built: a release gate where a per-request check was needed, or the reverse.
Gates judge a release on a test set. Runtime checks protect one request.
One threshold for every kind of query
A lookup and an exploratory "why" question get the same bar.
Set bars by query type. Segmentation in 4.4 shows how.
Knowledge checkL4.1 · 14

Three judgment calls. Write your answer before you read on.

01SQL correctness is blocking at >95%. A change adds 10 points to retrieval recall and 8 to answer completeness, and SQL correctness drops to 93%. Your PM says quality is clearly better. Do you release?
02Answer completeness is an optimization metric, target 80%, currently 78%. Six months later it's 83%, and the PM says users don't seem any happier. What might that tell you about your tree?
03A colleague sets latency p95 <5s as blocking "because that's our SLA." Is that a good reason? Which source is it, and what would make it stronger?
Next lessonL4.1 · 15

You now have a metric system. Next, you design each metric in it.

ClassificationBlocking or optimization for every metric, with a written reason for every blocking threshold.
Metric treeUser trust at the top, four drivers, and the trace-level metrics under each driver.
Release criteriaWhat has to improve, what has to hold steady, and which tradeoffs you'll accept for v2.
AI ANALYST LAB · aianalystlab.ai