AI Evals for Product DevelopmentL4.6 · 01

Metric specs and
release criteria

Writing down what "good enough to release" means, before you look at the results
Where we left offL4.6 · 02

What was the strongest driver you found, and how did it change what you'd fix first?

The driver
Query complexity, domain, user role, or an interaction between two of them.
Today
That finding becomes a threshold for a segment, not just a line in a brief.
Take 30 seconds and answer from memory.
The scenarioL4.6 · 03

v2 is ready. Your PM asks, "Are we shipping?" The room can't agree.

SQL success68% → 74%
Retrieval72% → 85%
Narrative quality0.82 → 0.79
Latencyup 8%
Engineering"74% SQL success is great."
Design"0.79 narrative quality is a regression."
Data science"The latency might breach the p95 target."
The metric specL4.6 · 04

A metric spec pins down eight things.

DefinitionThe exact calculation.
Unit of analysisEach question on its own, or grouped per user?
PopulationWhich traces count. Test accounts out?
SegmentsWhich breakdowns get their own threshold.
SamplingEvery trace or a sample, and why.
AggregationAverage, worst case, p95.
OwnershipWho runs the pipeline, owns the definition, and investigates.
VersioningHow often it's refreshed and how changes are tracked.
Release criteriaL4.6 · 05

Release criteria put every metric in one of three classes.

BlockingMust pass, or the release stops.SQL success at or above 60%
Narrative quality drop under 5%
GuardrailBlocking metrics that protect against harm to users or the business.Cost per query up less than 15%
p95 latency under 2,000 ms
OptimizationTracked and reported. Does not gate the release.Chart rendering speed
Query parse time
Keep blocking to 3 to 5 metrics.
Threshold patternsL4.6 · 06

Four ways to set a threshold.

Pattern
What it protects
Example
Absolute floor
A minimum below which the system isn't usable, whatever the baseline.
SQL success >= 60%
Delta
No meaningful regression from the last version, while allowing noise.
Narrative quality drop < 5%
Confidence-aware
The improvement is real and not measurement noise.
Interval on the difference excludes zero
Segment-aware
No segment hides under an acceptable average.
No segment < 50%
The orderL4.6 · 07

Set the thresholds before you look at the new version's numbers.

In this order
  1. Measure the v1 baseline and its interval.
  2. Set thresholds from the baseline and the business needs.
  3. Evaluate v2.
  4. Compare v2 with the thresholds.
Not this order
  1. Evaluate v2.
  2. Pick a threshold just under what v2 scored.
  3. v2 passes, because it was always going to.
Predict before the demoL4.6 · 08

v1 SQL success is 68%, with a 95% interval of 65% to 71%. Where would you put the blocking threshold?

68%, 65%, 70%, or something else? Tied to the baseline, or absolute?
What happens if v2 comes in at 67%? Write your threshold and your reason.
The demo: segment baselinesL4.6 · 09

The overall 68% hides a much lower rate for some users and some questions.

By user role
Executive55%[49%, 61%]
PM72%[68%, 76%]
DS71%[66%, 76%]
By query complexity
Ambiguous45%[38%, 52%]
Moderate68%[63%, 73%]
Simple82%[78%, 86%]
Scenario baselines for this lesson, given as stated inputs.
The demo: a stricter baselineL4.6 · 10

A query that runs is not the same as a query that returns the right answer.

Execution success
Did the SQL run without an error? Easy to measure on every trace. Can't catch a query that runs and returns wrong data.
Correctness against known answers
Did the SQL return the expected result? Needs a set of test questions where you already know the right answer.
The demo: the thresholdsL4.6 · 11

Three thresholds, each set from the v1 baseline before v2 was evaluated.

Absolute floorsql_success_rate, v1 at 68% [65%, 71%]>= 60%, below the interval
Deltanarrative_quality, v1 at 0.82 [0.79, 0.85]>= 0.779, a drop under 5%
Segment floorsql_success_rate by user roleno user role below 50%
PracticeL4.6 · 12

Write the narrative quality spec, classify your metrics, and write the release gate.

Base version, everyone
Complete the narrative quality spec: population, segments, thresholds, ownership. Classify your metrics as blocking (3 to 5), guardrail or optimization. Write the release gate as conditions joined by AND. Run v2 from slide 3 through it and make the call. Document one metric conflict, retrieval depth against cost.
Extended version, DS and engineering
Read the segment intervals on slide 9: which segments are clearly different, and which could share a threshold? Check that v1 passes its own gate. Write the spec for oracle_pass_rate, including how the known-answer set grows.
The artifactL4.6 · 13

The metric spec and release criteria for the AI Data Analyst.

Metric
Class
v1 baseline [95% CI]
Release threshold
sql_success_rate
Blocking
68% [65%, 71%]
>= 60% overall, >= 50% every user role
narrative_quality
Blocking
0.82 [0.79, 0.85]
>= 0.779 (drop < 5%)
oracle_pass_rate
Blocking
61% [57%, 65%]
>= 55%
cost_per_query
Guardrail
$0.08 [$0.07, $0.09]
< $0.092 (rise < 15%)
p95_latency
Guardrail
1,650 ms [1,500, 1,800]
< 2,000 ms
chart_rendering
Optimization
94% [91%, 97%]
tracked, not gating
Common mistakesL4.6 · 14

Four ways release criteria go wrong.

Mistake
What happens
Do this instead
Thresholds set after seeing v2
v2 passes because the line was drawn under it.
Write the criteria before you evaluate the candidate.
No segment thresholds
A fine average hides executives at 40%.
Add a floor that each key segment must clear.
Everything is blocking
With 15 blocking metrics, noise fails at least one on most releases.
Keep blocking to 3 to 5.
Ignoring the baseline's interval
A threshold inside v1's interval fails versions no different from v1.
Set floors below the interval, for business reasons.
Knowledge checkL4.6 · 15

Three judgment calls. Write your answer before you read on.

01v1 SQL success is 68% [65%, 71%]. Set the blocking threshold at 68%, 65%, 70% or something else? What happens if the threshold is inside the interval rather than below it?
02Your dashboard has 12 metrics and a colleague wants all 12 to be blocking. How many should be, how do you decide, and what's the risk of too many or too few?
03v2 narrative quality is 0.79. v1 was 0.82. The rule is "drop under 5%." Does v2 pass? Show the math. Does it change anything that v1's interval is 0.79 to 0.85?
Next lessonL4.6 · 16

You have a metric spec and release criteria. Next, a pipeline that runs them.

Metric specDefinition, unit, population, segments, sampling, aggregation, ownership, versioning. Every metric gets all eight.
Release criteria3 to 5 blocking metrics, guardrails for cost and latency, everything else optimization.
The ruleAll blocking pass, then all guardrails pass, then release. Thresholds set before the candidate is evaluated.
AI ANALYST LAB · aianalystlab.ai