AI Evals for Product DevelopmentL1.4 · 01

Distributional
thinking

Error bars, segments and variance sources for ship decisions
Where we left offL1.4 · 02

How many failure categories did you find, and which one showed up most?

Your taxonomy from 1.3
Categories discovered bottom-up from v0 traces, each with a severity and a triage label.
Today's shift
From "we found scoping errors" to a sentence like "scoping errors occur in 18% of complex queries, plus or minus 5, driven mostly by system variance."
The scenarioL1.4 · 03

Say 62 of 100 traces look acceptable. Your PM asks if that's good enough to ship.

Would a re-run give 58% or 67%?
No error bars, so you don't know.
Is it 5% catastrophic or 38% mediocre?
No segments, so you don't know the shape.
System randomness, or noisy measurement?
Variance sources not separated.
One number. Three unknowns. No principled way to decide.
Three errorsL1.4 · 04

Without distributional thinking, teams make three systematic errors.

Error
What the team does
What goes wrong
Optimistic single-sample estimate
"62% pass rate. Ship it."
The next batch comes back at 55%. The estimate never had error bars.
Capability confused with reliability
"pass@5 = 0.90, works great."
reliable@5 = 0.60. Nearly half of users hit at least one failure.
Variance source confused
"The system is too random. Lower the temperature."
85% of the variance was evaluator noise. A sprint spent fixing the wrong thing.
Made-up numbers, to show the pattern.
Error barsL1.4 · 05

A point estimate hides how much you don't know. Bootstrap it.

01
Take your observed outcomes: pass, fail, fail, pass, fail.
02
Draw a new sample of the same size, with replacement, so repeats are allowed.
03
Recompute the metric on that resample.
04
Repeat about a hundred times and sort the results.
05
The 2.5th and 97.5th percentiles are your 95% confidence interval.
No assumptions about the shape of the data. Works for any metric.
SegmentsL1.4 · 06

An aggregate of 0.78 can hide a segment at 0.45.

Simple lookup
pass@5 = 0.92
Revenue totals, single-table queries
Multi-table join
pass@5 = 0.45
Cross-segment comparisons, complex SQL
Trend analysis
pass@5 = 0.78
Time series, period over period
Made-up numbers. With half the traffic simple lookups and a fifth joins, the aggregate is 0.78.
Variance sourcesL1.4 · 07

System variance and evaluator variance need different fixes. Decompose before you mitigate.

System variance
Re-run the AI Data Analyst on the same query and get different outputs. Fix with temperature, retries, prompt changes.
Evaluator variance
Re-run the evaluator on the same fixed output and get different scores. Fix with calibration, rubric refinement, consensus scoring.
In this lesson the pass signal is a logged field, so re-scoring gives the same answer and evaluator variance is zero. Week 3 changes that.
Evidence sufficiencyL1.4 · 08

The evidence tells you which decision you can defend.

Full rollout
Narrow intervals, and every must-pass metric clears its threshold in every segment.
Gradual ramp
Wider intervals, but the direction is positive. Deploy to 10%, monitor, expand.
Hold or roll back
An interval includes the threshold, so you can't tell pass from fail. Or a guardrail failed.
The question is never "what is the metric." It is "does the evidence support the proposed action."
Predict before the demoL1.4 · 09

Five trials of one query. Predict pass@5, reliable@5, and the dominant variance source.

Trial 1Correct SQL, correct narrative
Trial 2Correct SQL, minor hallucination in the narrative
Trial 3SQL syntax error
Trial 4Correct SQL, correct narrative
Trial 5Correct SQL, overstated confidence in the narrative
"What was Q4 mobile checkout conversion?" Temperature 0.7. A made-up example. Write your three predictions before we advance.
The demo: compute the two metricsL1.4 · 10

Two formulas turn a single-trial rate into the two metrics.

pass@k
1 - (1 - p)^k
The chance that at least one of k trials succeeds, when each trial succeeds with probability p.
reliable@k
p^k
The chance that all k trials succeed. For one specific query, this always falls as k grows.
Here p is the single-trial success rate. For our five trials, p = 2/5.
The demo: add the error barsL1.4 · 11

Now bootstrap the five outcomes.

01
Outcomes: pass, fail, fail, pass, fail.
02
Resample five with replacement. Recompute pass@5 and reliable@5.
03
Repeat a hundred times, sort, and read the 2.5th and 97.5th percentiles.
Could you ship on this evidence? Decide before we advance.
The demo: 47 questions, five runs eachL1.4 · 12

In the v0 traces, capability reaches 1.0 and reliability falls to 0.36.

k
pass@k
95% CI
reliable@k
95% CI
1
0.81
[0.68, 0.91]
0.81
[0.68, 0.91]
3
1.00
[1.00, 1.00]
0.45
[0.32, 0.60]
5
1.00
[1.00, 1.00]
0.36
[0.23, 0.49]
47 questions asked at least five times, first five runs each, SQL executed counts as a pass. Before we discuss it: how much of the variation is the system, and how much is the evaluator?
PracticeL1.4 · 13

Interpret the segment distributions and write evidence-based ship criteria.

Base version, everyone
Use the made-up segments: pass@5 of 0.92, 0.45 and 0.78 from slide 6, and reliable@5 of 0.85 for simple lookups, 0.20 for multi-table joins and 0.55 for trends. Compute the gap for each and find the largest. Write a ship criterion in the form "ship if reliable@5 is above X with an interval narrower than Y." Recommend ship, ramp or hold per segment.
Going further
In the v0 traces, pass@5 is 1.0 in every question type and reliable@5 runs from 0.12 on trend questions to 0.50 on comparisons. Each type holds 7 to 11 questions. Say what segments that small do to the interval, and whether you'd make a per-segment call on them.
Common mistakesL1.4 · 14

Five ways teams misuse distributional evidence.

Confusing pass@k with reliabilityReporting pass@5 = 0.90 as "works 90% of the time" when reliable@5 = 0.60.
Point estimates without intervalsShipping on "62% success" when the interval, 0.52 to 0.72, includes the hold threshold.
Blaming the system for all varianceA sprint of re-prompting when 85% of the variance was evaluator noise.
pass@k alone for the ship decisionShipping on pass@5 = 0.98 when reliable@5 = 0.28.
Ignoring the worst segmentAggregate pass@5 = 0.85 looks fine. The multi-join segment is at 0.45.
Knowledge checkL1.4 · 15

Three judgment calls. Write your answer before you read on.

01The PM says "80% pass rate, ship it." The multi-join segment has reliable@5 = 0.22. What evidence do you show?
02pass@5 = 0.78 with a 95% interval of 0.68 to 0.88. The ship criterion is reliable@5 above 0.80. Can you ship?
03System variance is 15% of the total and evaluator variance is 85%. The team wants to spend a sprint re-prompting the model. What do you prioritize?
Next lessonL1.4 · 16

The gap between the two bars is user experience risk. Next, what it costs to close it.

Simple
Small gap
Multi-join
Large gap
Trend
Moderate gap
The made-up segment example from the exercise. Blue is pass@5 and grey is reliable@5: simple 0.92 and 0.85, multi-join 0.45 and 0.20, trend 0.78 and 0.55.
AI ANALYST LAB · aianalystlab.ai