AI Evals for Product DevelopmentL2.2 · 01

How little evidence
is enough

Sizing an eval set for the decision you are about to make
Where we left offL2.2 · 02

Which fields in v1 make it possible to measure SQL correctness?

v0 loggingrequest_iduser_queryfinal_answerlatency_mssuccesstimestamp
v1 logginguser_querygenerated_sqlsql_successsql_errororacle_sqlcorrectness_match
The scenarioL2.2 · 03

42 of 50 test queries are correct. Your PM wants a full rollout.

What you have
50 test queries with oracle SQL. 42 match. An 84% success rate on this sample.
The rollout rule
Before a full rollout, the accuracy estimate has to be within ±5 points, at 95% confidence.
The question: is 50 enough, and if not, how many do you need?
The inputsL2.2 · 04

Three things decide how many samples you need.

The decisionA full rollout, a limited ramp and an A/B comparison need different strength of evidence.
Confidence levelHow sure you need to be. Set by what it costs to be wrong.
NoiseHow much quality moves from query to query, and how consistent the evaluator grading it is.
The formulaL2.2 · 05

Three values in, one sample size out.

n = z² × p(1 − p) / ε²
zThe lookup value for your confidence level: 1.96 for 95%, 1.645 for 90%
pYour current success rate, as a decimal
εThe precision you need. ±5 points means ε = 0.05
Evaluator reliabilityL2.2 · 06

An inconsistent evaluator gives you less evidence than your sample count suggests.

Expert with an answer key
Kappa close to 1.0. Each sample counts as a full sample.
LLM evaluator
Kappa often 0.6 to 0.8. Rough rule of thumb in this course: divide the required n by kappa.
Kappa measures evaluator agreement, from 0 (no better than chance) to 1 (perfect).
Evidence by decisionL2.2 · 07

Different decisions need different precision.

Decision
Precision needed
Typical n
Full rollout
±4 to 5 points
200 to 400
Hold or ramp to a limited beta
±7 to 10 points
50 to 150
Experiment, v1 vs v2
80% chance of detecting a real improvement
Depends on the size of the improvement
Predict before the mathL2.2 · 08

Is 50 samples enough for ±5 points at 95% confidence?

42 of 50 correct. How wide is the interval right now, and how many samples would you need to get it to ±5?
Write down both guesses before we advance. Then note what would change your answer.
The math: where you areL2.2 · 09

With 50 samples, the interval is too wide to support a full rollout.

Plug in n = 50
Margin = 1.96 × √(0.84 × 0.16 / 50)
What the PM hears
"The true accuracy could be well below or well above 84%. We can't say which."
The math: where you need to beL2.2 · 10

How many samples get the margin to ±5?

n = 1.96² × 0.84 × 0.16 / 0.05²
Compare the result with your prediction before reading the notes.
The confidence leverL2.2 · 11

Dropping from 95% to 90% confidence cuts the samples you need.

95% confidence
z = 1.96
90% confidence
z = 1.645
Which confidence level fits is a product decision about risk.
The precision leverL2.2 · 12

Halving the margin roughly quadruples the samples.

Margin
±15 points
±10 points
±5 points
Samples needed
?
?
?
95% confidence, success rate 0.84. Fill in each count, then work out ±3, before reading the notes.
PracticeL2.2 · 13

Fill a six-cell decision matrix: three decisions at two confidence levels.

Base version, everyone
For full rollout (±5) and hold (±10), at 90% and 95%, work out n with the formula from slide 10 and p = 0.84. For the experiment row, start from the 730 per version in the knowledge check and say which way 90% moves it. Write a two-sentence evidence argument for each cell.
Going further
Adjust every cell for evaluator kappa of 0.65, 0.75 and 0.85 with the rule from slide 6. Mark the cells where a noisy evaluator changes which decision you can afford.
What you hand the PML2.2 · 14

Where you are, where you need to be, and where more samples stop helping.

n = 10A handful of checks
n = 50Where you are now
n for ±5The rollout rule
n = 400Diminishing returns
Sketch the margin of error against n from 10 to 500, at 84%.
Common mistakesL2.2 · 15

Three ways sample sizing goes wrong.

Mistake
What happens
Do this instead
One fixed number for everything
"We always use 100." A system near 50% gets under-sampled; a system near 95% gets over-sampled.
Calculate n for each decision and each success rate.
Ignoring the evaluator
A kappa 0.6 judge with no adjustment gives you about 40% less evidence than you think.
Adjust n for evaluator reliability.
Precision you don't need
Collecting 500 samples for ±3 when the PM would accept ±8 for a ramp.
Ask what decision you're making before sizing.
Knowledge checkL2.2 · 16

Three calculations. Work them out before you read on.

0135 of 50 queries are correct (70%). Can you roll out if the rule is 95% confidence that quality is above 65%? Calculate the 95% interval.
02You expect v2 to lift SQL accuracy from 84% to 89%. For an 80% chance of detecting that, about how many samples per version? At $0.50 per sample, what does the experiment cost?
03Your LLM judge for tone has kappa 0.68. The formula said 180 samples assuming perfect grading. How many do you need?
Next lessonL2.2 · 17

Decision, confidence, calculate and adjust, then the gap. Next, trace design.

What decision?Full rollout ±5, hold ±10, experiment uses power analysis
Confidence?90%, 95% or 99%, set by the cost of being wrong
Calculate and adjustApply the formula, then divide by evaluator kappa
OutputRequired n, current n, and the gap between them
AI ANALYST LAB · aianalystlab.ai