Week 2: Instrumentation and Reliability Engineering · Lesson 2.2

How much evidence is enough to decide

How many test cases do we need before the accuracy number can support this decision?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome back. In 2.1 we wrote the logging contract, the list of fields engineering adds so you can actually evaluate the system. Today's question comes right after that. You've got the fields, you've got a test set, and your PM asks how many test queries you need before you can trust the number. Every sample costs something. If an expert is writing the correct SQL for each test query and checking the output, that's real hours. But if you go with too few, you can't tell whether the feature is at 75 percent or 85 percent, and those are different products. At 75, one query in four is wrong. At 85, it's about one in seven. By the end of today you'll be able to work out the sample size for a specific decision at a specific confidence level, and explain it to the person asking.

About this lesson

You have 50 test queries with correct answers written ahead of time, and 42 of them come back right. That’s 84 percent. Your PM wants to roll out to everyone and asks whether 50 is enough. This lesson is how you answer that with a number instead of a guess.

Three things decide the sample size: the decision you’re making, how confident you need to be, and how noisy the measurement is. A full rollout needs a tight estimate, a limited ramp can live with a wider one, and comparing two versions needs a power analysis instead. Noise comes from the system and from the evaluator grading it.

The formula is n = z² × p(1 − p) / ε². You work through it on the 42-of-50 example against a rollout rule of plus or minus 5 points at 95 percent confidence, and see how far 50 samples falls short. Then you pull the two levers you control. Lowering the confidence level cuts the count. Tightening the margin raises it fast, because halving the margin roughly quadruples the samples.

An LLM evaluator isn’t perfectly consistent, and its kappa score tells you how much. The course uses a rough rule: divide the required sample size by kappa, so a noisy judge needs more samples to give you the same evidence.

The practice is a decision matrix: full rollout, hold and experiment, each at 90 and 95 percent confidence, with a sample size you work out from the formula and a two-sentence evidence argument in every cell. The extended version adjusts every cell for three evaluator kappa values.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→