Week 3: Rigorous Measurement of Output Success and Failure · Lesson 3.6

Correcting for an imperfect judge

The judge is wrong some of the time. What is the true pass rate, and how far can we trust it?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome back. In 3.5 you built an LLM judge and checked it against human labels with Cohen's Kappa. That answered one question: does the judge agree with people? Today we take the next step. The judge is going to score a lot of outputs from the AI Data Analyst and report a pass rate, and we know the judge makes mistakes. So what's the real pass rate of the system underneath? We'll work that out with a correction formula, put a confidence interval on the result, and then test the judge for two biases that don't show up in its accuracy numbers at all. Everything you find goes into a Judge Report Card, a one-page record of how far to trust the judge, and later weeks rely on it whenever a metric uses judge scores.

About this lesson

Your LLM judge scores the AI Data Analyst’s outputs and reports a 78% pass rate. You also know from Lesson 3.5 that the judge gets things wrong. It labels 90% of real passes as passes, so its true positive rate (TPR) is 0.90, and it catches 85% of real failures, so its true negative rate (TNR) is 0.85. The 78% has errors in both directions, and you don’t yet know which way the true number sits.

The Rogan-Gladen formula, borrowed from medical screening tests, uses the observed rate, TPR and TNR to estimate the true rate. You predict which way the correction goes, work it out, and check the answer by running it back through the judge. Because TPR and TNR come from a small labeled set, you bootstrap them to put an interval on the corrected rate, and you nudge each one by 5 points to see how much the answer moves. You also see why a judge’s accuracy has to be re-measured over time: the same hosted model can score very differently after a version update.

Accuracy isn’t the only problem. A judge can prefer whichever answer comes first in a pairwise comparison, or prefer the longer of two answers that say the same thing. You test for both with paired examples, and you see how correcting the same 78% can move the call from hold to ship once you compare it to the release bar.

The practice works a position bias score and corrects the pass rate again with test-split numbers, then checks whether the release call still holds. Everything goes into a Judge Report Card, the twelve-field record that any later metric using judge scores points back to.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→