Correcting for an imperfect judge
The judge is wrong some of the time. What is the true pass rate, and how far can we trust it?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Your LLM judge scores the AI Data Analyst’s outputs and reports a 78% pass rate. You also know from Lesson 3.5 that the judge gets things wrong. It labels 90% of real passes as passes, so its true positive rate (TPR) is 0.90, and it catches 85% of real failures, so its true negative rate (TNR) is 0.85. The 78% has errors in both directions, and you don’t yet know which way the true number sits.
The Rogan-Gladen formula, borrowed from medical screening tests, uses the observed rate, TPR and TNR to estimate the true rate. You predict which way the correction goes, work it out, and check the answer by running it back through the judge. Because TPR and TNR come from a small labeled set, you bootstrap them to put an interval on the corrected rate, and you nudge each one by 5 points to see how much the answer moves. You also see why a judge’s accuracy has to be re-measured over time: the same hosted model can score very differently after a version update.
Accuracy isn’t the only problem. A judge can prefer whichever answer comes first in a pairwise comparison, or prefer the longer of two answers that say the same thing. You test for both with paired examples, and you see how correcting the same 78% can move the call from hold to ship once you compare it to the release bar.
The practice works a position bias score and corrects the pass rate again with test-split numbers, then checks whether the release call still holds. Everything goes into a Judge Report Card, the twelve-field record that any later metric using judge scores points back to.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→