AI Evals for Product DevelopmentL3.6 · 01

Correcting for an
imperfect judge

Corrected pass rates, confidence intervals, and position and verbosity bias
RecallL3.6 · 02

The first-run judge in 3.5 got a Kappa of about 0.29. Where does that sit on the scale?

below 0.00
0.00 to 0.20
0.21 to 0.40
0.41 to 0.60
0.61 to 0.80
0.81 to 1.00
Poor
Slight
Fair
Moderate
Substantial
Almost perfect
Answer from memory before you look back at 3.5. Then place the 0.7 target.
The scenarioL3.6 · 03

Your judge reports a 78% pass rate. The judge itself is wrong some of the time.

What the judge reports
78% of AI Data Analyst outputs pass. Leadership asks whether that's good enough to ship.
What you know about the judge
TPR = 0.90: it labels 90% of real passes as passes and misses 10%. TNR = 0.85: it catches 85% of real failures and lets 15% through as passes.
The 78% is what an imperfect instrument reported. The true rate could be higher or lower.
Judges changeL3.6 · 04

A model's accuracy on one task can move a long way between versions.

GPT-4, March 2023
97.6% accuracy identifying prime numbers.
GPT-4, June 2023
2.4% on the same task.
Chen, Zaharia and Zou (2023). The TPR and TNR you measured describe the judge on the day you measured them.
The correctionL3.6 · 05

The Rogan-Gladen formula estimates the true rate from an imperfect judge.

true_rate = (observed_rate + TNR − 1) / (TPR + TNR − 1)
Inputs
The rate the judge reports, and the judge's TPR and TNR from your labeled calibration data.
Only valid when TPR + TNR > 1
The judge has to beat a coin flip. As the sum gets close to 1, the denominator shrinks toward zero and the estimate becomes unstable.
PredictL3.6 · 06

The judge reports 78% with TPR = 0.90 and TNR = 0.85. Is the true rate higher or lower?

Write down higher or lower, and one sentence on why.
If you're stuck, take 100 outputs where 70 truly pass and 30 truly fail. How many would this judge call a pass?
Work it outL3.6 · 07

Put the numbers into the formula.

01
Numerator: observed_rate + TNR − 1 = 0.78 + 0.85 − 1
02
Denominator: TPR + TNR − 1 = 0.90 + 0.85 − 1
03
Corrected rate: numerator divided by denominator
Check your answer against your prediction. Then check it by running the corrected rate back through the judge.
UncertaintyL3.6 · 08

TPR and TNR are estimates too. Bootstrap them to get an interval on the corrected rate.

01
Resample the labeled dev examples with replacement.
02
Recompute TPR and TNR on the resample, then the corrected rate.
03
Repeat 1,000 times. The 2.5th and 97.5th percentiles are the 95% interval.
Report the corrected rate with its interval, the sample size, and the dominant variance source.
SensitivityL3.6 · 09

Move TPR and TNR by 5 points each. Which one moves the corrected rate more?

Vary TPR, hold TNR at 0.85
TPR from 0.85 to 0.95, observed rate fixed at 78%.
Vary TNR, hold TPR at 0.90
TNR from 0.80 to 0.90, observed rate fixed at 78%.
A big swing means your calibration set is too small to pin the correction down.
Bias 1L3.6 · 10

Position bias: the judge prefers an answer because of where it appears.

Pair
Order 1 (A then B)
Order 2 (B then A)
Consistent?
1
Picks A
Picks A
Yes
2
Picks A
Picks B
No
Zheng et al. (2023): GPT-4 as a judge gave the same verdict after swapping answer order in 65% of cases.
Bias 2L3.6 · 11

Verbosity bias: the judge scores longer answers higher when they're no better.

Short correct answer
Q4 revenue, the number, and one line on how it was computed.
Long correct answer
The same number and method, plus several paragraphs restating them.
AlpacaEval: correcting judge scores for length raised the Spearman correlation with Chatbot Arena's human-preference ranking from 0.94 to 0.98.
Same score, different decisionL3.6 · 12

Say the release bar is 80%. The judge reports 78%. Does correcting it change the call?

On the raw number
78% is under the 80% bar. It looks like a hold.
After correction
Same judge, TPR 0.90 and TNR 0.85. Compare the corrected rate to the bar, then ask what its interval would have to show.
Work it out before we advance. You already have the formula.
PracticeL3.6 · 13

Check the judge for bias and correct its pass rate on the test split.

Base version, everyone
Say 3 of 12 position pairs flip when you swap the order. Write the bias score and whether it's acceptable for choosing between two prompts. Say the test split gives TPR 0.90 and TNR 0.80, lower than dev. Correct the 78% again and say whether the call against the 80% bar changes. Write a one-paragraph summary.
Going further
Put TPR = 0.55, TNR = 0.50 and an observed rate of 0.60 into the formula. Explain what goes wrong and what you'd tell the team. Say what happens to the interval if you label four times as many examples.
The artifactL3.6 · 14

The Judge Report Card: twelve fields that say how far to trust the judge.

Judge modelWhich model scores the outputs
Judge prompt versionSo a prompt change forces a re-check
Calibration setName, size and splits
TPR, with intervalFrom the test split
TNR, with intervalFrom the test split
Cohen's KappaAgreement with human labels
Observed pass rateWhat the judge reports
Corrected pass rateRogan-Gladen, with interval
Position biasScore and conclusion
Verbosity biasScore and conclusion
Recommended actionsFor example, run both orders
LimitationsWhat the card doesn't cover
Common mistakesL3.6 · 15

Three ways a judge score leads to the wrong call.

Mistake
What happens
Do this instead
Reporting only the point estimate
The corrected rate looks above threshold, but its interval crosses it.
Report the interval. If it crosses the threshold, label more examples first.
Skipping the bias tests
A prompt comparison favors whichever variant sat in the preferred position more often.
Run the position and verbosity tests before using the judge in an experiment.
Only validating on dev
The judge prompt was tuned against the dev split, and its TPR drops on new data.
Measure TPR and TNR on the held-out test split and compare.
Knowledge checkL3.6 · 16

Three judgment calls. Write your answer before you read on.

01TPR = 0.95, TNR = 0.70, observed rate = 80%. Without computing, will the corrected rate be higher or lower, and why?
02Your judge's position bias score is 0.35. A colleague says that's fine because it's under 0.5. Do you agree, and what would you do next?
03Dev-split TPR is 0.92 and test-split TPR is 0.78, a 14-point gap. Your checklist flags gaps over 15. Do you use this judge?
Next lessonL3.6 · 17

You can now say how far to trust a judge. Next, you decide which metrics block a release.

What you can do now
Correct a judge's pass rate for its errors, put an interval on it, test it for position and verbosity bias, and record all of it in a Judge Report Card.
Lesson 4.1: metric strategy
Blocking metrics that stop a release, optimization metrics you improve over time, and how they link to business outcomes.
AI ANALYST LAB · aianalystlab.ai