AI Evals for Product DevelopmentL3.5 · 01

Semantic metrics
and LLM judges

Building a judge and checking it against human labels
Where we left offL3.5 · 02

Retrieval found the right documents and the SQL ran. What can still go wrong?

What we measured last lesson
Which documents came back and how highly they ranked. Precision@k is high and the right table is in hand.
What the user reads
The written summary the model builds from those documents and query results.
Put one way the summary can fail in the chat.
The failureL3.5 · 03

The SQL is right and the chart is right. The summary says 12% when the data says 8%.

SQL executionReturns 8% growth for Q4 revenue.
ChartDraws an 8% bar.
Summary"Q4 revenue grew 12%."
SQL checks and retrieval metrics both pass. Only a check that reads the summary against the data catches this.
The judgeL3.5 · 04

An LLM judge is a model that grades another model's output against a rubric.

InputsThe user's question, the query results and the summary.
RubricWritten Pass and Fail criteria, with worked examples.
JudgeWrites its reasoning, then returns Pass or Fail.
A person with the same inputs and the same rubric is the reference the judge gets checked against.
Design choice 1L3.5 · 05

Pass or Fail gives more consistent verdicts than a 1 to 5 score.

1 to 5 scale
Same summary, two runs: a 3, then a 5. The rubric rarely says where a 3 ends and a 4 begins, so the judge lands somewhere different each time.
Pass or Fail
Same summary, two runs: Fail, then Fail. One boundary, and you write it down in the rubric.
Need more levels? Add a second Pass/Fail judge with a higher bar.
Design choice 2L3.5 · 06

The judge writes its reasoning first, then the verdict.

Evaluation: Mobile 3.2% and desktop 5.1% match the result set. The 1.9 point gap is correct. The claim that desktop led in every month is supported by the data.
Verdict: Pass
When the judge disagrees with a person, the reasoning shows you why.
The rubricL3.5 · 07

The rubric says what Pass and Fail mean, down to the edge cases.

Pass
Every key finding stated accurately. Numbers match the result set. No claims the data doesn't support. Every change larger than 5% is mentioned.
Fail
Any one of: a number not in the result set, an unsupported claim, an omitted change larger than 5%, a wrong direction or size.
Edge cases
Rounding 7.8% to "about 8%": Pass.
Three of four findings, the missing one at 6%: Fail.
A plausible conclusion the data doesn't show: Fail.
One dimension: narrative faithfulness. Does the summary describe the query results accurately and completely?
Before you iterateL3.5 · 08

Split the labeled examples before you start changing the prompt.

Train, about 10%
Try the first rubric. Pick the few-shot examples for the prompt.
Dev, about 45%
Change the prompt, rerun, and track agreement for each version.
Test, about 45%
Run once at the end. No changes to the judge after you've seen it.
Only 50 labels? Roughly 5, 20 and 25. Keep test the largest split.
Measuring agreementL3.5 · 09

Percent agreement can look high by chance. Cohen's Kappa corrects for that.

below 0.00Poor
0.00 to 0.20Slight
0.21 to 0.40Fair
0.41 to 0.60Moderate
0.61 to 0.80Substantial
0.81 to 1.00Almost perfect
Course target: Kappa of 0.7 or higher. Below 0.6, don't use the judge to support a ship, ramp, hold or roll back call.
Predict before the runL3.5 · 10

We're about to run the first-draft judge. Write down two predictions first.

The rubric from this lesson. The train split: 10 summaries a person already labeled.
01What percent of the time will the judge's verdict match the human's?
02When it's wrong, will it be too strict or too lenient?
The first runL3.5 · 11

An example first run on the train split: judge verdicts against human labels.

Judge: Pass
Judge: Fail
Human: Pass
6
0
Human: Fail
3
1
Work out percent agreement and Kappa from the grid before we advance.
Reading the missesL3.5 · 12

The three misses have the same shape: the judge passes summaries that leave out a finding.

What the human saw
Three of four findings stated correctly. The fourth, a 6% change, left out. That's over the 5% line, so Fail.
What the judge did
Checked the three findings it could see, found them correct, and passed the summary. It never asked what was missing.
Fix: spell out the omission rule in the Fail criteria, add a Fail example that shows it, and rerun on dev.
PracticeL3.5 · 13

Fix the judge for the misses, then decide what counts as done.

Base version, everyone
Rewrite the Fail criteria from slide 7 so they catch the misses on slide 12, and write the new Fail example. Say a rerun on 10 dev summaries, 6 human Pass and 4 human Fail, misses only one Fail. Fill in the grid, compute percent agreement and Kappa, and write the log entry. Then write your stop rule.
Going further
Compute Kappa for a judge that passes all 10 summaries on the same split. Explain why percent agreement and Kappa tell such different stories about it.
The overfit checkL3.5 · 14

Compare dev Kappa with test Kappa. A big gap means the prompt was tuned to the dev examples.

Reading the gap
A few hundredths is normal sampling noise. A gap of 0.10 or more is a red flag: go back to the rubric and plan on a fresh test set.
Semantic Rubric with Calibration Plan
The rubric, the few-shot examples, dev and test Kappa, and the cases where the judge still gets it wrong.
Common mistakeL3.5 · 15

A judge below 0.6 can't tell you whether a pass rate measures quality or noise.

How it goes wrong
A PM hears "the judge says 85% of summaries pass" and approves the launch. Nobody checked the judge against people.
If you can't get past 0.6
Simplify the rubric until each criterion is concrete. Get more labels and look for labeling mistakes. Or use the judge as a filter and send uncertain cases to a person.
Knowledge checkL3.5 · 16

Three judgment calls. Write your answer before you read on.

01Your judge has dev Kappa 0.75 and test Kappa 0.58. What does the gap suggest, and what do you do?
02A colleague wants a 1 to 5 scale instead of Pass/Fail because "it gives us more information." What are the tradeoffs?
03Your judge has Kappa 0.72 and reports a 68% pass rate. A PM asks, "Can we trust this 68%?" What do you tell them?
Next lessonL3.5 · 17

You can build a judge and measure it against people. Next, you correct for what it still gets wrong.

What you can do now
Write a Pass/Fail rubric, split labeled data, iterate a judge on dev, measure agreement with Kappa, and validate once on test.
Lesson 3.6: correcting for an imperfect judge
Correcting the pass rate with Rogan-Gladen, testing the judge for position and verbosity bias, and re-measuring it on a schedule.
AI ANALYST LAB · aianalystlab.ai