Week 3: Rigorous Measurement of Output Success and Failure · Lesson 3.5

Semantic metrics with human and model judges

How do we score a written answer with a judge, and how do we know the judge agrees with people?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome back. In 3.4 we evaluated retrieval, whether the right documents came back, with metrics like precision@k and MRR. Today we move one step later in the pipeline, to the written summary the user actually reads. Checking a summary means reading it and deciding whether it tells the truth about the data. That's a judgment call, and we want to make it a thousand times a week. So today we build a judge, a model that makes that call against a rubric, and then we check it against human labels before anyone relies on its numbers. Back in Week 1 we said uncertainty comes from two places: the system and the judge. This is the lesson where we measure the judge.

About this lesson

The SQL is right and the chart is right, and the summary still tells the PM that Q4 revenue grew 12% when the data says 8%. Retrieval metrics and SQL checks both pass, because the mistake only exists in the words. To catch it, something has to read the summary against the data.

At a thousand questions a week, that something is an LLM judge: a model that gets the question, the query results, the summary and a rubric, and returns Pass or Fail. You design one for narrative faithfulness. The rubric is binary, because a 1 to 5 score drifts between runs. It spells out a 5% materiality threshold and the edge cases. The judge writes its reasoning before the verdict, so you can see why it disagrees with a person.

Then you check it against human labels. You split about 100 labeled summaries into train, dev and test, and measure agreement with Cohen’s Kappa, which corrects for the agreement you’d get by chance. You read an example first run on 10 summaries where the misses all fall the same way, and use that pattern to decide what to change in the prompt. The course target is a Kappa of 0.7. Below 0.6, the judge’s numbers shouldn’t support a ship, ramp, hold or roll back call.

The practice is one round of refinement on paper: rewrite the Fail criteria to catch the misses, work out Kappa for a rerun grid, write the log entry, and set a stop rule. The result is a Semantic Rubric with Calibration Plan: the rubric, the few-shot examples, dev and test Kappa, and the cases where the judge still gets it wrong.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→