Semantic metrics with human and model judges
How do we score a written answer with a judge, and how do we know the judge agrees with people?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
The SQL is right and the chart is right, and the summary still tells the PM that Q4 revenue grew 12% when the data says 8%. Retrieval metrics and SQL checks both pass, because the mistake only exists in the words. To catch it, something has to read the summary against the data.
At a thousand questions a week, that something is an LLM judge: a model that gets the question, the query results, the summary and a rubric, and returns Pass or Fail. You design one for narrative faithfulness. The rubric is binary, because a 1 to 5 score drifts between runs. It spells out a 5% materiality threshold and the edge cases. The judge writes its reasoning before the verdict, so you can see why it disagrees with a person.
Then you check it against human labels. You split about 100 labeled summaries into train, dev and test, and measure agreement with Cohen’s Kappa, which corrects for the agreement you’d get by chance. You read an example first run on 10 summaries where the misses all fall the same way, and use that pattern to decide what to change in the prompt. The course target is a Kappa of 0.7. Below 0.6, the judge’s numbers shouldn’t support a ship, ramp, hold or roll back call.
The practice is one round of refinement on paper: rewrite the Fail criteria to catch the misses, work out Kappa for a rerun grid, write the log entry, and set a stop rule. The result is a Semantic Rubric with Calibration Plan: the rubric, the few-shot examples, dev and test Kappa, and the cases where the judge still gets it wrong.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→