Week 3: Rigorous Measurement of Output Success and Failure · Lesson 3.3

Deriving evaluation signals from available ground truth

Given imperfect ground truth, what signals can we compute and what are they for?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome to lesson 3.3. Last lesson you built a regression suite out of production traces and synthetic examples. Back in Week 2 you instrumented the AI Data Analyst, so every stage now logs something: retrieval scores, the SQL it wrote, execution results, the summary text, latency, token counts. That's a lot of data. Today we turn it into evaluation signals, which are the specific numbers you'd point to when someone asks whether the system is ready. We'll sort signals two ways. By type, meaning how you compute the signal. And by role, meaning which decision it feeds. Then you'll predict a retrieval number, see what it does to the rest of the pipeline, and build a signal catalog for the AI Data Analyst.

About this lesson

By Week 3 the AI Data Analyst logs something at every stage: retrieval scores, the SQL it wrote, execution results, the summary text, latency and token counts. The PM asks whether v1 is ready to ship, and a dashboard full of numbers still doesn’t answer that. What’s missing is a decision, made ahead of time, about which numbers you measure and what each one is for.

An evaluation signal is the number you get when you compare the system’s output to ground truth, the known right answer. Every signal has a type and a role. The type is how you compute it. Execution-based oracles run the output, for example running the generated SQL and a reference query and comparing the result sets. Structural signals check format, like whether the JSON parses or the required fields are there. Semantic signals need a judge and a rubric, and the judge brings its own variance on top of the system’s. The role is the decision the signal feeds. A gate blocks a release, a diagnostic helps you find where a failure came from, and a driver explains where quality varies. If you can’t name the decision a signal feeds, you drop it.

You take an example hit rate on 300 test queries, work out how many queries reach the SQL step with nothing relevant, and look at what a retrieval miss does to the SQL and the summary that come after it.

The practice is a signal catalog. You work out what Precision@5 and MRR would add to the hit rate, describe what over-retrieving and under-retrieving look like, and write down at least five signals with their type, what they’re checked against, their role, their cost and how they’re computed. The extended version adds a judge-based signal with its variance written down and grows the catalog to ten signals.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→