AI Evals for Product DevelopmentL3.3 · 01

Deriving evaluation
signals

Signal types, signal roles, and which numbers decide a release
Where we left offL3.3 · 02

What made one example worth keeping in your regression suite, and another worth leaving out?

Write down the rule you actually used for each call.
Keep your answer. We'll come back to it in a minute.
The scenarioL3.3 · 03

Every stage is logged. The PM asks "is v1 ready to ship?" and you can't answer.

What the traces hold
Retrieval scores, document IDs, SQL strings, execution results, narrative text, latencies and token counts for every query.
What hasn't been decided
Which of those numbers block a release, which ones help you debug, and which ones are just there because they were easy to log.
DefinitionL3.3 · 04

A signal compares the system's output to ground truth. Every signal has a type and a role.

TracesWhat the system logged: retrieval results, SQL, execution output, narrative text.
SignalsType: how you compute it. Role: which decision it feeds.
DecisionsShip, ramp, hold, or roll back.
Type 1: execution-based oraclesL3.3 · 05

If the output can run, run it and compare results.

String matching
The generated query uses a CTE. The reference query uses a subquery. The text differs, so the answer is marked wrong.
Oracle
Run both queries against the warehouse and compare the result sets. Same rows back, so the answer is marked correct.
Works for anything executable: SQL, code, API calls, tool calls.
Type 2: structural signalsL3.3 · 06

Structural signals check the form of the output. They can't tell you if the content is right.

What they check
The JSON parses. The SQL parses without syntax errors. Required fields are present. The response length is within bounds.
What they miss
A valid JSON response with a wrong number in it passes every structural check.
Fast, cheap, and no ground truth needed beyond a schema.
Type 3: semantic signalsL3.3 · 07

Semantic signals need a judge and a rubric, and the judge adds its own variance.

Judge variance
Score the same fixed output twice with an LLM judge and get two different scores. The measuring tool is noisy.
System variance
Run the AI Data Analyst twice on the same input and get two different outputs. The system is inconsistent.
Relevance, completeness, factual accuracy, tone. Always say which variance you're measuring.
Signal rolesL3.3 · 08

A signal's role is the decision it feeds.

Gateblocks a releaseSQL correctness must be at least 90%. Below that, you hold.
Diagnostichelps you debugRetrieval Hit Rate@5 tells you why SQL correctness dropped. It doesn't block anything on its own.
Driverexplains where quality variesSQL correctness split by query complexity shows which kinds of questions pull the average down.
Choosing a signalL3.3 · 09

Start from the ground truth you have. Then pick the type, then the role.

What ground truth do you have?Known result sets, a schema, or a rubric a judge can apply.
Which type fits?Executable output: oracle. Format and schema: structural. Relevance or tone: semantic.
Which decision does it feed?Gate, diagnostic, or driver. If you can't name one, drop the signal.
Predict before the demoL3.3 · 10

Say Hit Rate@5 comes back at 0.68 on 300 test queries. What happens to the queries it missed?

Hit Rate@5: the share of queries where at least one relevant document shows up in the top five results.
Write down how many of the 300 reach the SQL step with nothing relevant, and what the SQL step does with them.
The resultL3.3 · 11

When retrieval misses, every stage after it works without the context it needs.

Retrieval missNo relevant document in the top five.
Missing contextNo metric definitions, schema details or business rules.
SQL without contextThe model guesses at tables, columns and definitions.
Wrong answerBad SQL, and a summary that describes the wrong numbers.
Did your count match, and did you predict what the SQL step does next?
PracticeL3.3 · 12

Choose more signals and build a signal catalog.

Base version, everyone
Say what Precision@5 and MRR would tell you that Hit Rate@5 doesn't. Describe what over-retrieving and under-retrieving would look like in those three numbers. Fill a signal catalog with at least five signals covering all three types and all three roles.
Going further
Add a semantic signal and write down where its variance comes from. Describe how you'd check whether poor retrieval predicts poor SQL correctness. Grow the catalog to ten signals.
The artifactL3.3 · 13

A signal catalog: what you measure, what you check it against, and what it decides.

Signal
Type
Checked against
Role
Cost
How it's computed
SQL correctness
Execution
Oracle queries
Gate
Medium
Run both queries, compare result sets
Retrieval Hit Rate@5
Execution
Relevant document IDs
Diagnostic
Low
Share of queries with a relevant document in the top five
JSON validity
Structural
Schema definition
Diagnostic
Low
Parse and validate against the schema
Query complexity
Structural
SQL structure
Driver
Low
Parse the SQL, count joins and aggregations
Narrative accuracy
Semantic
Rubric
Driver
High
Human or LLM judge
Common mistakesL3.3 · 14

Five ways signal design goes wrong.

Measuring everything that's availableA model change moves half the metrics up and half down, and nobody can decide. Assign roles first.
Using similarity when you could run itScoring SQL by text overlap when you could execute it and compare results.
Reporting a point estimate alone"Correctness is 76%" with no interval doesn't say how much to trust it.
Mixing up the two variancesBlaming the system for scores that move because the judge is noisy, or the reverse.
Gating on an unproven signalBlocking releases on a metric that hasn't been shown to be stable.
Knowledge checkL3.3 · 15

Three judgment calls. Write your answer before you read on.

01You have 50 oracle queries. 38 of the generated queries return the same result set as the reference. What's the SQL correctness point estimate, and what else do you need before it's decision-ready evidence?
02The PM says "ship if Hit Rate@5 is above 70%." You measure 71%, 95% CI 66% to 76%, n = 300. Do you ship on this signal alone?
03Your catalog includes "narrative accuracy, LLM judge score 0 to 10." Which two sources of variance affect that score?
Next lessonL3.3 · 16

You can sort signals by type and role. Next, you measure the retrieval step.

What you can do now
Pick a signal from the ground truth you have, assign it a role, and build a catalog that says which numbers block a release.
Lesson 3.4: retrieval metrics
Why a similarity score isn't the same as relevance, and precision, recall, MRR and NDCG for the retrieval step.
AI ANALYST LAB · aianalystlab.ai