AI Evals for Product DevelopmentL3.4 · 01

Similarity and
retrieval metrics

Measuring whether the system found the right context before it wrote anything
Where we left offL3.4 · 02

In 3.3, retrieval hit rate was a diagnostic. What was it helping you explain?

From 3.3
Hit Rate@5: for what share of queries did at least one relevant document appear in the top five results?
Today
Four metrics that ask more specific questions about the ranked list: how much noise, how much coverage, how early, and in what order.
The scenarioL3.4 · 03

"What was mobile checkout conversion in Q4?" The answer looks right and describes desktop users.

RetrievalReturns documents about desktop conversion and general engagement definitions.
SQL generationValid SQL, real tables and columns. It answers the wrong question.
SummaryFaithful to the SQL result. A chart and specific numbers, all for desktop.
End-to-end metrics: SQL execution success 0.91, summary faithfulness 0.78. The PM asks why users are reporting wrong answers.
Two familiesL3.4 · 04

Similarity tells you how close two texts are. Retrieval metrics tell you whether the right documents came back.

Similarity metrics
Cosine similarity, Euclidean distance. How close are two embeddings? Used to choose an embedding model and tune a cutoff threshold.
Retrieval ranking metrics
Precision@k, recall@k, MRR, NDCG. Did the right documents come back, and did they come back near the top?
A mobile conversion query and a desktop conversion document can score as very similar. The document is still wrong for this question.
Precision@kL3.4 · 05

Precision@k: of the top k documents retrieved, what fraction are relevant?

precision@k = relevant in top k / k
Rank 1doc_schema_usersRelevant
Rank 2doc_metrics_conversionRelevant
Rank 3doc_schema_desktopIrrelevant
Rank 4doc_schema_sessionsRelevant
Rank 5doc_metrics_engagementIrrelevant
Recall@kL3.4 · 06

Recall@k: of all the relevant documents that exist, what fraction made it into the top k?

recall@k = relevant in top k / all relevant documents
Relevant documents in the catalog: 4
doc_schema_users, doc_schema_sessions, doc_metrics_conversion, doc_context_mobile
Same top 5 as the last slide
Three of the relevant documents came back. doc_context_mobile did not.
MRRL3.4 · 07

MRR: how high up is the first relevant document?

reciprocal rank = 1 / rank of first relevant document
First relevant at rank 1
1 / 1
First relevant at rank 2
1 / 2
First relevant at rank 5
1 / 5
MRR, mean reciprocal rank, is the average of that number across every query in the test set.
NDCGL3.4 · 08

NDCG: are the most relevant documents ranked above the partly relevant ones?

Graded relevanceEach document gets a grade instead of yes or no: 2 highly relevant, 1 partly relevant, 0 irrelevant.
Position discount
DCG = Σ grade / log2(rank + 1)
A grade counts for less the further down it appears. NDCG divides by the DCG of the best possible ordering, so a perfect ordering scores 1.0.
Choosing a metricL3.4 · 09

Pick the metric from what the product does with the documents.

Product
Metric
Why
Legal or research search
recall@k
Missing a relevant document is the expensive mistake.
Question-answering bot
MRR
The user reads the first result, so it has to be right.
Small LLM context window
precision@k
Every irrelevant document takes a slot and misleads generation.
Recommendations, multi-document summaries
NDCG
Relevance comes in grades and the order of results matters.
PredictL3.4 · 10

Predict precision@3 and recall@3 for this result before we compute them.

RankDocumentLabel
1doc2Relevant
2doc5Relevant
3doc1Irrelevant
4doc8Irrelevant
5doc3Irrelevant
Query: "What was mobile checkout conversion in Q4?" There are 3 relevant documents in the catalog.
The answerL3.4 · 11

The two numbers can match and still answer different questions.

precision@3
relevant in top 3 / 3
Divides by k, the number of documents you looked at. It measures noise.
recall@3
relevant in top 3 / 3 relevant total
Divides by the relevant documents in the catalog. It measures coverage.
They are equal only when k happens to equal the number of relevant documents. Change k and they move apart.
One query, four metricsL3.4 · 12

Compute all four for one query. The catalog has 3 relevant documents.

Rank
Document
Grade
1
doc4
2, highly relevant
2
doc7
0, irrelevant
3
doc9
1, partly relevant
4
doc6
0, irrelevant
5
doc11
0, irrelevant
Precision@3, recall@3, reciprocal rank, NDCG@5.
PracticeL3.4 · 13

Build a Retrieval Quality Report for the AI Data Analyst.

Base version, everyone
Treat the queries on slides 10 and 12 as a two-query test set. Compute precision and recall at 3 and at 5, and reciprocal rank, for each. Then take the mean of each across both. Say what the precision and recall tradeoff tells you, and write a recommendation: which metric this product should track, and why.
Going further
Say adversarial queries drop precision@3 from 0.80 to 0.60. Write the drop as points and as a percentage. Then describe, step by step, how one retrieval miss turns into the desktop answer in the scenario on slide 3.
The reportL3.4 · 14

What your finished report should let someone read in a minute.

Overall retrievalprecision@3, recall@3, MRR, NDCG@5 across all queries
Adversarial segmentthe same four on adversarial queries, and the size of the drop
Oracle schema matchshare of oracle queries whose expected schema was retrieved
One traced failurea retrieval miss and the SQL error it caused
Recommendationwhich metric this product should track, and why
Common mistakesL3.4 · 15

Four ways teams misread retrieval.

Mistake
What happens
Do this instead
Measuring only end to end
Answer quality drops from 0.82 to 0.71. Nobody can tell whether retrieval or generation changed, so the team edits the prompt.
Track retrieval metrics separately.
Precision when you need recall
A legal search tool reaches precision@3 of 0.92. Recall@10 is 0.40.
Match the metric to the product.
Only looking at the aggregate
Precision@5 is 0.88 overall. On ambiguous entity queries it is 0.52.
Segment by query difficulty.
Unvalidated similarity cutoff
A 0.75 threshold drops relevant documents at 0.72 and lets in distractors at 0.76.
Test cutoffs against labeled data.
Knowledge checkL3.4 · 16

Three judgment calls. Write your answer before you read on.

01A retrieval system has precision@5 = 0.90 and recall@5 = 0.45. A PM asks if it's good enough to ship. What would you tell them for a question-answering bot, and for a legal research tool?
02Precision@3 is 0.85 and NDCG@5 is 0.78, and SQL generation still fails often. How do you use oracle queries to find out whether retrieval or generation is at fault?
03Embedding model A gets higher cosine similarity scores and lower NDCG@5 than model B. Which number decides between them?
Next lessonL3.4 · 17

You can now measure whether the right context came back. Next, you measure whether the answer helped.

What you can do now
Compute precision@k, recall@k, MRR and NDCG, choose the one that fits the product, and separate a retrieval failure from a generation failure.
Lesson 3.5: semantic metrics with judges
Scoring free-form output with human and model judges, rubrics with calibration examples, and judge variance.
AI ANALYST LAB · aianalystlab.ai