Week 3: Rigorous Measurement of Output Success and Failure · Lesson 3.4

Retrieval metrics: precision, recall, MRR and NDCG

Did the system find the right context before it wrote anything, and which retrieval metric fits this product?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome to lesson 3.4. So far this week we've mostly been evaluating what the AI Data Analyst produces at the end: the SQL, the result, the summary. Today we go one step earlier. Before the system writes any SQL, it goes and looks up schema context, the table and metric definitions it thinks are relevant to the question. That lookup is called retrieval. If it pulls the wrong definitions, everything after it works from the wrong starting point. So we're going to look at how to measure retrieval on its own. We'll cover four ranking metrics, precision at k, recall at k, MRR and NDCG, how to compute each one by hand for a single query, and how to pick the one that fits what your product does with the documents it retrieves. Then you'll build a Retrieval Quality Report on the AI Data Analyst's labeled retrieval data.

About this lesson

A user asks the AI Data Analyst for mobile checkout conversion in Q4. Retrieval, the step that looks up schema context before any SQL is written, comes back with documents about desktop conversion. The SQL runs, the summary describes it accurately, and the numbers are for desktop users. SQL execution success is 0.91 and summary faithfulness is 0.78, so nothing you are measuring points at the problem. That is why retrieval gets evaluated on its own.

Similarity scores like cosine similarity tell you how close a query and a document are. A mobile question and a desktop document can be very close and still wrong for each other. Retrieval metrics compare the ranked list against labels a person assigned. Precision@k measures noise in the top k. Recall@k measures how many of the relevant documents came back at all. MRR looks at how high the first relevant document sits. NDCG uses graded labels and rewards putting the most relevant documents first.

You compute each one by hand for a single query, see why precision and recall can match in one example and split apart as soon as k changes, and learn to pick the metric from what the product does with the documents: recall for legal search, MRR for a one-answer bot, precision for a small context window, NDCG for graded, ordered results.

The practice is a small Retrieval Quality Report on the two example queries from the lesson: mean precision and recall at 3 and at 5 and MRR, a read on the precision and recall tradeoff, and a recommendation. The extended version works an adversarial drop as points and as a percentage, and traces one retrieval miss to the wrong answer it caused.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→