Retrieval metrics: precision, recall, MRR and NDCG
Did the system find the right context before it wrote anything, and which retrieval metric fits this product?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
A user asks the AI Data Analyst for mobile checkout conversion in Q4. Retrieval, the step that looks up schema context before any SQL is written, comes back with documents about desktop conversion. The SQL runs, the summary describes it accurately, and the numbers are for desktop users. SQL execution success is 0.91 and summary faithfulness is 0.78, so nothing you are measuring points at the problem. That is why retrieval gets evaluated on its own.
Similarity scores like cosine similarity tell you how close a query and a document are. A mobile question and a desktop document can be very close and still wrong for each other. Retrieval metrics compare the ranked list against labels a person assigned. Precision@k measures noise in the top k. Recall@k measures how many of the relevant documents came back at all. MRR looks at how high the first relevant document sits. NDCG uses graded labels and rewards putting the most relevant documents first.
You compute each one by hand for a single query, see why precision and recall can match in one example and split apart as soon as k changes, and learn to pick the metric from what the product does with the documents: recall for legal search, MRR for a one-answer bot, precision for a small context window, NDCG for graded, ordered results.
The practice is a small Retrieval Quality Report on the two example queries from the lesson: mean precision and recall at 3 and at 5 and MRR, a read on the precision and recall tradeoff, and a recommendation. The extended version works an adversarial drop as points and as a percentage, and traces one retrieval miss to the wrong answer it caused.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→