Week 1: Foundations and Economics · Lesson 1.1

What AI evaluation is and why it requires a different approach

Why don't the evaluation methods I already know work for AI-powered product features?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome to AI Evals for Product Development. This is the first lesson, and we're going to start with the thing that makes evaluating an AI feature different from testing normal software. We're going to start with a scenario where everything passes review and users still complain, work out why that happens, and put two numbers on it that you can bring to a product manager. Then we'll look at one question run five times, and at what the course's own v0 traces show when users ask the same question again and again. If you've ever had an AI feature that worked in the demo and felt flaky once real people used it, this lesson gives you the language for that.

About this lesson

Your team launches an AI Data Analyst. It takes a question in plain English, writes SQL, runs it, and hands back a chart and a short summary. It passes code review. It works in the demo. Then say that in the first week about 15% of users report that the same question gives different answers, and engineering finds nothing wrong.

Nothing is broken. A language model samples every answer, so the same input can come back different with every component doing what it was designed to do. The testing habits you already have, run it once, assert the exact output, release, do not carry over.

You look at one query run five times, a made-up example, and then at the 47 questions the course’s v0 traces hold five runs of. Then you put two numbers on it. pass@k asks whether at least one of k runs succeeded. That tells you the system is capable. reliable@k asks whether all k runs succeeded. That tells you what your users experience. In the v0 traces pass@5 is 1.0 and reliable@5 is 0.36. The gap between the two is the number a product manager needs to see, and its size tells you whether you have a consistency problem or a capability problem.

You also see why temperature zero reduces variance without removing it, why a benchmark score screens a model but does not make the ship decision, and why “evaluation” means three different things to a PM, a data scientist and an engineer. The course treats it as evidence for four decisions: ship, ramp, hold, or roll back.

The practice is a Non-Determinism Report you do from the slides: pass@5 and reliable@5 for the five-run demo query and for the 47 repeated v0 questions, where each result lands on the gap table, and two sentences on what the gap means for the ship decision.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→