AI Evals for Product DevelopmentL1.1 · 01

Why AI evaluation
is different

Distributional thinking and the gap between capability and reliability
The scenarioL1.1 · 02

Everything passes review. Users still get different answers to the same question.

What the team checked
Code review passed, the API returns 200s, the SQL is valid, and the data pipeline is healthy.
What users reported in week one
"The same question gives different answers." Say about 15% of users said this.
The running example for this course is an AI Data Analyst: a natural language question in, SQL, a chart and a written summary out.
The shiftL1.1 · 03

One run tells you the system can do this. It does not tell you the system will.

Traditional testing asks
Does this input produce the correct output? Run it once, assert equality, and release.
AI evaluation asks
How often does this input produce a correct output? Run it many times, count, and reason about the pattern.
The question changes from "does it work" to "what fraction of the time does it work."
Two metricsL1.1 · 04

pass@k measures capability. reliable@k measures consistency.

pass@k
At least one of k trials succeeds. Can the system do this task at all?
reliable@k
All k trials succeed. Does the system do this task every time?
The gap between them is the difference between what the system can do and what your users experience.
A common fix that doesn't workL1.1 · 05

Temperature zero reduces variance. It does not remove it.

What people expect
Set temperature to 0, the randomness setting, and the model becomes deterministic.
What production models show
Re-runs still differ at temperature 0, because floating-point math differs across GPU types, batch sizes and parallelism.
You cannot assume the variance away. You have to measure it.
Two sources of varianceL1.1 · 06

The system can vary. So can the thing you use to judge it.

System variance
Re-run the same query through the AI Data Analyst and get different SQL or a different summary.
Judge variance
Score the same output twice with an LLM judge, or two human reviewers, and get different scores.
This week is about system variance. Judge variance comes in Week 3, when we build scoring systems.
One word, three meaningsL1.1 · 07

When a PM says "evaluation" they mean rubrics. A data scientist means benchmarks. An engineer means tests.

Role
What "evaluation" means
The question they ask
PM
Rubrics, quality criteria
Does it feel high quality to users?
Data scientist
Benchmarks, scored metrics
How does it compare to baseline?
Engineer
Test suites, CI checks
Did the tests pass?
This course treats evaluation as evidence for four decisions: ship, ramp, hold, or roll back.
BenchmarksL1.1 · 08

A benchmark tells you the model is in the ballpark. It does not tell you it works for your users.

Google Bard, February 2023
Gave a factually wrong answer about the James Webb Space Telescope in its launch demo. Alphabet lost roughly $100B in market value in a day.
IBM Watson for Oncology
Promising results in controlled settings. Internal documents reported in 2018 showed unsafe and incorrect treatment recommendations.
Benchmarks screen. Product evaluation on your users' real queries decides.
Predict before the demoL1.1 · 09

Here is one query, run five times. Write down three predictions first.

"What was Q4 revenue?" Five runs. Temperature 0.7. A made-up example, not from the course data.
01How many of the five outputs will be identical?
02Will any output be factually wrong, a wrong number or a wrong time period?
03Will the generated SQL be the same across all five runs?
The demoL1.1 · 10

Five runs of the same query.

Run 1SELECT SUM(revenue) FROM sales WHERE quarter = 'Q4'
Run 2SELECT total_revenue FROM quarterly_summary WHERE q = 4
Run 3SELECT SUM(amount) FROM transactions WHERE date >= '2025-10-01'
Run 4SELECT SUM(revenue) FROM sales WHERE quarter = 'Q3'
Run 5SELECT revenue_total FROM revenue_summary WHERE period = 'Q4-2025'
Read the five WHERE clauses before we discuss them.
Put the numbers on itL1.1 · 11

What do the two metrics say about this query?

pass@5
Did at least one of the five runs succeed?
reliable@5
Did all five runs succeed? Across a query set, what fraction of queries did?
Work it out for the Q4 query. Then for the v0 traces: 47 questions asked at least five times, first five runs each.
Reading the gapL1.1 · 12

The size of the gap tells you what kind of problem you have.

Metrics
Gap
Decision
What it means
pass@5 = 1.0, reliable@5 = 0.8
0.2
Ship candidate
Capable and mostly consistent.
pass@5 = 1.0, reliable@5 = 0.4
0.6
Hold
Capable but unreliable. Needs mitigation first.
pass@5 = 0.4, reliable@5 = 0.0
0.4
Hold
Not capable. A model quality problem, not a consistency problem.
The rows are reading guides. The course's v0 traces land in the second row.
PracticeL1.1 · 13

Build a Non-Determinism Report.

Base version, everyone
Mark each of the five runs on slide 10 right or wrong for the question asked. Compute pass@5 and reliable@5 for that query. Then take the v0 result: 47 repeated questions, all 47 worked at least once, and 17 worked on all five runs. Place both results in a row of slide 12 and write two sentences to your PM on what the gap means for shipping.
Going further
Sort the differences between the five runs into ones that change the answer and ones that don't. Decide how you'd score runs like these automatically: compare the SQL text, or run each query and compare what comes back. Write down one case where the other method gets it wrong.
Common mistakesL1.1 · 14

Three ways teams ship on the wrong number.

Mistake
What happens
Do this instead
Ship on pass@k alone
Say pass@5 = 0.95 and reliable@5 = 0.3. About 65% of queries pass on some runs and fail on others.
Always compute both. If the gap is above 0.2, add mitigation before shipping.
Assume temperature 0 is deterministic
One test per query passes. In production, re-runs still differ because of hardware differences.
Measure variance by running multiple trials.
Use a benchmark as the ship decision
Say 92% on a standard benchmark, then failures on real user queries.
Benchmarks screen. Product evaluation decides.
Knowledge checkL1.1 · 15

Three judgment calls. Write your answer before you read on.

01A system has pass@5 = 0.92 and reliable@5 = 0.41. A PM asks whether it's ready to ship to all users. What do you say, and why?
02You run one query three times at temperature 0.7 and get three different SQL queries. All three return the same result set. Is that a failure?
03A benchmark shows 89% accuracy. Your PM says "great, let's ship." What question do you ask first?
Next lessonL1.1 · 16

You can now measure how inconsistent a system is. Next, you map where it fails.

What you can do now
Run a query many times, compute pass@k and reliable@k, read the gap, and connect it to a ship, ramp or hold decision.
Lesson 1.2: the evaluation surface map
Every place an AI feature can fail, from the SQL to a misleading summary, classified by severity and matched to an evaluation method.
AI ANALYST LAB · aianalystlab.ai