Week 1: Foundations and Economics · Lesson 1.4

Distributional thinking for AI quality

How do we reason about AI behavior when outputs vary across inputs, contexts, and users?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome back. Over three lessons you've built up a picture of what the AI Data Analyst actually does. In 1.1 you measured non-determinism. In 1.2 you mapped where the system can fail. In 1.3 you discovered the failure categories that are really there. Now we need to know whether the evidence is good enough to make a call. A single success rate is not evidence. It's one number from one sample. Today we treat quality as a distribution. Error bars on every metric, a breakdown by segment, and a way to separate the system's variance from your evaluator's. Then we connect all of that to a ship, ramp or hold decision. Let's start with where 1.3 left you.

About this lesson

Say you have 100 traces and 62 of them look acceptable. Your PM asks whether 62% is good enough to ship, and you cannot answer. You do not know whether a re-run would give 58 or 67. You do not know whether the system fails badly on 5% of cases or is mediocre on 38%. You do not know whether the variation comes from the system or from your evaluation method.

Three tools fix that. Bootstrap resampling puts a confidence interval on any metric without assuming anything about the shape of the data: resample your outcomes with replacement, recompute, repeat, and read the 2.5th and 97.5th percentiles. Segment breakdowns show what an aggregate hides. In a made-up example an aggregate pass@5 of 0.78 hides multi-table joins at 0.45, and in the v0 traces reliable@5 runs from 0.12 on trend questions to 0.50 on comparisons. Variance decomposition separates the system’s randomness from the evaluator’s noise, because the two need different fixes.

Then the decision. Full rollout when the intervals are narrow and every must-pass metric clears its threshold in every segment. Gradual ramp when the intervals are wider but the direction is positive. Hold or roll back when an interval includes the threshold, because you cannot tell pass from fail.

You compute pass@5 and reliable@5 for five trials of one query using the two formulas, bootstrap the interval, and see why five trials is not enough evidence. Then you look at the 47 questions the v0 traces hold five runs of. pass@5 reaches 1.0, reliable@5 falls to 0.36, and most of the variation is the same question passing on one run and failing on the next.

The practice is a distributional quality report: the gap between pass@5 and reliable@5 for each segment, a ship criterion written with a threshold and an interval width, and a ship, ramp or hold call per segment with the evidence behind it.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→