AI Evals for Product DevelopmentL6.6 · 01

Communicating
AI product impact

Writing the same evidence for the executive, the PM and the engineer, so each of them can act on it
Where we left offL6.6 · 02

In 6.3 you scored every fix on seven dimensions. Which one was hardest to score objectively?

User harm
Frequency
Business criticality
Confidence
Fixability
Time-to-learn
Reversibility
Write your answer and one reason before we move on.
The scenarioL6.6 · 03

Monday morning. Two requests, one set of evidence.

QualitySQL success +2.7 points, task completion +5.4, retrieval precision +7.9
Latency+17.6% on average, 10.6% of users over the 2-second SLA
Cost per query+20.2%
The callHold. Both guardrails failed
The PM"Can you put together an exec brief on why v2 isn't going out?"
The engineering manager"I need a regression update on latency and cost. What's the scope?"
AudienceL6.6 · 04

Each reader is making a different decision with the same evidence.

Executive
Should we keep investing in this direction?

Needs what changed for users.
Product
What do we build next, and for which users?

Needs the result by segment.
Engineering
What do I fix first, and how fast?

Needs the component, the metric and the latency numbers.
Claim disciplineL6.6 · 05

Match the strength of each claim to the strength of the evidence.

Past tense"Improved," because the experiment is over and describes what happened.
A specific number2.7 points, instead of "significantly" or "dramatically."
ScopeWhich users, which experiment, which period.
UncertaintyA confidence interval, a p-value or a sample size.
We'll call this claim discipline. Every metric claim in a brief gets all four.
Claim disciplineL6.6 · 06

The same results, written two ways.

Overclaimed
With claim discipline
"v2 improves accuracy."
SQL success rose 2.7 points, from 71.9% to 74.6% (95% CI 2.3 to 3.0), in our 10,000-user experiment.
"Task completion is significantly better."
Task completion rose 5.4 points (95% CI 5.0 to 5.8), with 5,013 users on v2 and 4,987 on v1.
"Latency slightly increased."
Average latency rose 17.6%, from 1,416 to 1,665 ms. Users averaging over the 2-second SLA went from 1.6% to 10.6%.
"The change is positive overall."
Every user group gained on SQL success. Every group also got 16 to 20% slower and 19 to 22% more expensive, past both guardrails.
Distribution reportingL6.6 · 07

An average can hide who got worse.

Average
One number. Can't tell you whether everyone got a little slower or a few people got a lot slower.
p10
The fastest 10% of queries. What the best case looks like.
p50
The median. What a typical user sees.
p90 and p95
The slowest 10% and 5%. Where SLAs break.
By segment
Which group of users those slow queries belong to.
For latency, load time, error rate and anything else with a tail, report the median, the tail and the segments.
Predict before the demoL6.6 · 08

Which result do you lead the exec brief with?

QualitySQL success +2.7 points, task completion +5.4, retrieval precision +7.9
Latency+17.6% on average, 10.6% of users over the 2-second SLA
Cost per query+20.2%
The callHold. Both guardrails failed
Write the first sentence of your impact brief. Keep it. You'll rewrite it at the end of the practice.
The demo: step 1L6.6 · 09

Before you write "all users," check each segment.

What we compute
The treatment effect, v2 minus v1, for each metric within each user role: PM, DS, engineering, executive. Each with a 95% confidence interval.
What we look for
Does every segment's interval sit above zero? Then run the same check on latency and cost, where "every group" can be true of the regression too.
The demo: step 2L6.6 · 10

Then split latency and cost by segment, and check each one against its limit.

What we compute
Latency percentiles and the share of users over the 2,000 ms SLA, by user role and by question complexity. Median and average cost per query.
What we look for
Is the breach the same for everyone, or worse in one group? The answer changes who has to act.
The demo: step 3L6.6 · 11

Every line in the brief says what changed, how much, for whom, and how sure you are.

SQL successRose 2.7 points, from 71.9% to 74.6% (95% CI 2.3 to 3.0), across 10,000 users. Every role gained; executives least, 1.5 points.
Task completionRose 5.4 points (95% CI 5.0 to 5.8).
LatencyAverage up 17.6% against a 10% guardrail. 10.6% of users now average over the 2-second SLA; 31% of users who mostly ask complex questions.
Cost per queryAverage up 20.2% (95% CI 18.4% to 22.1%) against a 15% guardrail. About 2% of v2 users average over 8 cents a query.
The demo: step 4L6.6 · 12

The same evidence, written three times.

For the executiveWhat changed for usersv2 got more questions right, but it was slower and cost about a fifth more per query, past the limits we set before the test. It's on hold while the team fixes speed and cost, and then it gets tested again.
For the PMWhere to go nextEvery user group gained, executives least, and every group got slower and more expensive. The slowness lands on complex questions: 31% of those users average over 2 seconds. Recommendation: fix complex-question latency first, then rerun.
For engineeringWhat to fixMedian user latency up 197 ms; p99 from 2.05s to 4.55s. About 2% of v2 users average over 3 seconds and about 2% over 8 cents a query, and no v1 user does. Three quarters of users over the 2-second SLA mostly ask complex questions.
Two documentsL6.6 · 13

Every release that changes quality gets two write-ups from the same evidence.

Impact briefFor executives and productSummary with the main result and the deployment scopeKey results with claim disciplineWhat it means for usersTradeoffs: what improved, what got worse, the mitigation
Regression updateFor engineering and on-callProblem and severity, with distribution detailRoot cause as currently understoodRepro steps: the query pattern that triggers itMitigation, fix plan and rollback readiness
PracticeL6.6 · 14

Write the impact brief and the regression update.

Base version, everyone
Write the full impact brief from the numbers in this lesson, every claim with claim discipline. Add a distribution section that marks each segment's SLA status and names the worst one. Write the full regression update. Then rewrite your first sentence from the prediction and compare the two.
Extended version, DS and engineering
Write the rule that labels a segment's SLA status, with the cutoffs, and apply it to every segment in the lesson. Explain what the gap between median and average cost tells engineering. Describe, in words, the two repro pulls: the slowest users overall and among complex-question users.
Common mistakesL6.6 · 15

Three ways impact reporting goes wrong.

Mistake
What happens
Do this instead
Overclaiming
"v2 improves accuracy." The CTO asks: by how much, for whom, and what's the margin of error?
Past tense, a specific number, scope and uncertainty on every metric claim.
One report for everyone
Engineering gets a brief with no repro steps. Executives get percentiles and can't tell whether to back the hold.
Write both documents, each for the decision its reader makes.
Average only
"Latency up a quarter of a second" reads as manageable. It hides the users now waiting over 3 seconds.
Report the median, the tail and the segments.
Knowledge checkL6.6 · 16

Three judgment calls. Write your answer before you read on.

01A colleague says "v2 dramatically improves retrieval quality" is too strong. Rewrite it with past tense, a number, the users it applies to and how sure you are.
02A different experiment: quality rose 8% overall, but executive users, 10% of the population, saw a 3% regression. You're writing the brief for the CEO. Do you mention it, and how?
03Three days into a ramp, p95 latency jumps by half a second and support tickets triple. The engineering manager wants a regression update. Which three things go first, in order?
End of the courseL6.6 · 17

Before any result leaves your hands, check three things.

Claim disciplinePast tense, a specific number, the scope and the uncertainty, on every metric claim.
The distributionThe median, the tail and the segments, for anything where users can have very different experiences.
The readerWhat decision this person makes next, and whether the document gives them what they need to make it.
AI ANALYST LAB · aianalystlab.ai