AI Evals for Product DevelopmentL5.3 · 01

Experiment design
for stochastic systems

Online tests for a system that gives different answers to the same question
Where we areL5.3 · 02

v2 did better offline. Your PM asks if it can go to every user.

What you have
A higher corrected pass rate for v2 on the held-out test set. Frozen questions, frozen schema, no follow-ups.
What you don't have
Any evidence about real users: follow-ups, rephrasing, abandoned sessions, questions your test set never covered.
The answer today is "not yet." The experiment is how you get to a real answer.
Why it's harderL5.3 · 03

Same question, three runs, three different queries.

Question"Show sales by region"
Run 1SELECT region, SUM(sales) FROM ...
Run 2SELECT region, total_sales FROM ... GROUP BY 1
Run 3SELECT r.name, SUM(s.amount) FROM ... JOIN ...
The system's own randomness sits on top of the randomness in who ends up in each group.
Before any dataL5.3 · 04

Decide what you're measuring and who gets randomized.

What you're measuring
The average effect of v2 on sql_success_rate. One sentence, written before the test starts.
The unit of randomization
The user. Each user sees v1 or v2 for the whole experiment.
User, session or query: the choice changes both what can go wrong and how many people you need.
The hidden assumptionL5.3 · 05

Randomizing by user assumes one user's version can't change another user's result.

Holds
Your result depends only on whether you got v1 or v2. Two email subject lines sent to people who never interact.
Breaks
Users share something. The AI Data Analyst has a shared cache that keeps retrieved schema definitions for 60 seconds.
The formal name is SUTVA, the stable unit treatment value assumption.
MetricsL5.3 · 06

Three kinds of metric, and the guardrails double as stop conditions.

Primarysql_success_rate. The thing v2 is supposed to improve.
Secondarytask_completion_rate. Does better SQL turn into people finishing their analysis?
Guardrailsavg_latency_ms must not rise more than 10% against control. avg_cost_usd must not rise more than 15%. Breach either one and the test stops.
PredictL5.3 · 07

v2 improves retrieval. Which guardrail is more likely to move?

avg_latency_ms
Limit: no more than 10% above control
avg_cost_usd
Limit: no more than 15% above control
Think about what better retrieval does mechanically. Write your pick and one sentence of reasoning.
Sample sizeL5.3 · 08

Three inputs decide how many users each group needs.

Baselinesql_success_rate is about 0.72 in production today
Smallest effect worth detecting3 percentage points
NoiseHigher than a normal product, because the same question gets different answers
How many users per group would you guess? Write a number before we compute it.
Decision rulesL5.3 · 09

Write the call for every outcome before the results exist.

Ship
The whole 95% interval for the primary sits at or above 3 points, guardrails pass, no segment gets worse
Ramp
The interval excludes zero but reaches below 3 points, guardrails pass. Roll out to 10% and keep measuring
Hold
The interval crosses zero, guardrails pass. Collect more data or look for what's hiding the effect
Roll back
Any guardrail fails, or the primary gets significantly worse
The resultL5.3 · 10

The standard A/B result: read the interval, then the estimate.

Metric
Role
What to read
sql_success_rate
Primary
Effect, 95% interval, p-value
task_completion_rate
Secondary
Same, as supporting evidence
Which of the four rules does the result match?
Guardrails and the callL5.3 · 11

Check the guardrails before anyone celebrates the primary.

Guardrail
Limit
Status
avg_latency_ms
+10% vs control
Check against your prediction
avg_cost_usd
+15% vs control
Check against your prediction
Then apply the rule you wrote before the test and record it on the Experiment Design One-Pager.
Back to the cacheL5.3 · 12

To take the shared cache out, switch everyone at once.

Block 1All users v1
Block 2All users v2
Block 3All users v1
Block 4All users v2
Block 5All users v1
Block 6All users v2
A switchback: everyone gets the same version in each block of time. The course data uses six-hour blocks, alternating, for 30 days. Will its effect come out larger, smaller or the same?
What the switchback foundL5.3 · 13

Same system, a different design. Does the call change?

Standard A/B, split by user
Control shares the cache with treatment, so control can look better than v1 really is.
Switchback, six-hour blocks
No one in a v1 block reads v2's cache entries. Check what carries over when a v1 block follows a v2 block.
Run both results through the rules from slide 9. Does the call change?
PracticeL5.3 · 14

Design the experiment and make the call.

Base version, everyone
Use the sample size inputs from slide 8 and say whether the test had enough users. Take the effects and intervals from today's results and match each to a rule from slide 9. Work out both guardrail changes as percent changes. Fill in the Experiment Design One-Pager: effect, interval, guardrails, decision, next action.
Extended version, DS and engineering
Add the switchback result and the carryover gap to the one-pager, and write what each would have to show to change the call. Then write the design for the rerun on the fix: unit, blocks, sample size and segment checks.
Common mistakesL5.3 · 15

Four ways an experiment on an AI system goes wrong.

Mistake
What happens
Do this instead
Ignoring a guardrail because the primary won
The primary clears and a guardrail doesn't. Finance finds it on the next monthly bill.
Check every guardrail before anyone discusses the primary.
Reading a flat, underpowered result as "no effect"
A good change gets dropped because the test was too small to see it.
Run the power analysis first. Treat a wide interval as "can't tell."
Trusting the overall number
A positive average hides a segment that got worse.
Break the effect down by segment before the call.
Randomizing users who share something
The shared cache shrinks the measured effect.
Ask whether one user's version can change another's result.
Knowledge checkL5.3 · 16

Three calls. Write your answer before you read on.

01An experiment shows +2.5 points on sql_success_rate, 95% interval from -0.5 to +5.5, p = 0.10. All guardrails pass. Ship, ramp, hold or roll back?
02PM users improved by 5 points, engineering users dropped by 3, and the overall effect is +2. A colleague says the overall is positive, so ship it. What's wrong, and what do you recommend?
03The power analysis called for 8,000 users per group. The test ran with 3,000 per group and came back at p = 0.25. Can you conclude v2 has no effect?
Next lessonL5.3 · 17

Design, then analyze, then decide. Next, what a fixed v2 has to show before users see it.

Design, before dataWrite down what you measure and the unit, the three kinds of metric, the sample size, the stop conditions and the decision rules.
AnalyzeRead the effect and its interval, check guardrails as a blocking step, break it down by segment, and ask whether users shared anything.
DecideMatch the result to the rule you wrote: ship, ramp, hold or roll back.
AI ANALYST LAB · aianalystlab.ai