AI Evals for Product DevelopmentL4.5 · 01

Driver analysis

Explaining the gaps between segments and choosing what to change first
Where we left offL4.5 · 02

Which dimension did you put first in your 4.4 schema, and why?

What you chose
Query complexity, domain, user role, or a dimension of your own.
What today tests
Whether that dimension explains the most of the gap in the overall number.
Write down your top dimension and one sentence on why you ranked it first.
The scenarioL4.5 · 03

SQL correctness is 74% and the blocking threshold is 70%. Release now, or spend two more weeks?

If failures are spread evenlyThe system is weak at everything. Something core needs work: the prompts, the retrieval, the model.
If failures are concentratedOne kind of query is failing. You know where to aim, and a targeted fix could move the overall number a lot.
Both look like 74% from the outside.
Why rankL4.5 · 04

The worst-looking segment isn't always the one worth fixing.

Multi-hop reasoningQuestions that pull information across several tables.
40% correct · 2% of queries
Complex single-domainHarder questions within one area of the business.
65% correct · 30% of queries
Which one would you spend two engineering weeks on?
The impact scoreL4.5 · 05

Rank segments by how far they drag the overall number down.

impact = (baseline - segment rate) × segment share
Baseline is the overall rate. Share is the segment's fraction of queries. The result is in percentage points of the overall number. A segment above the baseline gets a negative score. It's pulling the number up.
The workflowL4.5 · 06

Three steps from one overall number to a list of things to try.

1. Segment and measureSplit by dimension, compute the metric and an interval per segment. Flag any segment under 30 queries.
2. Rank by impactCompute impact scores and sort, highest first.
3. Write hypothesesFor the top segments: a likely cause, a proposed fix, the expected gain, and how you'd check it worked.
A warningL4.5 · 07

Driver analysis finds associations. Every finding is a hypothesis until you test it.

What you see
Executive users get lower narrative quality than analysts.
What might be going on
Executives ask harder questions. The driver could be complexity, and user role just comes along with it.
Week 5 is where you test these hypotheses with experiments.
Predict before the demoL4.5 · 08

500 test queries, 250 simple and 250 complex, 74% correct overall. What will you find?

01Simple query correctness: ___%
02Complex query correctness: ___%
03Your conclusion: failures spread evenly, or complexity drives them?
The demo: split by complexityL4.5 · 09

Complexity drives the gap.

Simple queries
90%
95% CI [86%, 93%] · n = 250
Complex queries
58%
95% CI [52%, 64%] · n = 250
The demo: split by domainL4.5 · 10

Split by domain, inventory looks worst. Is domain the real driver?

Domain
Correctness
n
Inventory
65%
160
Sales and customer
78%
340
Are inventory questions harder, or do inventory users ask more complex questions?
The demo: both at onceL4.5 · 11

Complex queries fail everywhere, and much more in inventory.

Domain
Simple
Complex
n per cell
Inventory
85%
45%
80
Sales and customer
92%
64%
170
When two factors together produce a worse result than either one alone, that's an interaction effect.
PracticeL4.5 · 12

Score the cells, rank by impact, check the interaction, write the hypothesis.

Base version, everyone
From the tables on slides 10 and 11, compute the impact score for the inventory domain and for all four cells. Say whether the complexity driver holds in both domains. Rank the top two cells by impact, and fill the intervention hypothesis template for the worst one.
Extended version, DS and engineering
Say a fix lifts complex inventory queries to 64%, level with complex queries elsewhere. Work out the new overall rate. Write a second hypothesis, for complex queries outside inventory. Plan how you'd check that the drivers hold month to month.
The briefL4.5 · 13

A driver analysis brief: two findings and a hypothesis you can test.

Biggest dragComplex queries, the single dimension with the largest impact score.
Where to aimComplex inventory queries, the interaction cell.
EvidenceSegment rates with intervals, and the ranked impact table.
Likely causeYour theory for why that cell fails.
Proposed fix and expected gainWhat you'd change, and how much it should move the overall number.
CheckHow you'll know it worked, decided before you build it.
Common mistakesL4.5 · 14

Four ways driver analysis goes wrong.

Mistake
What happens
Do this instead
Too many slices
20 or more segments, some with under 10 queries. A new sample reorders the list.
Flag segments under 30 queries. Merge sparse ones.
Correlation read as cause
"Executives get worse answers," when executives just ask harder questions.
Split by pairs of dimensions before blaming one factor.
Ranking by gap alone
Multi-hop at 40% looks urgent at 2% of queries.
Rank by impact score, gap times share.
Stale drivers
The analysis used month-old data and usage has changed.
Check that the drivers hold over time before you build the fix.
Knowledge checkL4.5 · 15

Three judgment calls. Write your answer before you read on.

01Simple queries: 95% correct, n = 300. Complex: 60% correct, n = 200. What's the overall rate, and what are the two impact scores?
02Your PM asks whether to delay the release to improve complex queries. What evidence supports your answer, and what else would you want to know?
03Complex inventory queries are your worst cell. Write one intervention hypothesis with all four parts: cause, fix, expected gain and check.
Next lessonL4.5 · 16

Driver analysis tells you what to fix first. Next: how you'll know it's good enough.

Segment and measureRates and intervals. Under 30 queries: merge or collect more.
Rank by impactGap from baseline times share.
Write hypothesesCause, fix, expected gain, check.
Test in Week 5Run the experiment before committing the work.
Run it againConfirm the gain and find the next driver.
AI ANALYST LAB · aianalystlab.ai