AI Evals for Product DevelopmentL4.4 · 01

Segmentation
strategy

Finding where the system works, where it fails, and what to fix first
Where we left offL4.4 · 02

If multi-join queries start failing, how would you know from one overall number?

Your blocking metric
SQL correctness, the one you designed in 4.2: run the query and compare the result with a known answer.
The change
Quality drops only for multi-join queries, the ones that join three or four tables.
Answer from memory. Write one sentence.
The scenarioL4.4 · 03

SQL correctness is 86%, above the 85% threshold. The team ramps from 10% to 50% of users.

What the PM saw
86%
Aggregate SQL correctness across 1,000 traces. Above the line, so the release went ahead.
What happened next
?
Three days after the ramp, the support queue filled up with complaints about wrong numbers.
SegmentationL4.4 · 04

Break the metric down by the groups that matter for users and for decisions.

By query complexitySimple lookups, multi-join queries, advanced queries with window functions.
By domainSales, marketing, finance, operations.
By user roleWho is asking, and what they use the answer for.
Two uses: find groups doing worse than the average, and find which groups explain a change over time.
Choosing dimensionsL4.4 · 05

Segment by things you can see before the system runs.

Good dimensions
Query complexity, domain, how specific or vague the question is, time period, user tenure. All known before the model answers.
Circular dimensions
Whether the query succeeded or failed. That's the metric itself, so splitting by it tells you nothing new.
If you'd do the same thing for every value of a dimension, drop it.
Measuring each segmentL4.4 · 06

For every segment, report the rate, the sample size and a confidence interval.

Segment
Rate
n
95% CI
one row per value
metric
count
[low, high]
A segment with fewer than 20 traces gets flagged as unreliable. Its rate can move a long way on a new sample.
PrioritizingL4.4 · 07

Rank segments by volume times severity. Safety-critical segments jump the queue.

Fix firstHigh volume, high severity
Safety-critical overrideLow volume, high severity, and a wrong answer causes real harm
MonitorHigh volume, low severity
BacklogLow volume, low severity
Predict before the demoL4.4 · 08

About 1,000 traces, aggregate 86%. Split by query complexity. What will you see?

01Simple (about 600), multi-join (about 350), advanced (about 50). Which segment is highest, and which is lowest?
02Advanced queries are about 5% of traffic. How much can that segment move the aggregate?
03How much wider will the interval be for advanced (n of about 50) than for simple (n of about 600)?
The demo: the splitL4.4 · 09

The 86% average hides a segment where most answers are wrong.

Segment
Rate
n
95% CI
Simple
0.96
600
[0.94, 0.98]
Multi-join
0.78
350
[0.73, 0.82]
Advanced
0.34
50
[0.21, 0.48]
The demo: how sure are we?L4.4 · 10

With 50 traces, the advanced rate could sit anywhere in a wide range.

Simple[0.94, 0.98]
Multi-join[0.73, 0.82]
Advanced[0.21, 0.48]
00.250.50.751.0
PracticeL4.4 · 11

Rank four segments and build your prioritization schema.

Base version, everyone
Rank these by volume times severity and write an action for each: multi-join 0.78 (n 350), finance 0.71 (n 80), advanced 0.34 (n 50), weekend 0.75 (n 12). Then pick three dimensions for your own product that pass both tests from slide 5, and fill the prioritization template.
Extended version, DS and engineering
Work out how many points each segment on slide 9 adds to the aggregate. Design a dimension you'd derive from the raw trace, like number of tables joined. Then estimate how many advanced finance queries there are, and say whether you could act on that cell.
The schemaL4.4 · 12

A segment prioritization schema, ranked by volume times severity.

Segment
Rate
n
95% CI
Priority and action
Multi-join
0.78
350
[0.73, 0.82]
1. High volume and high severity, so fix it first.
Finance domain
0.71
80
[0.60, 0.80]
2. Pulled up as safety-critical.
Advanced
0.34
50
[0.21, 0.48]
3. Broken, but 5% of traffic. Collect more traces before sizing the fix.
Weekend queries
0.75
12
[0.43, 0.93]
Backlog. n under 20, so the rate can't be trusted yet.
Common mistakesL4.4 · 13

Four ways segmentation goes wrong.

Mistake
What happens
Do this instead
Too many dimensions
8 dimensions with 5 values each gives about 390,000 combinations, most with one or two traces.
Start with 3 to 5 dimensions tied to decisions.
Small-sample alarms
n = 8 with a 75% failure rate. The interval runs from 35% to 97%.
Report intervals. Flag anything under 20 traces.
Splitting by the outcome
"Succeeded vs failed" restates the metric.
Split by what you can see before the system runs.
Stopping at the pattern
"Finance is worse" gets reported and nothing happens.
Find out why. That's driver analysis in 4.5.
Knowledge checkL4.4 · 14

Three judgment calls. Write your answer before you read on.

01Aggregate SQL correctness is 87%. Simple n = 200 at 95%, multi-join n = 150 at 84%, advanced n = 50 at 62%. The PM says "87% is above our 85% threshold, let's ship." What do you say, and which evidence do you cite?
02A new segment, 8 traces, shows a 75% failure rate. Your team wants to stop other work to fix it. What do you recommend?
03A colleague proposes segmenting by eight dimensions "to be thorough." What's the risk, and where would you start instead?
Next lessonL4.4 · 15

Segmentation finds where the system fails. Next, you find out why.

DefineStart from the number and ask who it's for. Pick 3 to 5 dimensions you can see before the system runs.
ComputeRate, n and a confidence interval for every segment. Flag anything under 20 traces.
PrioritizeVolume times severity, with safety-critical segments pulled up. The schema feeds the release criteria in 4.6.
AI ANALYST LAB · aianalystlab.ai