Segmentation strategy for AI systems
Where does the system work well, where does it fail, and how do we structure segments to see it?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
The AI Data Analyst’s SQL correctness is 86 percent, just above its 85 percent threshold, so the team ramps it to half of its users. Days later the complaints start. The number was right. It was an average, and the easy queries that make up most of the traffic were carrying it.
This lesson is about breaking one metric into the groups that matter: query complexity, domain, user role, or anything else you can see before the system answers. Splitting by whether the query succeeded doesn’t count, because that only restates the metric. For each segment you report the rate, the sample size and a confidence interval, and you flag any segment with fewer than 20 traces as too small to act on alone.
Then you rank what you found by volume times severity. Big segments with real problems come first, and safety-critical segments like finance get pulled up the list even when they’re small.
In the worked example you predict what a split by query complexity will show, then look at the segment rates and their intervals. One segment turns out to be badly broken, and its interval is so wide that the first step is collecting more traces before sizing the fix.
In the practice you rank four segments by volume times severity, one of them too small to act on, then check your ranking against ours. You pick three dimensions for your own product and assemble a segment prioritization schema that feeds the release criteria in 4.6. On the extended track you work out how much each segment adds to the average, design a dimension you’d derive from the raw trace, and see how few traces are left when you combine two dimensions.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→