The cost-latency-quality frontier
How do we reason about quality improvements that change cost and latency?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
Your product team hands you three constraints. Respond in under 2 seconds at the p95. Spend no more than $500 a month at 10K queries a day. Keep SQL correctness at 85% or above. In the lesson’s scenario the current system clears the quality floor at 88% and costs about $0.08 a query, $24,000 a month, with a p95 of 3.2 seconds. Cheaper models meet the latency target and fall below the quality floor, and at 10K queries a day none of them fits the budget. No single model satisfies all three.
Cost, latency and quality are not independent dials. They form a surface. Some configurations are dominated, meaning another option is better on at least one dimension and no worse on the rest, and you replace those without debate. Others are true tradeoffs, and there your constraints break the tie. Whether GPT-4o is dominated by GPT-4o-mini depends entirely on whether the 85% floor is hard.
Routing gets you closest. Send simple queries to the cheap model and reserve the expensive one for the complex ones, and route on observable features like query length and question type rather than spending a model call to judge every query. Evaluation has a cost too. A $500 a month feature cannot carry a $5K a month evaluation pipeline, so you validate routing on a sample. Even a routed setup misses this budget, so the decision ends with taking the cost projection back to the product team.
You predict how much a partial switch saves before you see the numbers, and you see why a model that is ten times cheaper per call does not make the whole pipeline ten times cheaper. Only the share of the work you switched gets cheaper.
The practice is a benchmark table for three configurations from the scenario figures, a note on what the v0 traces could tell you (no tokens or cost at all, a measured p95 of 7.56 seconds and 78.8% SQL success), and a Model Selection Decision Template: the configuration you recommend, a pass or fail on each constraint, what you give up and why, and the risks that would change your mind.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→