AI Evals for Product DevelopmentL1.5 · 01

The cost, latency
and quality frontier

Choosing a model configuration under constraints
Where we left offL1.5 · 02

If pass@k is high and reliable@k is low, what does retrying cost you?

Capability high
The system can produce the right answer. At least one of k trials succeeds.
Consistency low
Most sessions include at least one failure. The instinct is "retry until it works."
Say each retry costs $0.08 and takes 3 seconds. Cost and latency grow with every retry.
The scenarioL1.5 · 03

Three constraints, and no single configuration satisfies all of them.

<2s
Latency95% of responses, the p95, under 2 seconds.
$500
Budget$500 a month at 10K queries a day.
85%
Quality floorSQL correctness, does the AI write queries that work, at least 85%.
In the lesson scenario v0 costs $0.08 a query, $24,000 a month, with a 3.2 s p95 and 88% quality. The v0 traces log no cost, a 7.56 s p95 and 78.8% SQL success.
The default playbookL1.5 · 04

Picking the best model once and never revisiting is how teams overspend.

Pick a modelThe most capable one you can afford.
Release itMove on to the next feature.
Never revisitBudget blown. Latency target missed. Quality degrading on hard queries.
Cost, latency and quality are not independent dials. They form a surface, and improving one usually costs you another.
The frontierL1.5 · 05

Some configurations are simply worse. Others are real tradeoffs.

On the frontier
Improving one dimension requires giving up another. These are the tradeoffs worth debating.
Below the frontier
Another configuration is better on at least one dimension and no worse on the rest. Replace it without debate.
The question is not "which model is best." It is "which position on the frontier satisfies my constraints."
Quick checkL1.5 · 06

GPT-4o-mini is 10x cheaper and twice as fast, and scores 6 points lower on quality.

Is GPT-4o dominated by mini? Decide before we advance.
Think about what you'd need to know to answer.
Dominated vs. true tradeoffL1.5 · 07

Constraints decide whether a comparison is dominance or a true tradeoff.

Dominated
Configuration A is dominated by B if B is better on at least one dimension and no worse on the others. Switch, with no debate.
True tradeoff
Both sit on the frontier. Improving one dimension means losing on another. Judgment enters here, and constraints break the tie.
Switching from GPT-4o to mini cuts cost 10x and halves latency, and drops quality from 88% to 82%, three points below the floor.
RoutingL1.5 · 08

Route most queries to a cheap model. Reserve the expensive one for the queries that need it.

Incoming query
Complexity filter
query length, question type, tables referenced
Say 80% simpleGPT-4o-mini at $0.008
Say 20% complexGPT-4o at $0.08
Route on observable features. Do not spend a model call judging every query.
Evaluation has a cost tooL1.5 · 09

A $500 a month feature cannot have a $5K a month evaluation pipeline.

Evaluation cost blindness
Say routing saves $500 a month on inference and needs $2K a month in automated quality checks. Net: $1,500 a month worse.
Constraint-aware evaluation
Route on observable features. Validate with a 10 to 20% sample. The evaluation cost stays inside the budget.
Budget limits how many automated checks you can afford, which limits your metric coverage.
Predict before the demoL1.5 · 10

Switch SQL generation only to GPT-4o-mini. How much does total cost drop?

A. About 10%
SQL is a small share of the cost.
B. About 50%
SQL is roughly half the work.
C. About 90%
SQL is the dominant cost driver.
v0 uses GPT-4o for SQL, narrative and charts. Mini is 10x cheaper. Narrative and charts stay on GPT-4o. Write your answer and your reasoning.
The demo: the benchmark tableL1.5 · 11

No single model satisfies all three constraints.

Configuration
$ per query
Monthly
p95 latency
Quality
All GPT-4o (v0)
$0.08
$24,000
3.2s
88%
All GPT-4o-mini
$0.008
$2,400
1.6s
82%
All Gemini Flash
$0.005
$1,500
1.2s
80%
Scenario figures for teaching, not measured from the course data. Blue meets the constraint and red misses it.
The demo: where the cost goesL1.5 · 12

Ten times cheaper per call is not ten times cheaper overall.

SQL generation, about 40% of compute
Narrative and charts, about 60%, still on GPT-4o
You only save on the fraction of the work you switched.
PracticeL1.5 · 13

Benchmark three configurations and fill in the decision template.

Base version, everyone
Build the benchmark table for GPT-4o, mini and Flash from the scenario figures on slide 11. Mark pass or fail on each constraint and mark which configurations are dominated. Write down what the v0 traces could and couldn't have told you. Complete the Model Selection Decision Template.
Going further
State a query mix, for example the share of simple lookups. Write a routing rule in plain words. Compute the weighted cost, latency and quality for the routed configuration and add it as a row. Check it against all three constraints.
v0 logs no tokens and no cost on any trace, and no model name on 328 of 500. Latency is on every trace. Work with what you have.
Common mistakesL1.5 · 14

Three ways teams make bad model decisions.

Mistake
What happens
Do this instead
Optimize one dimension
Pick the cheapest. All mini saves $21,600 a month and still misses the budget. Quality comes in at 82%, below the 85% floor.
Check all three dimensions in the template every time.
Confuse averages with tails
"Latency improved to 1.5s" is the mean. The p95 went from 3.2s to 4.1s.
Report p50, p95 and p99. Never the mean alone.
Evaluation cost blindness
Routing saves $500 a month. The quality checks cost $2K. Net: $1,500 a month worse.
Deterministic routing features. Sample-based monitoring.
Knowledge checkL1.5 · 15

Three judgment calls. Write your answer before you read on.

01A model has mean latency 1.2s and p95 latency 2.8s. The target is "under 2 seconds for 95% of queries." Does it meet the target?
02Switching from GPT-4o to mini saves $4K a month on inference and costs $1.5K a month in automated quality checks. What's the net, and what have you traded?
03A colleague says "use the cheapest model that meets our quality floor." Name two risks that heuristic misses.
Next lessonL1.5 · 16

No row clears all three. At 10K queries a day, the budget rules every one of them out.

Configuration
$ per query
p95 latency
Quality
Verdict
All GPT-4o
$0.08
3.2s
88%
Fails budget, latency
All GPT-4o-mini
$0.008
1.6s
82%
Fails budget, quality
All Gemini Flash
$0.005
1.2s
80%
Fails budget, quality
80/20 routing
~$0.022
~1.9s
~83%
Fails budget, quality
AI ANALYST LAB · aianalystlab.ai