AI Evals for Product DevelopmentL4.3 · 01

Cost-aware
evaluation

Spending a fixed evaluation budget where the decisions are
Where we left offL4.3 · 02

Which metrics did you classify as blocking, and why does that matter for a release?

Blocking
Has to pass its threshold before the release goes out.
Optimization
Tracked over time. Doesn't stop a release.
Answer from memory. Name two of each for the AI Data Analyst.
The scenarioL4.3 · 03

Running every metric on every query costs about eleven times what the product costs to run.

Metric
Type
Cost per eval
Per day
Per month
PII detection
B
$0
$0
$0
Response latency
O
$0
$0
$0
SQL correctness
B
$0.001
$10
$300
Retrieval precision
O
$0.01
$100
$3,000
Narrative faithfulness
B
$0.05
$500
$15,000
Chart appropriateness
O
$0.05
$500
$15,000
10,000 queries a day. Inference costs $3,000 a month. Leadership: "Cut evaluation cost by 90%, and don't lose visibility into quality."
The first instinctL4.3 · 04

Sampling 10% of everything saves money in the wrong places.

Rare, high-stakes failures
PII leaks happen in about 0.1% of traffic. A 10% sample sees one query in ten, so it sees about one leak in ten.
Routine traffic
Low-stakes queries get the same coverage as the dangerous ones.
Uniform sampling treats every query as equally important.
Strategy oneL4.3 · 05

Sample more where a missed failure would hurt more.

10%
Safety-criticalPII exposure, policy risk, high-value users
30%
Business-criticalOnboarding, core use cases, big customers
50%
GeneralRoutine analytics and reporting
10%
TailLow-risk, low-volume queries
Share of the 10,000 daily queries in each segment.
Strategy twoL4.3 · 06

Blocking metrics get paid for first, because they gate the release.

Blocking
SQL correctness, narrative faithfulness, PII detection. Higher coverage, the more accurate judge, checked more often.
Optimization
Retrieval precision, chart appropriateness, response latency. Sampled, cheaper methods, updated less often.
The classification from 4.1 is now a budget rule.
The budget mathL4.3 · 07

Work backward from the budget to the coverage you can afford.

monthly cost = Σ cost per eval × daily queries × coverage × 30
$3,000 a monthThe evaluation budget
10,000 queries a dayThe traffic you are sampling from
Six metricsEach with its own cost per evaluation
The demo: the coverage matrixL4.3 · 08

Blocking metrics on high-risk segments get the most coverage.

Metric
Type
Safety 10%
Business 30%
General 50%
Tail 10%
PII detection
B
100%
100%
100%
100%
SQL correctness
B
100%
50%
30%
10%
Narrative faithfulness
B
100%
50%
20%
5%
Retrieval precision
O
50%
20%
10%
5%
Chart appropriateness
O
20%
10%
5%
1%
Response latency
O
100%
100%
100%
100%
Predict before the demoL4.3 · 09

Will stratified sampling alone get you under $3,000 a month?

Your prediction
Yes or no, and roughly what the monthly bill comes to.
To estimate it
Weight each metric's coverage by the traffic in each segment. Then use the formula.
Full coverage was $33,300. Write your number before we advance.
The demo: stratified sampling, costedL4.3 · 10

Stratified sampling alone doesn't get you under budget.

Where the money still goes
Narrative faithfulness. It's blocking, it costs five cents a call, and it runs on every safety-critical query and half the business-critical ones.
Why you can't just sample it less
It gates the release, and business-critical traffic is 30% of volume. You need each evaluation to cost less.
Strategy threeL4.3 · 11

Run cheap checks first. Pay for the expensive judge only when the cheap one isn't sure.

$0
Rule checksDoes a narrative exist? Is it long enough? Does it reference the chart? Obvious failures stop here.
$0.003
Small judgeA small model scores the clear cases. If it's confident, you keep its score.
$0.05
Frontier judgeOnly the cases the small judge wasn't confident about.
This is a judge cascade.
The demo: the cascade, costedL4.3 · 12

Most sampled narratives never reach the expensive judge.

15%
Fail the rulesMarked as failures at no cost.
60%
Resolved by the small judgeConfident score at $0.003.
25%
EscalatedScored by the frontier judge at $0.05.
A made-up split for the demo. What is the average cost per evaluation now, and what does the monthly bill come to?
PracticeL4.3 · 13

Tune the cascade on paper and write the Cost Allocation Plan.

Base version, everyone
The demo split escalates 25%. Pick a new split for rules, small judge and frontier that escalates under 20%. Work out the new cost per evaluation and the monthly bill with the formula. Write what you'd check before trusting the lower cutoff. Fill in the Cost Allocation Plan.
Extended version, DS and engineering
Cost out a two-tier cascade (rules, then frontier). Rerun the plan for three scenarios: budget cut in half, traffic doubles, a new five-cent judge metric is added. Write how you'd use the cases where the two judges disagree.
Common mistakesL4.3 · 14

Three ways a cost plan loses the quality signal.

Mistake
What happens
Do this instead
Uniform sampling
Rare safety failures are sampled at the same rate as routine traffic, so most of them go unseen.
Stratify by risk. Blocking metrics get full coverage on safety-critical traffic.
Cheap judge everywhere
A small model that misses a share of real failures gives you a low bill and false confidence.
Measure agreement with the frontier judge before trusting the cheap tier.
Safety segment drawn too wide
Half the traffic counts as safety-critical, and 100% coverage there blows the budget.
Narrow it to the queries where a miss is really costly.
Knowledge checkL4.3 · 15

Three judgment calls. Write your answer before you read on.

01Your budget covers 20% average coverage for the narrative judge. Onboarding is 5% of traffic, the core use case 70%, the tail 25%. How do you split it, and why?
02Next quarter the evaluation budget drops from $3,000 to $1,500. What do you cut first, and what do you keep?
03The small judge agrees with the frontier judge 90% of the time, and you're still over budget. What can you change, and what does it cost you?
Next lessonL4.3 · 16

You can afford the evaluation now. Next, you define the segments it depends on.

Stratified samplingCoverage follows the cost of a missed failure. Full coverage on safety-critical traffic for every blocking metric.
Blocking firstThe metrics that gate the release get the coverage and the more accurate judge.
Judge cascadeRules, then a small judge, then a frontier judge only for the unclear cases.
AI ANALYST LAB · aianalystlab.ai