AI Evals for Product DevelopmentL2.6 · 01

Instrumentation
at scale

How much to log when you can't afford to log everything
Where we left offL2.6 · 02

Your 2.3 trace spec goes live at 100,000 queries a day. What breaks first?

The platform bill
Ingestion, indexing, querying and retention for every trace.
The evaluators
An automated check that calls an LLM on every stored trace.
The users
Logging that sits in the request path adds waiting time.
Pick one and write down why before we look at the numbers.
The scenarioL2.6 · 03

Full logging at production volume is a line item somebody will ask about.

Volume
100,000 queries a day, one trace each. About 3 million traces a month.
Assumed all-in cost per trace
About $0.02 on a managed platform, covering ingestion, indexing, querying and retention.
Multiply them before anything else. That's the bill before a single evaluator runs.
The methodL2.6 · 04

A sampling plan has to satisfy three constraints at once.

Evidence sufficiency
Enough sampled traces to catch the failures that matter.
Cost
Platform budget, retention, and evaluator spend.
Operational reliability
Logging never slows a user down and never loses a trace.
A plan that fixes one of these and ignores the others fails in production.
EvidenceL2.6 · 05

The rarest failure you care about sets the minimum sample.

Common failures
Retrieval errors15%
Hallucination15%
SQL syntax12%
Empty results8%
Rare failures
Multi-turn context loss2%
Security violations1%
Numeric overflows0.5%
Example base rates for the AI Data Analyst, organized like your Lesson 1.3 taxonomy.
CostL2.6 · 06

Cost rises in a straight line with the sampling rate.

Full logging
Every trace stored. The whole $0.02 × 3 million.
Sampling at rate r
You pay r times the full cost. Halve the rate, halve the bill.
So the question becomes: what is the smallest rate that still meets the evidence constraint?
OperationsL2.6 · 07

Trace collection has to stay out of the user's request.

Request handledResponse goes to the user right away
Trace queuedAdded to a background queue without waiting
Workers writeBackground workers send traces to storage
Drain on shutdownThe queue empties before a worker stops
If logging is in the request path, a slow disk or a network hiccup becomes a slow answer.
Predict before the demoL2.6 · 08

What sampling rate do you need to catch a 0.5% failure with 90% confidence?

The setup
100,000 queries a day. The seven failure categories from the base-rate slide. The rarest happens in 0.5% of queries, one in 200.
The target
If the failure is happening, you see it in a day's sample at least 9 times out of 10.
Write down a sampling rate before we advance.
The demo: detection mathL2.6 · 09

The floor is a count of sampled traces, and it comes from one formula.

n = ln(1 − confidence) / ln(1 − base rate)
What n means
Sampled traces you need to see a failure at least once, at that confidence.
Turning n into a rate
Divide n by daily volume. The rarest failure gives the highest rate.
Floor and depthL2.6 · 10

Seeing a failure once and being able to debug it are different requirements.

Detection floor, about 0.5%
A few examples a day of the rarest failure. It proves the failure exists.
Operating rate, 10%
Dozens of examples a day. Enough to find the pattern and track the rate.
Detection coverage stops improving at the floor. Everything above it buys depth.
Dynamic samplingL2.6 · 11

Raise the rate when the risk goes up, then bring it back down.

Steady state10%The base rate
After a deploy50%First 48 hours
Error spike25%24 hours after error rate passes 5%
Seasonal surge15%Dec 20 to 31
Each rule needs a trigger, a temporary rate, a duration, a ramp-down and a cost estimate.
Two decisionsL2.6 · 12

Which traces you store and which ones you evaluate are separate policies.

Trace sampling
What gets stored. Driven by platform cost.
Eval sampling
Which stored traces get an automated quality check. Driven by compute cost.
If eval sampling is written as "all stored traces," every change to trace sampling changes the evaluator bill too.
RetentionL2.6 · 13

Keep full detail while people debug. Keep summaries after that.

7 days
Full traces. Every field, every span.
90 days
Summary stats. Counts, latency percentiles, error rates.
Indefinitely
Trends only.
Compare it with keeping every sampled trace in full for 90 days.
PracticeL2.6 · 14

Write a sampling strategy your infrastructure team could implement.

Base version, everyone
Write three risk-triggered rules. For each one: trigger, temporary rate, duration, ramp-down and cost, using the per-event costs on slide 11. Say how many deploys and error spikes you expect in a month, add them to the base cost, and check the total against a budget you state.
Going further
Stratify: sample multi-turn sessions and executive users at 20%, PM users at 10%, simple lookups at 5%. State a traffic mix, compute the effective rate, and check the rarest failure is still covered.
The spec: base rate, detection table, cost breakdown, dynamic rules, retention, operational requirements, review cadence.
Common mistakesL2.6 · 15

Three ways a sampling plan goes wrong.

Mistake
What happens
Do this instead
Pick a rate without the volume
A rate that works at one traffic level leaves the rarest failure invisible at another.
Calculate the sampled count you need, then divide by your real daily volume.
One rate for every situation
Thin evidence during launches and incidents, when failures matter most.
Write risk-triggered rules with a duration and a ramp-down.
Tie eval sampling to trace sampling
The evaluator bill moves every time storage changes.
Set the two rates separately, on purpose.
Knowledge checkL2.6 · 16

Two judgment calls. Write your answer before you read on.

01100,000 queries a day, ten failure modes: nine between 5% and 10%, one at 0.3%. You need 90% confidence of seeing all of them, and the budget allows 5% sampling. Do you (a) sample 5% uniformly, (b) stratify toward risky queries, (c) ask for more budget, or (d) accept the 0.3% failure as invisible?
02You store 10% of traces and run a $0.001 judge on every stored trace. What does evaluation cost per day? What happens when trace sampling goes to 20%?
Next lessonL2.6 · 17

A plan you can defend on evidence, cost and operations. Next, what users actually need.

Strategy
Evidence
Cost
Operations
Full logging
Everything
Fails the budget
Simple
10% uniform
Covers the rarest failure with depth
A tenth of full
Simple
10% base plus dynamic rules
Extra depth in risky windows
Base plus per-event cost
More rules to run
AI ANALYST LAB · aianalystlab.ai