Week 2: Instrumentation and Reliability Engineering · Lesson 2.6

Sampling traces at production scale

We can't afford to keep every trace. How many do we keep and still catch the rare failures?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome back. In 2.3 you designed a trace with a span for every stage, and in 2.5 you added the fields each system type needs. That's a lot of data per request. In development, with fifty test queries, it costs nothing. Today we take that same trace to production traffic and ask what it costs, and how much of it you really need to keep. The answer is a sampling plan: which traces you store, for how long, and which ones you spend evaluation money on. By the end you'll have one you could hand to an infrastructure team.

About this lesson

The trace you designed earlier in the week is cheap with fifty test queries. At 100,000 queries a day it is about 3 million traces a month, and at an assumed two cents per trace on a managed observability platform, full logging costs about $60,000 a month before a single evaluator runs. So you keep a sample, and the question is how big.

A sampling plan has to meet three constraints together. Evidence: enough sampled traces to see the failures you care about. Cost: platform, retention and evaluator spend, which rise in a straight line with the sampling rate. Operations: trace collection runs on a background queue, so it never slows a user’s request and never loses a trace when a server shuts down.

The rarest failure you care about sets the minimum. To see a failure at least once with 90% confidence you need ln(0.1) / ln(1 − base rate) sampled traces. For a failure in one query out of 200, that is about 460 a day, which is about half a percent of traffic. The lesson runs at 10% anyway, because a few examples a day only prove a failure exists. Dozens a day let you debug it and track its rate, and everything above the floor buys depth. Detection coverage stops improving there.

You also cover dynamic sampling that raises the rate after a deploy, during an error spike and in a seasonal surge; why the rate at which you store traces and the rate at which you run automated evaluators on them should be set separately; and tiered retention that keeps full traces for seven days and summaries after that.

The practice is a Sampling Strategy Specification an infrastructure team could implement: three risk-triggered sampling rules with their costs, added up for a month and checked against a budget you state. The extended version adds stratified sampling, keeping more of the risky traffic such as multi-turn sessions and executive users.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→