Week 2: Instrumentation and Reliability Engineering · Lesson 2.3

Trace design and reproducibility

If we need to explain or reproduce a behavior, what must be true about our trace design?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome back. In 2.1 you wrote the logging contract, the list of fields engineering agrees to capture. In 2.2 you worked out how many examples you need before a decision holds up. Today we get specific about the shape of the record itself. When a request comes through the AI Data Analyst, what gets written down, stage by stage, so that later you can say what happened, measure how often it happens, and run the same request again to check a fix? That record is called a trace, and the design of it decides which questions you can answer three months from now. Let's start with one field from 2.1 that changed a lot.

About this lesson

A PM asks why a user got a conversion rate of 4.2% when the dashboard says 3.8%. You open the trace and it says success true, latency 2,300 milliseconds. Was the SQL wrong, did retrieval pull the wrong metric definition, or did the summary misread the number? A trace that only records that the system ran can’t tell you.

A trace is the record of one request’s path through the system. It is split into spans, one per stage: retrieval, SQL, chart and narrative. Each span holds the fields that stage needs, because each stage fails in its own way. Retrieval gets the most fields, since a retrieval failure doesn’t throw an error. It just produces a confident wrong answer.

Traces are stored as nested JSON and flattened into one row per request for analysis, using spans_to_row. From there you can measure completeness, the share of schema fields that actually have data. Against a 20-field schema, a v0 trace fills 6 and a v1 trace fills 19.

You predict which of two retrieval schemas lets you prove retrieval caused a wrong DAU definition, then compare a v0 and a v1 trace for the same request. You also see what it takes to replay a request after a fix: the version fields from 2.1 plus what each span took in and put out.

The lesson closes on three mistakes: logging everything, changing the schema without a schema_version field, and deploying logging nobody tested.

The practice is an Instrumentation Readiness Report built from the v0 and v1 traces in the lesson: completeness for each, a schema comparison table where every field names the measurement it enables, and a chart rendering span you design yourself. The extended version writes the rules a checker would apply to traces against your schema.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→