Trace design and reproducibility
If we need to explain or reproduce a behavior, what must be true about our trace design?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
A PM asks why a user got a conversion rate of 4.2% when the dashboard says 3.8%. You open the trace and it says success true, latency 2,300 milliseconds. Was the SQL wrong, did retrieval pull the wrong metric definition, or did the summary misread the number? A trace that only records that the system ran can’t tell you.
A trace is the record of one request’s path through the system. It is split into spans, one per stage: retrieval, SQL, chart and narrative. Each span holds the fields that stage needs, because each stage fails in its own way. Retrieval gets the most fields, since a retrieval failure doesn’t throw an error. It just produces a confident wrong answer.
Traces are stored as nested JSON and flattened into one row per request for analysis, using spans_to_row. From there you can measure completeness, the share of schema fields that actually have data. Against a 20-field schema, a v0 trace fills 6 and a v1 trace fills 19.
You predict which of two retrieval schemas lets you prove retrieval caused a wrong DAU definition, then compare a v0 and a v1 trace for the same request. You also see what it takes to replay a request after a fix: the version fields from 2.1 plus what each span took in and put out.
The lesson closes on three mistakes: logging everything, changing the schema without a schema_version field, and deploying logging nobody tested.
The practice is an Instrumentation Readiness Report built from the v0 and v1 traces in the lesson: completeness for each, a schema comparison table where every field names the measurement it enables, and a chart rendering span you design yourself. The extended version writes the rules a checker would apply to traces against your schema.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→