What to log and why it matters for AI evaluation
Which fields does each trace need before we can evaluate the feature and tell which change caused a regression?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
When a v0 system fails, the log usually holds the question, the final output and a latency number. A retrieval error, a bad SQL query and a made-up number in the summary all look the same in that record, so you can’t tell where it broke and you can’t measure how often it happens.
This lesson answers the question engineering will ask you: what do you need us to log? The answer is an instrumentation contract between product, engineering and evaluation. Ask for too little and evaluation isn’t possible. Ask for everything and you pay for data you never use.
The contract has four categories. Correlation identifiers (trace_id, session_id, user_id) link each trace to a user and a conversation. Version metadata records which model, prompt and retrieval config produced each output, so you can tell which change caused a regression. Pipeline observability logs input, output, latency and status for each of the six stages. Join keys connect what the AI did to what the user did next, like retrying or finishing the task.
You compare a v0 trace and a v1 trace for the same failed request. In v1, two fields, the generated SQL and the database error, show that SQL generation joined on a column that doesn’t exist. You also see why “we can add fields later” doesn’t work: fields added next month don’t exist for the requests that already happened.
The practice is a Minimum Evaluability Logging Spec. For each failure type in your Lesson 1.3 taxonomy, list the fields that make it diagnosable, justify each field in one sentence, mark it minimum viable or expansion, and check the spec against three scenarios. The extended version writes the rules a checker would apply to every incoming trace.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→