Week 2: Instrumentation and Reliability Engineering · Lesson 2.1

What to log and why it matters for AI evaluation

Which fields does each trace need before we can evaluate the feature and tell which change caused a regression?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome to Week 2. At the end of last week you built a benchmark table for three model configurations, and you ran into a problem while you did it. No trace had token counts or cost. Most didn't say which model answered. Latency was there, and not much else you needed. So the table sat on incomplete evidence, and there wasn't much you could do about it, because the data was never captured. This week is about fixing that at the source. Sooner or later engineering asks you a question: what do you need us to log? By the end of the lesson you'll have an answer they can build from.

About this lesson

When a v0 system fails, the log usually holds the question, the final output and a latency number. A retrieval error, a bad SQL query and a made-up number in the summary all look the same in that record, so you can’t tell where it broke and you can’t measure how often it happens.

This lesson answers the question engineering will ask you: what do you need us to log? The answer is an instrumentation contract between product, engineering and evaluation. Ask for too little and evaluation isn’t possible. Ask for everything and you pay for data you never use.

The contract has four categories. Correlation identifiers (trace_id, session_id, user_id) link each trace to a user and a conversation. Version metadata records which model, prompt and retrieval config produced each output, so you can tell which change caused a regression. Pipeline observability logs input, output, latency and status for each of the six stages. Join keys connect what the AI did to what the user did next, like retrying or finishing the task.

You compare a v0 trace and a v1 trace for the same failed request. In v1, two fields, the generated SQL and the database error, show that SQL generation joined on a column that doesn’t exist. You also see why “we can add fields later” doesn’t work: fields added next month don’t exist for the requests that already happened.

The practice is a Minimum Evaluability Logging Spec. For each failure type in your Lesson 1.3 taxonomy, list the fields that make it diagnosable, justify each field in one sentence, mark it minimum viable or expansion, and check the spec against three scenarios. The extended version writes the rules a checker would apply to every incoming trace.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→