Week 2: Instrumentation and Reliability Engineering · Lesson 2.5

How instrumentation requirements differ by system type

What must be captured for this system type so we can diagnose the failures that matter?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome back. By now your v1 traces log every pipeline stage: the input, the output, the latency and whether it succeeded. That gets you a long way. But an LLM app, a retrieval system and an agent break in different places, and a generic span with a status field can't tell you why each one broke. Today we work out what each type of system needs you to log on top of the basics, and then we check the AI Data Analyst against it, stage by stage. Let's start with something you already built.

About this lesson

A v1 span can tell you that SQL generation failed. It can’t tell you whether the system picked the wrong tool, wrote bad SQL, or had its query rejected by the database. Those three causes need different fixes, and they all show up as the same status field.

What you need to log depends on the kind of system. An LLM app needs the prompt template version, the model config, a schema validation result, token counts and finish_reason, so you can tell a truncated answer from a wrong one. A RAG system needs every step between the question and the context the model saw: the search query, the candidate documents with their scores and rank, any reranking, the chunks passed to the generator and the citations. An agent needs each stage of a tool call logged separately (which tool it chose, the arguments it wrote, whether the call ran, and what it did with the result) plus the transitions between states.

The AI Data Analyst uses all three patterns. Query parsing, charts and the narrative are LLM apps, context retrieval is RAG, and SQL generation and execution are tool calls. You predict which stage has the biggest gap in v1, then check it. In SQL generation only execution is fully logged. The decisions before it, which tool and which arguments, are missing.

You also pull the state transitions out of the traces and look for paths that were never designed. In v1, 124 traces go from running the SQL straight to the narrative with no chart step. You write hypotheses for that path and name the field that would settle it.

The practice is a Delta Instrumentation Spec built from the lesson’s v1 findings: each stage classified by system type, the evaluation questions you can and can’t answer today, three hotspots from the transition counts, and a table of missing fields with the evaluation each one unblocks.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→