How instrumentation requirements differ by system type
What must be captured for this system type so we can diagnose the failures that matter?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
A v1 span can tell you that SQL generation failed. It can’t tell you whether the system picked the wrong tool, wrote bad SQL, or had its query rejected by the database. Those three causes need different fixes, and they all show up as the same status field.
What you need to log depends on the kind of system. An LLM app needs the prompt template version, the model config, a schema validation result, token counts and finish_reason, so you can tell a truncated answer from a wrong one. A RAG system needs every step between the question and the context the model saw: the search query, the candidate documents with their scores and rank, any reranking, the chunks passed to the generator and the citations. An agent needs each stage of a tool call logged separately (which tool it chose, the arguments it wrote, whether the call ran, and what it did with the result) plus the transitions between states.
The AI Data Analyst uses all three patterns. Query parsing, charts and the narrative are LLM apps, context retrieval is RAG, and SQL generation and execution are tool calls. You predict which stage has the biggest gap in v1, then check it. In SQL generation only execution is fully logged. The decisions before it, which tool and which arguments, are missing.
You also pull the state transitions out of the traces and look for paths that were never designed. In v1, 124 traces go from running the SQL straight to the narrative with no chart step. You write hypotheses for that path and name the field that would settle it.
The practice is a Delta Instrumentation Spec built from the lesson’s v1 findings: each stage classified by system type, the evaluation questions you can and can’t answer today, three hotspots from the transition counts, and a table of missing fields with the evaluation each one unblocks.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→