AI Evals for Product DevelopmentL2.5 · 01

Instrumentation
by system type

What an LLM app, a RAG system and an agent each need you to log
Where we left offL2.5 · 02

In 2.3 you split the trace into spans. Which fields told SQL generation apart from execution?

Tool selectionWas the right tool chosen?
Argument generationWas the SQL itself malformed?
ExecutionDid the database reject it?
Write down the fields from memory before we go on.
The problemL2.5 · 03

The v1 span tells you SQL generation failed. It doesn't tell you why.

stage"sql_generation"
status"failed"
tool_selection_alternativesnot logged
schema_validationnot logged
execution_error_typenot logged
Your PM: is it the wrong tool, bad SQL, or the database rejecting the query?
Three system typesL2.5 · 04

Each type of system breaks in its own places, so each needs its own fields.

LLM appprompt → responseBreaks on output format, token limits and config changes.
RAG systemretrieval → generationBreaks when nothing relevant comes back, or the wrong thing gets ranked first.
Agentreasoning → tools → stateBreaks on the wrong tool, bad arguments, execution errors and skipped steps.
LLM appsL2.5 · 05

An LLM app needs the config that produced the response, and why the response stopped.

Prompt
Prompt template and its version
Model config
Model, temperature, max tokens, stop sequences
Output check
Did the response pass schema validation?
Tokens
Input, output and total token counts
Stop reason
finish_reason: completed normally, or hit the limit?
RAG systemsL2.5 · 06

A RAG system needs every step between the question and the context the model saw.

Retrieval
retrieval_query, transformed_query if the query gets rewritten
Candidates
Candidate document IDs, retrieval scores, rank order
Reranking
Reranking scores and which ranker was used
Context
Chunks passed to the generator, and citations in the output
Latency
Retrieval time and generation time, logged separately
AgentsL2.5 · 07

A tool call has four stages, and each one fails differently.

Selectiontools_considered
tool_selected
selection_rationale
Argument generationtool_params_generated
schema_validation_result
Executiontool_call_success
execution_latency
tool_output
Output handlinghow_output_used
next_state
Agents also move between states. Log each transition, so you can see a step that got skipped.
Combined systemsL2.5 · 08

The AI Data Analyst uses all three patterns.

Query parsingLLM app
Context retrievalRAG
SQL generationAgent
SQL executionAgent
Chart renderingLLM app
NarrativeLLM app
Each stage needs the fields for its own pattern.
Predict before the demoL2.5 · 09

Which stage is missing the most instrumentation in v1?

Query parsing
Retrieval
SQL generation
SQL execution
Charts
Narrative
Pick one stage and name the fields you think it's missing, using the pattern that stage follows.
The demo: checking v1 against the patternL2.5 · 10

SQL generation, checked stage by stage against the agent pattern.

Selection
Is the choice logged? Are the alternatives?
Argument generation
Is the SQL logged? Is a validation result?
Execution
Are success and latency logged?
Output handling
Is the next state logged? Is how the output was used?
The demo: state transitionsL2.5 · 11

One path in v1 jumps over a stage: 124 traces skip the chart.

sql_executed→chart_rendered
sql_executed→narrative_written
sql_generated→(stops)
The skip is a bug or a shortcut. The trace doesn't say which.
From finding to hypothesisL2.5 · 12

The skipped chart gets hypotheses you can check.

Intended shortcut
Hypothesis: the question didn't need a chart, so the step was skipped on purpose.
Silent failure
Hypothesis: the chart step failed without an error and the narrative went out anyway.
Logging bug
Hypothesis: the chart ran but the stage list didn't record it. 52 of the 124 still log a chart.
The field that would settle it goes into your delta spec.
PracticeL2.5 · 13

Build a Delta Instrumentation Spec.

Base version, everyone
Classify the six stages by system type. From the SQL generation check on slide 10, list which evaluation questions you can answer today and which are blocked. Take the transition counts from slide 11, pick three hotspots, write a hypothesis for each, and write the spec.
Going further
Check the v1 retrieval span from Lesson 2.3 against the RAG pattern on slide 6. Write a strict and a lenient definition of failure and say how the hotspots change. Propose a span schema change for one high-priority missing field.
Common mistakesL2.5 · 14

Two ways system-type instrumentation goes wrong.

Mistake
What happens
Do this instead
Over-instrumenting
Forty-plus fields, most never queried. Storage grows, queries slow down, and the logging format drifts without anyone noticing.
Log the fields your system type needs for the evaluations you plan to run.
Under-instrumenting
Five fields. Evaluations stay blocked and every failure is debugged by hand.
Check each stage against its pattern before launch.
If you can't name the evaluation question a field enables, don't log it yet.
Knowledge checkL2.5 · 15

Your PM asks if retrieval is good enough to release. Can you answer?

In the traces
retrieval_query
context_provided
Not in the traces
candidate_documents
ranking_scores
Write your answer and name the evaluation that's blocked before you read on.
Next lessonL2.5 · 16

Match the fields to the architecture. Next, keep them affordable at scale.

LLM appPrompt template version, model config, schema validation, token counts, finish_reason.
RAGRetrieval query, candidates with scores and rank, reranking, context provided, citations.
AgentSelection, argument generation, execution, output handling, and state transitions.
AI ANALYST LAB · aianalystlab.ai