AI Evals for Product DevelopmentL2.1 · 01

What to log

The fields your traces need before evaluation is possible
Where we left offL2.1 · 02

Which failure from your taxonomy was hardest to diagnose from the input and the output alone?

What the v0 traces gave you
The user's question, the final output, a latency number, a success flag, and the generated SQL on some traces (207 of 500).
What you needed
Write down the one piece of intermediate data that would have made that failure obvious.
The problemL2.1 · 03

With only the input and output logged, every failure looks the same.

user_query: "Show revenue by region"
final_output: "error"
latency_ms: 1847
Parsing?
Retrieval?
SQL gen?
Execution?
Chart?
Narrative?
The contractL2.1 · 04

"What do you need us to log?" starts a negotiation between three groups.

ProductNeeds to diagnose what users complain about.
EngineeringWants to keep storage cost and build work down.
EvaluationNeeds specific fields to compute quality metrics.
Log too little
Evaluation isn't possible. The data you need doesn't exist.
Log too much
Storage cost grows, queries slow down, and it gets hard to tell which fields matter.
The frameworkL2.1 · 05

The Evaluability Logging Contract has four required categories.

Correlation identifiersLink each trace to its user, its session and what happened next.
Version metadataTie each output to the model, prompt and retrieval config that produced it.
Pipeline observabilityShow what happened at each stage, so you can see where it broke.
Join keysConnect what the AI did to what the user did afterward.
If a category is missing, a whole class of evaluation goes with it.
Category 1L2.1 · 06
trace_idOne per request. The ID everything else hangs off.
session_idShared across a multi-turn conversation.
user_idLets you segment by user and join to what they did.
Category 2L2.1 · 07

Version metadata tells you which change caused a regression.

Field
Before the jump
After the jump
model_id
gpt-4
gpt-4
prompt_template_version
v2.3
v2.4
retrieval_config_version
v1.1
v1.1
Error rate goes from 2% to 12% overnight. Without these fields, you can't say which change did it.
Category 3L2.1 · 08

Pipeline observability: four fields for each of the six stages.

Parsing
Retrieval
SQL gen
Execution
Chart
Narrative
input
output
latency
status
With these, the trace shows which stage broke.
Category 4L2.1 · 09

Join keys connect what the AI did to what the user did next.

From the trace
trace_id, session_id and user_id, carried into your product and business data.
What it lets you ask
Did the user retry right after this answer? Did they finish the task they came to do?
This is how you check that a quality metric reflects what users actually experience.
Predict before the demoL2.1 · 10

Engineering proposes two schemas. Which one supports each metric?

Schema Arequest_id
user_query
final_answer
latency_ms
success
timestamp
Schema Brequest_id, user_query, final_answer,
latency_ms, session_id, retrieval_query,
num_docs_retrieved, generated_sql,
sql_success, model_version,
prompt_template_version
1. Which supports a SQL correctness metric? 2. Which supports breaking results down by user role? Write both answers down.
The demo: v0L2.1 · 11

A failed request, sketched the v0 way.

request_id: abc123
user_query: "Show revenue by region"
final_answer: "error"
latency_ms: 1847
success: false
timestamp: 2025-01-15T14:23:01Z
A simplified sketch made for this slide. The real v0 file has 37 fields per trace, and 19 are empty on every trace. A user just reported this one. What went wrong?
The demo: v1L2.1 · 12

The same request, sketched with the full contract logged.

trace_id: abc123   session_id: sess_789   user_id: u_456
model_version: gpt-4   prompt_template_version: v2.4
retrieval_query: "revenue region sales data"   num_docs_retrieved: 3
generated_sql: SELECT revenue FROM sales s INNER JOIN regions ON s.region = regions.id
sql_error: "column regions.id does not exist"
sql_success: false
Which stage failed, and which field told you?
PracticeL2.1 · 13

Build a Minimum Evaluability Logging Spec engineering can implement.

Base version, everyone
For each failure type in your 1.3 taxonomy, list the fields that make it diagnosable. Fill all four categories. Justify each field in one sentence. Mark each one minimum viable or expansion. Check the spec against three scenarios.
Going further
Write the validation rules a trace checker would apply: for each required field, what counts as missing, and what happens when it is.
Common mistakesL2.1 · 14

Three ways a logging spec goes wrong.

Mistake
What happens
Do this instead
Log the endpoints only
Every failure looks like "output doesn't match expected." You can't tell retrieval from SQL from narrative.
Log input, output, latency and status for every stage.
Log everything
Storage cost climbs, queries slow down, and the useful fields get buried.
Keep a field only if you can name the metric it enables.
Skip versions
A regression shows up and you can't say which change caused it.
Log model, prompt and retrieval config versions on every trace.
Knowledge checkL2.1 · 15

Engineering says "we can add fields later." Write your answer before you read on.

01The spec logs trace_id, user_query and final_output. Which Week 3 capabilities does this block?
02Why doesn't adding the missing fields next month solve the problem?
Next lessonL2.1 · 16

Four categories, one contract. Next: how much evidence is enough to decide.

Correlation identifierstrace_id, session_id, user_id. Multi-turn linking, segmentation, outcome joins.
Version metadatamodel, prompt and retrieval config versions. Regression attribution and reproducibility.
Pipeline observabilitySix stages, four fields each. Failure diagnosis by stage.
Join keysTrace IDs carried into user behavior and business outcomes.
AI ANALYST LAB · aianalystlab.ai