AI Evals for Product DevelopmentL2.3 · 01

Trace design and
reproducibility

What each request has to record so you can explain it, measure it and replay it
Where we left offL2.3 · 02

v0 logged six fields. Which field did v1 add that let you see SQL fail?

request_id
user_query
final_answer
latency_ms
success
timestamp
Answer from memory before you advance.
The scenarioL2.3 · 03

A user says conversion is 4.2%. The dashboard says 3.8%. The trace says success.

What the PM asks
"User reports a conversion rate of 4.2%, but the BI dashboard says 3.8%. What went wrong?"
What the v0 trace shows
success: true
latency_ms: 2300
No SQL, no retrieved context, nothing per stage.
DefinitionL2.3 · 04

A trace is the record of one request's path through the system.

User queryWhat was asked
RetrievalDefinitions and context
SQLGenerate and run
ChartRender
NarrativeWrite the answer
Final answerWhat the user saw
A trace schema is the list of fields you capture at each step.
SpansL2.3 · 05

Each stage gets its own section of the trace, called a span.

Retrievalretrieval_querynum_docs_retrievedtop_doc_scoreretrieved_doc_idsretrieval_latency_ms
SQLgenerated_sqlsql_successsql_errorsql_latency_ms
Chartchart_typechart_successchart_error
Narrativenarrative_textmodel_usedprompt_tokens
Which span carries the most fields, and why would that be?
Storage and analysisL2.3 · 06

Store traces nested. Flatten them into rows to analyze.

As logged
{ "request_id": "req_0127",
  "spans": {
    "retrieval": {...},
    "sql": {...} } }
Flattened
One row per request, one column per span field. Filter, group and count as you would any table.
CompletenessL2.3 · 07

Completeness is the share of schema fields that actually have data.

completeness = fields populated ÷ fields in the schema
v0 trace
Six fields populated out of the 20 the v1 schema defines.
v1 trace
19 of 20 populated. The empty one is an error field.
Predict before the demoL2.3 · 08

A user got the wrong DAU definition. Which schema lets you prove retrieval caused it?

Option A
num_docs_retrieved
retrieval_latency_ms
Option B
num_docs_retrieved
retrieval_latency_ms
retrieval_query
top_doc_score
retrieved_doc_ids
Write your pick and one reason before we advance.
The demo: a v0 traceL2.3 · 09

The v0 trace for a mobile DAU question.

request_idreq_0127
user_query"What's our DAU trend for mobile app last 30 days?"
final_answer"Your mobile app DAU averaged 47,320..."
latency_ms2300
successtrue
timestamp2025-01-15T14:23:01Z
What can you check about this answer from these six fields?
The demo: the same request in v1L2.3 · 10

The v1 trace records each stage separately.

Retrieval span
retrieval_query"mobile app DAU definition and calculation"
retrieved_doc_idsfive document IDs, values elided here
top_doc_score0.84
SQL span
generated_sqlSELECT date, COUNT(DISTINCT user_id) FROM events WHERE platform = 'mobile' ...
sql_successtrue
Chart and narrative spans
chart_type · model_usedline · GPT-4
What becomes measurableL2.3 · 11

For every field you add, name what it lets you measure.

Field
v0
v1
What becomes measurable
retrieval_query
No
Yes
Whether the question gets turned into a useful search
retrieved_doc_ids
No
Yes
Whether the right documents come back, and which never do
generated_sql
No
Yes
SQL correctness against what the user asked
sql_success
No
Yes
SQL error rate, by query type, user group and time
ReproducibilityL2.3 · 12

To replay a request, you need the versions and what each span took in.

From 2.1
model_version, prompt_template_version, retrieval_config_version: which system produced the answer.
From today
Each span's input and output: what the system saw and did at every step.
Missing either one, you can see the failure but can't run it again.
PracticeL2.3 · 13

Build an Instrumentation Readiness Report.

Base version, everyone
Put the v0 trace on slide 9 next to the v1 trace on slide 10 and write the schema comparison. Use slide 7 for completeness. Fill in what each SQL span field lets you measure. Design the chart rendering span from scratch. Assemble the report with a justification for every field.
Going further
Write the rules a trace checker would apply to your schema. Say which fields may be blank and when, and what share of traces would have to match before you trust the data.
Common mistakesL2.3 · 14

Three ways a trace schema goes wrong.

Mistake
What happens
Do this instead
Log everything
Storage grows, queries slow down, and personal data ends up inside logged prompts.
Log a field only if you can name the metric that needs it.
No schema version
Old traces lack new fields. Trends across the change break.
Put schema_version on every trace. Leave missing fields empty, never a fake number.
Deploy it untested
A typo in a field name silently loses days of data.
Run test traces in staging and check they match the schema first.
Knowledge checkL2.3 · 15

Three judgment calls. Write your answer before you read on.

01A trace has generated_sql filled in but sql_success is empty. What can you infer, and what can't you measure for that request?
02A teammate wants to log the full 2,000-character prompt on every trace. You're at 10,000 requests a day. How do you decide?
03The v0 trace shows latency_ms 2300 and nothing per stage. The PM says "it's too slow, fix it." What's the problem?
Next lessonL2.3 · 16

A trace schema you can diagnose and replay from. Next, a test suite that runs before every deploy.

v0
Six fields. Tells you the system ran.
v1
Four spans. Tells you what each stage did, how long it took, and lets you replay it.
AI ANALYST LAB · aianalystlab.ai