AI Evals for Product DevelopmentL5.7 · 01

Capstone lab:
run the pipeline by hand

Two evaluation runs on real AI Data Analyst traces, from sampling to the alert
What you needL5.7 · 02

Everything for this lab is on these slides.

For the base lab
The two trace tables on slides 6 and 11. The run record on slide 8. The report table on slide 9. A spreadsheet, or paper and a calculator.
Not needed
No code and no API key. A full pipeline adds an oracle SQL check and an LLM judge, and those do need code and a key.
The traces are real rows from the course's v1 trace set of 1,000 AI Data Analyst runs.
The six stagesL5.7 · 03

You'll run all six stages, twice.

SampleTwelve traces, chosen at random
JudgePass or fail on each trace, from recorded outcomes
AggregateFour numbers per run
StoreA run record with its metadata
ReportEach number against its gate
AlertWhich gate fails, and what you check first
Run A samples from all traces. Run B samples only from multi-table join questions.
The rulesL5.7 · 04

Decide what counts as a pass before you look at the data.

Completion rate
Share of traces that reached the narrative. A blank in the "failed at" column means it completed.
SQL success
Share of traces that did not fail at SQL execution. Gate: 0.75 or higher.
Mean latency
Average of the latency column. Gate: 2,000 ms or lower.
Mean cost
Average of the cost column. Gate: $0.15 or lower per query.
Completion rate is tracked. The other three are blocking gates for this lab.
Predict firstL5.7 · 05

Run A is twelve random traces. Will it pass all three gates?

Write down pass or fail for SQL success, latency and cost, and which gate you expect to be closest to failing.
You know the full trace set fails at SQL execution about 15% of the time. Use that.
Run A · random sample of 12 from all tracesL5.7 · 06

Run A: count the passes.

#
Trace
Question type
Failed at
Latency ms
Cost $
1
d04e1bd2
multi_table_join
1866
0.0108
2
cad46696
comparison
1674
0.0117
3
dd08b84d
simple_lookup
1396
0.0081
4
2c4b55c2
trend_analysis
1757
0.0100
5
edf7fe79
trend_analysis
1530
0.0139
6
6b178422
comparison
1972
0.0079
7
4a4a52c1
trend_analysis
2034
0.0085
8
b8a9c83c
trend_analysis
SQL execution
1765
0.0124
9
5596806c
multi_table_join
Chart rendering
1474
0.0081
10
25680b66
comparison
1518
0.0117
11
25911040
multi_table_join
1554
0.0113
12
d4febcb2
simple_lookup
Context retrieval
1592
0.0067
Check your predictionL5.7 · 07

Which gate was closest to failing in Run A?

What you predicted
Pass or fail on each gate, and the one you expected to be tightest.
What to compare
For each gate, how far the Run A number sits from its line, in the same units as the gate.
A gate that passes by a small margin on twelve traces can fail on the next twelve.
Stage 4 · storeL5.7 · 08

Write the run record before you report anything.

run_idyour name for this run
datethe day you ran it
datasettraces_v1_full, v1.0.0
modelgpt-4o, 2024-05-13
prompt_template1.0.0
samplingstrategy and seed
sample_sizenumber of traces
judgerecorded outcomes, and your pass rules
metricsthe four numbers
notesanything you decided along the way
Stage 5 · reportL5.7 · 09

Put each number next to its gate.

Metric
Gate
Run A
Run B
SQL success
0.75 or higher
____
____
Mean latency
2,000 ms or lower
____
____
Mean cost
$0.15 or lower
____
____
Completion rate
Tracked only
____
____
Sample size
Always reported
____
____
Predict before Run BL5.7 · 10

Run B samples only multi-table join questions. What happens to SQL success?

Up, down, or about the same as Run A? Write one sentence with your reason.
Same dataset, same model, same pass rules. The only change is which traces can be sampled.
Run B · random sample of 12 from multi-table join tracesL5.7 · 11

Run B: count the passes.

#
Trace
Question type
Failed at
Latency ms
Cost $
1
7e3b534b
multi_table_join
SQL execution
1740
0.0118
2
c06da2bd
multi_table_join
1861
0.0115
3
a14f33be
multi_table_join
1895
0.0070
4
81d8f990
multi_table_join
1836
0.0092
5
8e5aaea3
multi_table_join
SQL execution
1768
0.0136
6
18f969a7
multi_table_join
SQL execution
1343
0.0093
7
95ebddba
multi_table_join
1693
0.0070
8
966a66e9
multi_table_join
1880
0.0106
9
94b2c8ba
multi_table_join
2194
0.0124
10
7ff00009
multi_table_join
1412
0.0060
11
42b8d038
multi_table_join
1705
0.0136
12
109539e0
multi_table_join
SQL execution
2072
0.0100
Stage 6 · alertL5.7 · 12

The SQL gate failed. Check the record and the sample size first.

01What changed between the runs? Put the two run records side by side and find every field that differs.
02How many traces is the number built on? Ask how far one trace moves it.
03Then decide: is this a problem with the system, a problem in one segment, or noise from a small sample?
Your deliverableL5.7 · 13

Hand in two run records, one report and one paragraph.

Base version, everyone
Run records for A and B. The completed report table. One paragraph on the Run B alert: what differs between the runs, what the sample size allows you to say, and what you would do next.
Extended version, your own logs
Pull 50 traces from an AI feature you work on. Write the pass rules first. Run all six stages, then a second run with one change, and compare the records.
Common mistakesL5.7 · 14

Three ways a first pipeline run goes wrong.

Mistake
What happens
Do this instead
Setting pass rules after counting
Edge cases like trace 12 get scored whichever way makes the number look better.
Write the rules, and each judgment call, into the run record first.
Reporting a number without its record
Next month the team can't tell whether a change came from the system or the measurement.
Store dataset, model, sampling, seed, sample size and judge with every run.
Acting on an alert from a small sample
A rollback gets triggered by four traces out of twelve.
Check what changed and the sample size, then rerun larger before deciding.
Knowledge checkL5.7 · 15

Three judgment calls. Write your answer before you read on.

01A teammate reruns Run A next week with a new judge version, and completion falls. Nothing else in the record changed. What can you conclude about the system?
02Run B's SQL gate failed. Your PM asks for one sentence on whether the system got worse. What do you say?
03Both runs used recorded outcomes as the judge. What could still be wrong in a trace that completed every stage?
Next lessonL5.7 · 16

The pipeline gives you evidence. Next, you decide when the evidence is unclear.

What you did this week
Sampled, judged, aggregated, stored, reported and alerted, on two runs of real traces, and read an alert against its run record.
Lesson 6.1
Decision-making under uncertainty: what to do when the numbers sit between roll back and ramp.
AI ANALYST LAB · aianalystlab.ai