Week 5: Pipelines, Experiments, and Continuous Validation · Lesson 5.7

Capstone lab: run the pipeline by hand

When an alert fires, can you tell from the run records whether the system got worse or the sample changed?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome to the Week 5 lab. Last lesson you saw how the pieces connect: the offline pipeline, the gate check before each rollout step, and monitoring once the feature is live. Today you run the pipeline yourself, twice, from start to finish. You'll take a sample of traces, score them, roll the scores up, write a run record, compare the numbers against the gates, and decide whether an alert should fire. The second run is set up to fire one, and most of the lab is about what you do after it fires. You don't need code for any of this. Everything is on the slides. Let's start with what you'll need.

About this lesson

This is the Week 5 lab. You run the evaluation pipeline from Lesson 5.1 yourself, twice, on real traces from the AI Data Analyst: sample, judge, aggregate, store, report and alert. You don’t need code. The two trace tables, the run record and the report table are all in the slides, and a spreadsheet or paper is enough.

Run A is a random sample of twelve traces from the full set of 1,000. Run B is a random twelve from the multi-table join questions only. You write the pass rules before you count, score each trace from the outcomes the pipeline recorded (did it reach the narrative, and if not, where did it stop), and compute completion rate, SQL success, mean latency and mean cost. Then you fill in a run record for each run and compare the numbers against three gates.

Run B fails the SQL gate. Most of the lab is about what you do next: put the two run records side by side to find what changed, look at how much one trace moves a twelve-trace number, and decide if the cause is a worse system, a weaker segment or a small unlucky sample.

The full version of the judging stage adds the oracle SQL check and the LLM judge from Week 3. Those need code and an API key, so the slide version uses recorded outcomes, and the knowledge check asks what that leaves out.

The practice is two run records, the completed report table, and one paragraph on the Run B alert. If you can get logs from an AI feature you work on, the extended version runs the same six stages on 50 of your own traces, with a second run that changes one thing.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→