Week 2: Instrumentation and Reliability Engineering · Lesson 2.4

The regression safety net: evaluation in CI/CD

How do we prevent regressions before a deploy when outputs are non-deterministic?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome back. So far this week you've built the logging contract, sized how much evidence a decision needs, and designed traces with a span for every stage. All of that tells you what happened after the fact. Today we use it before the fact. The idea is a regression suite: a set of test cases that runs automatically on every change and stops the deploy when something that used to work breaks. The hard part with AI systems is that the same input doesn't always give the same output, so a suite that works fine for normal software can go red for no reason. By the end of this lesson you'll have a way to tell real breakage from run-to-run noise, and a 20-case suite with gating rules you could hand to your engineers. Let's start with the fields the suite reads.

About this lesson

A prompt change can pass code review, go out, and break something the team fixed months ago. The traces show the problem right away, but by then users have hit it. A regression suite runs a set of test cases on every change and stops the deploy before that happens.

The harder part with AI systems is that the same input can give a different output, so a suite can pass on one run and fail on the next with no code change. This lesson is about telling a real regression from that run-to-run variation.

You build the suite from your Lesson 1.3 failure taxonomy: one real trace for each category that needs ongoing measurement, plus edge cases, adversarial inputs and queries with a known correct answer. Then you sort metrics into three tiers. Blocking metrics, like SQL correctness checked against a known answer, stop the deploy on any failure. Optimization metrics, like narrative quality scored by a judge, block only when they regress by more than normal variation. Informational metrics are logged and never block. Noisy checks run several times and are tracked with pass@3 and reliable@3.

In a worked example you rerun a suite with no code change. The known-answer SQL checks give the same result both times. The judge checks move. That movement is why the tolerance band for a noisy metric has to come from your own no-change reruns, and why a case that always passes is still worth keeping.

The practice is a 20-case regression suite: at least five failure categories and five known-answer queries, each metric classified, a tolerance band for the noisy checks set from the worked rerun example, gating rules in plain English tested on five pull requests you describe, and a policy for adding and retiring cases. The extended version writes the gating rules as pseudocode and estimates how many cases you would need to detect a 5-point regression.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→