The regression safety net: evaluation in CI/CD
How do we prevent regressions before a deploy when outputs are non-deterministic?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
A prompt change can pass code review, go out, and break something the team fixed months ago. The traces show the problem right away, but by then users have hit it. A regression suite runs a set of test cases on every change and stops the deploy before that happens.
The harder part with AI systems is that the same input can give a different output, so a suite can pass on one run and fail on the next with no code change. This lesson is about telling a real regression from that run-to-run variation.
You build the suite from your Lesson 1.3 failure taxonomy: one real trace for each category that needs ongoing measurement, plus edge cases, adversarial inputs and queries with a known correct answer. Then you sort metrics into three tiers. Blocking metrics, like SQL correctness checked against a known answer, stop the deploy on any failure. Optimization metrics, like narrative quality scored by a judge, block only when they regress by more than normal variation. Informational metrics are logged and never block. Noisy checks run several times and are tracked with pass@3 and reliable@3.
In a worked example you rerun a suite with no code change. The known-answer SQL checks give the same result both times. The judge checks move. That movement is why the tolerance band for a noisy metric has to come from your own no-change reruns, and why a case that always passes is still worth keeping.
The practice is a 20-case regression suite: at least five failure categories and five known-answer queries, each metric classified, a tolerance band for the noisy checks set from the worked rerun example, gating rules in plain English tested on five pull requests you describe, and a policy for adding and retiring cases. The extended version writes the gating rules as pseudocode and estimates how many cases you would need to detect a 5-point regression.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→