AI Evals for Product Development
How to test an AI feature: find where it breaks, measure how often, and decide when it is good enough to release. 36 lessons over six weeks, a short quiz each week, and a certificate when you pass the final exam.
By Shane Butler. Free, self-paced, no account needed to read.
You cannot test an AI feature the way you test normal software. Ask it the same question twice and you can get two different answers, and a wrong answer looks just like a right one. An eval is a test you can rerun every time you change the prompt, the model or the data, and see whether things got better or worse.
The example running through the course is an AI data analyst: someone asks a question in plain English, and it writes the SQL, runs it and explains the result. That makes the failures concrete, like a number that changes between runs or a query on the wrong definition.
The 36 lessons
Foundations and Economics
Why an AI feature can pass every test you have and still give users a different answer each time, and how to weigh quality against cost and speed.
Week 1 quiz, 5 questionsInstrumentation and Reliability Engineering
What to log so you can replay a bad answer later, and how to catch a change that makes answers worse before users notice.
Week 2 quiz, 5 questionsRigorous Measurement of Output Success and Failure
How to score answers: against a known right answer when there is one, and with a rubric, a person or a model judge when there is not.
Week 3 quiz, 5 questions- 3.1 Grounding evaluation in user value
- 3.2 Ground truth sources, regression suites, and synthetic data
- 3.3 Deriving evaluation signals from available ground truth
- 3.4 Retrieval metrics: precision, recall, MRR and NDCG
- 3.5 Semantic metrics with human and model judges
- 3.6 Correcting for an imperfect judge
Metric Design and Business Outcome Linkage
Deciding which metrics can hold a release, spending a fixed evaluation budget where it matters, finding where quality drops and why, and writing release criteria.
Week 4 quiz, 5 questions- 4.1 Metric strategy: blocking metrics vs optimization metrics
- 4.2 Metric design patterns for AI features
- 4.3 Cost-aware evaluation on a fixed budget
- 4.4 Segmentation strategy for AI systems
- 4.5 Driver analysis: explaining variance and choosing what to change first
- 4.6 Metric specifications, thresholds, baselines, and release criteria
Pipelines, Experiments, and Continuous Validation
Running evals on every change, keeping the test set current, A/B testing a feature whose output varies, and watching it after launch.
Week 5 quiz, 5 questions- 5.1 Evaluation pipeline architecture and environments
- 5.2 Test set strategy and dataset lifecycle
- 5.3 Experiment design for stochastic systems
- 5.4 Launch readiness and rollout gates
- 5.5 Monitoring for drift and regressions
- 5.6 Building evaluation automation end-to-end
- 5.7 Capstone lab: run the pipeline by hand
Decision-Making and Organization
Deciding whether to release, hold or roll back, who owns the evals, and how to explain the results to leadership.
Week 6 quiz, 5 questionsThe certificate
Pass the final exam, 16 of 20 questions across all six weeks, and your certificate is made on the page with your name, score and date. You can retake the exam.
Questions
- Is it free?
- Yes. All 36 lessons, the six quizzes and the final exam are free, with no paid tier and no deadline.
- Do I get a certificate?
- Each weekly quiz is five questions and gives you a completion certificate when you pass. The final exam is 20 questions across all six weeks; score 16 or more and you get the course certificate. Both are made on the page with your name and date, and download as an image.
- What is in a lesson?
- A slide deck with speaker notes under each slide, so you can read it like a chapter. Each lesson ends with a practice you can do from the slides, with no code or data files to download.
- How long does it take?
- Five to seven lessons a week for six weeks if you take it a week at a time. It is self-paced, so you can go faster.
- Who made it?
- Shane Butler, co-founder of AI Analyst Lab, from ten years of product data science at Stripe, Nextdoor, AppFolio, Ontra and PwC.
Prefer email? AI evals by email sends the same six weeks as 30 short emails.