Free email course

AI evals by email

30 short emails, one a day, following the same six weeks as the free AI Evals for Product Development course. Each email gives you one thing to build, from an evaluation brief on day 1 to an evaluation cadence on day 30.

By Shane Butler. Updated Sep 26, 2026.

The lessons, the weekly quizzes and the certificate are on the course page. The emails are the short version, one a day in your inbox.

The emails

Week 1

Foundations & Economics

  1. Day 0 Welcome to the AI Evals for Product Development Email Course
  2. Day 1 Your Evaluation Brief: pick the feature and the decision
  3. Day 2 The Evaluation Surface Map: what are you measuring end to end
  4. Day 3 Failure Taxonomy: turn messy traces into categories you can act on
  5. Day 4 Distribution Thinking: why averages hide risk
  6. Day 5 Cost, Latency, Quality: pick your trade-offs on purpose
Week 2

Instrumentation & Reliability

  1. Day 6 The Minimum Logging Spec: what to capture so evals are possible
  2. Day 7 How Little Instrumentation Is Enough?
  3. Day 8 Trace Design: make behavior reproducible
  4. Day 9 Regression Suite and CI Gates: stop regressions before deploy
  5. Day 10 LLM App vs RAG vs Agent: what you must capture to diagnose failures
Week 3

Measurement

  1. Day 11 Define User Value: what 'good' means outside the model
  2. Day 12 Ground Truth and Coverage: where your eval cases come from
  3. Day 13 Signal Catalog: gates vs diagnostics vs drivers
  4. Day 14 Similarity and Retrieval Metrics: isolate where failure lives
  5. Day 15 Rubrics and Judges: measure quality when similarity falls short
Week 4

Metric Design

  1. Day 16 Metric Strategy: build a suite that supports decisions
  2. Day 17 Metric Patterns: pick the archetype that matches the workflow
  3. Day 18 Metric Validation: prove your metric is decision-ready
  4. Day 19 Segmentation: find where it works and where it breaks
  5. Day 20 Driver Analysis: explain variance and pick interventions
Week 5

Pipelines & Experiments

  1. Day 21 Evaluation Pipelines: make eval repeatable and queryable
  2. Day 22 Dataset Lifecycle: iteration sets vs test sets vs holdouts
  3. Day 23 Experiments for Stochastic Systems: design for noise and tails
  4. Day 24 Rollout Gates: what must be true before exposure
  5. Day 25 Monitoring: catch silent failures without evaluating everything
Week 6

Decisions & Operations

  1. Day 26 Ship, Ramp, Hold, Roll Back: what the evidence justifies
  2. Day 27 Findings to Actions: what to change next and how you'll know
  3. Day 28 Prioritization: what to fix first
  4. Day 29 Ownership: who owns quality, metrics, and the ship decision
  5. Day 30 Evaluation Cadence: how this becomes a real operating system

Get the six weeks one email at a time

Free. A welcome email arrives right after you sign up.