Deriving evaluation signals from available ground truth
Given imperfect ground truth, what signals can we compute and what are they for?
Browse lessons
Week 1: Foundations and Economics
Week 2: Instrumentation and Reliability Engineering
Week 3: Rigorous Measurement of Output Success and Failure
Week 4: Metric Design and Business Outcome Linkage
- 4.1Metric strategy: blocking metrics vs optimization metrics
- 4.2Metric design patterns for AI features
- 4.3Cost-aware evaluation on a fixed budget
- 4.4Segmentation strategy for AI systems
- 4.5Driver analysis: explaining variance and choosing what to change first
- 4.6Metric specifications, thresholds, baselines, and release criteria
Week 5: Pipelines, Experiments, and Continuous Validation
- 5.1Evaluation pipeline architecture and environments
- 5.2Test set strategy and dataset lifecycle
- 5.3Experiment design for stochastic systems
- 5.4Launch readiness and rollout gates
- 5.5Monitoring for drift and regressions
- 5.6Building evaluation automation end-to-end
- 5.7Capstone lab: run the pipeline by hand
Week 6: Decision-Making and Organization
Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.
By Week 3 the AI Data Analyst logs something at every stage: retrieval scores, the SQL it wrote, execution results, the summary text, latency and token counts. The PM asks whether v1 is ready to ship, and a dashboard full of numbers still doesn’t answer that. What’s missing is a decision, made ahead of time, about which numbers you measure and what each one is for.
An evaluation signal is the number you get when you compare the system’s output to ground truth, the known right answer. Every signal has a type and a role. The type is how you compute it. Execution-based oracles run the output, for example running the generated SQL and a reference query and comparing the result sets. Structural signals check format, like whether the JSON parses or the required fields are there. Semantic signals need a judge and a rubric, and the judge brings its own variance on top of the system’s. The role is the decision the signal feeds. A gate blocks a release, a diagnostic helps you find where a failure came from, and a driver explains where quality varies. If you can’t name the decision a signal feeds, you drop it.
You take an example hit rate on 300 test queries, work out how many queries reach the SQL step with nothing relevant, and look at what a retrieval miss does to the SQL and the summary that come after it.
The practice is a signal catalog. You work out what Precision@5 and MRR would add to the hit rate, describe what over-retrieving and under-retrieving look like, and write down at least five signals with their type, what they’re checked against, their role, their cost and how they’re computed. The extended version adds a judge-based signal with its variance written down and grows the catalog to ten signals.
Go deeper with AI Analytics for Everyone
5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.
Book 1-on-1 with Shane
30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.
Finished all 36 lessons? Take the exam and get your free AI Evals certification.
→