Week 4: Metric Design and Business Outcome Linkage · Lesson 4.1

Metric strategy: blocking metrics vs optimization metrics

Which metrics should be able to hold a release on their own, and why?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 15

Speaker notes

Welcome to Week 4. For three weeks you've been building individual metrics for the AI Data Analyst. Retrieval recall, SQL correctness, judges for the narrative, and last lesson a Judge Report Card that tells you how far to trust those judges. Each of those measures something real. But a pile of good metrics doesn't answer the question your PM actually brings to the room, which is whether v2 goes out. More numbers can make that meeting longer, because now everyone has a number to point at. So this week is about turning measurements into decisions. Today's piece is the first step: deciding which metrics get to stop a release and which ones you track and keep improving. That one choice changes how you set thresholds, how often you watch each metric, and who owns it. Let's start with the judges you just calibrated.

About this lesson

By Week 4 the AI Data Analyst has twelve evaluation metrics, and your PM still can’t get an answer to whether v2 should go out. Twelve good numbers don’t make a decision on their own, because the team hasn’t said which ones matter for a release.

This lesson gives every metric one of two jobs. A blocking metric has a threshold, and if it misses, the release waits. An optimization metric has a target, and missing it starts work without holding the release. To sort them you ask one question of each metric: if this one fails and everything else passes, do we release?

The metrics also get a shape. User trust sits at the top, tracked through a proxy. Four drivers sit under it: answer correctness, response latency, cost per query and multi-turn coherence. The metrics you compute from traces sit under those. When a driver drops, the tree tells you where to look first.

Every blocking threshold needs a written reason, meaning a sentence on what breaks below that number. The reasons usually come from one of four places: the output has to work at all, a cost or revenue limit, what users already have today, or the point where people give up. You predict how many of the twelve metrics should be blocking, then compare your number with a worked classification.

The practice is a metric system spec. You classify four metrics that are hard to call, write release criteria for a v2 retrieval change, and, on the extended track, put numbers on your threshold reasons and test one tradeoff: recall up 5 points, latency up 20 percent.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→