Week 6: Decision-Making and Organization · Lesson 6.2

Translating evaluation signals to product actions

Given what we observed, what should we change next, and how will we know it helped?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 17

Speaker notes

Welcome back. Last lesson was about the decision itself. You had conflicting evidence and you picked ship, ramp, hold or roll back, and you wrote down why. In 6.1 the evidence on v2 said hold, and the team went to work on v2.1. Now jump ahead and say v2.1 fixed latency and cost, cleared its rollout, and is now live for everyone. The numbers in this lesson are a worked scenario for that release. What happens next is the part we're covering today. Your dashboard says everything is fine, and your users say it isn't. We'll go through how to spot that gap, how to trace it back to something your metrics aren't measuring, and how to turn it into a specific change with a test attached. Let's start with the decision you just made.

About this lesson

v2 was held in 6.1. Say the fix, v2.1, goes live. Two days later, in this worked scenario, every tracked metric is green and support tickets have tripled. Users say it takes forever and keeps asking them to clarify obvious questions. When user behavior and your evaluation metrics disagree like this, believe the users first, then find out what your metrics are not measuring.

You track four behaviors next to your automated scores: edit rate, reformulation rate, abandonment rate and escalation rate. A simple grid of behavior against metrics shows which case you are in. High behavioral signals with passing metrics means the metrics are incomplete, so you fix the measurement before you change the system.

In the demo you plot how much users edited each answer against the judge’s score. A cluster of answers the judge passed and users rewrote leads to one trace: a correct answer buried in 87 words of jargon. The judge scores faithfulness and completeness and has no dimension for readability.

From there you pick what to change. There are six places to change an AI system: prompts, retrieval, model config, UX constraints, guardrails and data quality. Each candidate fix gets a hypothesis (the change, the metric, how much it should move, why, and how you will check), a priority score of impact times confidence divided by effort, and staged validation: test set, shadow mode, a 10% experiment and full rollout, with a rollback if any stage fails.

The practice is a findings-to-actions plan for the v2.1 scenario, using the numbers and complaints in the lesson: explain why the dashboard missed the slow answers and the needless clarifying questions, then propose at least three ranked interventions with hypotheses and validation plans, plus at least one change to your metric suite.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→