Week 6: Decision-Making and Organization · Lesson 6.1

Decision-making under uncertainty

Given conflicting evidence, what decision is justified?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 15

Speaker notes

Welcome to Week 6. For five weeks we've been building the evaluation side of the AI Data Analyst. Instrumentation, metrics, judges, experiments, and last week, launch readiness. So you can now tell whether a change made the system better, worse, or did nothing. What that doesn't do is make the decision for you. Once you can see everything, you can see the tradeoffs too. Quality goes up on three measures, and latency and cost go up with it. One group of users gains less than the others and pays the same price. And somebody in the room still has to say what happens next. That's this week. Today we build a way to turn mixed evidence into a decision you can defend, write down, and reverse if it goes wrong. Let's start with where Week 5 left you.

About this lesson

By Week 6 you can tell whether a change made the AI Data Analyst better or worse. That doesn’t make the decision for you. In the course’s experiment, v2 raised SQL success by 2.7 points, task completion by 5.4 and retrieval precision by 7.9. It also raised average latency 17.6 percent and cost per query 20 percent, past both guardrails, and the share of users averaging over the 2-second SLA went from 1.6 to 10.6 percent. Your PM wants to know if it goes out anyway.

With only “ship” and “don’t ship” to choose from, a result like that turns into an argument. This lesson gives you six decision types instead: ship, ramp, hold, roll back, scope-restrict, and deploy with a human in the loop. Each one needs a different strength of evidence.

To pick between them you ask four questions. Which way did each metric move? By how much? How much can you trust the numbers? And if it goes wrong after deploy, how quickly would you find out? Then you write the tradeoff down in one sentence: what you recommend, what improved, what got worse, and what limits the risk.

You work through v2 by user group. Every group gained, executives least, and every group paid the same latency and cost, so there’s no group to ship to and scope-restrict is out. The slowness lands on users who ask complex questions, 31 percent of whom average over 2 seconds. The evidence points to hold: v2 stays off while the team fixes latency and cost, starting with complex questions.

The practice is a decision memo for v2 with six fields: recommendation, primary evidence, tradeoff acknowledgment, risk mitigation, rollback trigger and decision confidence. You answer confidence and risk containment from the numbers in the lesson and check that every field names a metric and a number. The extended version works out the daily cost at 100,000 queries, checks whether any user group stays inside both guardrails, and writes a monitoring spec for v2.1.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→