Week 6: Decision-Making and Organization · Lesson 6.4

Ownership model: the AI Reliability Lead

Who owns the data, the metric, and the ship decision, and how do we prevent ambiguity?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome back. Last lesson you took the evaluation evidence and turned it into a ranked backlog: what to fix first, what done looks like for each item, and how you'll know it worked. That backlog assumes somebody picks up each item. It also assumes the evidence behind the ranking is still good next month. Somebody has to keep the judges calibrated, keep the regression suite current, and watch the dashboard after the release goes out. Today is about who that somebody is. We'll look at an AI Data Analyst team six months after launch, see what happened to its evaluation system when nobody owned it, and then build the ownership model that would have prevented it: who owns the data, who owns the metrics, who owns the ship decision, and what happens when something breaks.

About this lesson

An evaluation system that works at launch can stop working six months later without anyone deciding anything. The judges never get recalibrated, the regression suite misses the new query types, nobody watches the dashboard, and the thresholds end up in a Slack thread. Each person assumed someone else had it.

The lesson starts with that audit on the AI Data Analyst. You predict what happened to a SQL judge’s false positive rate after the system moved to a new model, and work out why nobody can answer without recalibrating it.

Then you build the ownership model. Someone owns the data (the rows), someone owns what each metric means and whether it is measured correctly (the columns), and someone owns the decision to ship, ramp, hold or roll back. The AI Reliability Lead looks after the health of the whole evaluation system and tells the ship review how much the evidence can be trusted.

A RACI matrix writes the ownership down, with one Accountable owner per activity, never zero and never two. An evaluation debt register tracks the maintenance that did not happen, each item with a risk level, an owner and a target date. Escalation triggers set in advance what gets escalated, to whom and how fast.

The practice is an ownership model for your capstone: the lesson’s RACI extended with at least three new activities and one A on every row, at least five debt items from the audit and your own Week 2 to 5 work in a register, and at least three review types and three escalation triggers.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→