Week 1: Foundations and Economics · Lesson 1.2

Product evaluation framework for AI systems

What is the full evaluation surface for an AI feature, from inputs to user outcomes?

← All lessons
Browse lessons

Free, self-paced. Read the deck with its speaker notes, work the practice from the slides, then take the week's quiz for a certificate.

Slide 1 of 16

Speaker notes

Welcome back. Last lesson you ran the same query five times, saw four right answers and one wrong one, and put pass@k and reliable@k on it. So you can now say how inconsistent a system is. What you can't say yet is where it goes wrong. Was the SQL wrong? Did the retrieval pull the wrong context? Did the summary make a number up? You can't fix what you can't locate. Today we build a map of every place the AI Data Analyst can fail, every way someone could deliberately misuse it, and every blind spot in what you can currently observe. Let's start with what you already measured.

About this lesson

Knowing that a system fails one time in five does not tell you where. Was the SQL wrong? Did retrieval pull the wrong context? Did the summary make a number up? With the thin logs a v0 system writes, all three look the same from the outside.

The AI Data Analyst has six stages: query parsing, context retrieval, SQL generation, SQL execution, chart rendering and narrative synthesis. Each one fails in its own way, and you cannot test SQL correctness the way you test narrative accuracy. So you build a map organized by stage.

The map has three layers. The functional surface is what can go wrong when a regular user asks a regular question. The adversarial surface is what someone could do on purpose: prompt injection, jailbreaks, policy violations, private data leaks, resource exhaustion. The coverage gaps are what you cannot see yet: failures you suspect but have not observed, rare scenarios not in your sample, and behavior your logging cannot record at all.

The method is to predict failure modes from the architecture first, by looking at each stage’s inputs, outputs and dependencies, and then read real traces to confirm the predictions and find what you missed. You work through SQL generation as the example, then look at what the 500 v0 traces can show. The visible failures are in SQL and the narrative. Retrieval never appears, because v0 logs nothing about it.

The last step splits the map into what you can evaluate with today’s logging and what needs better instrumentation first. That second list becomes your Week 2 priorities.

The practice is your own evaluation surface map: failure modes for SQL execution and chart rendering predicted from the architecture, each attack type matched to the stages it hits, at least three coverage gaps, and the evidence split.

Go deeper with AI Analytics for Everyone

5-week course: metrics, root cause analysis, experimentation, and storytelling. Think like a Product Data Scientist.

Book 1-on-1 with Shane

30-minute AI evals Q&A. Talk through your specific evaluation challenges and get hands-on guidance.

Finished all 36 lessons? Take the exam and get your free AI Evals certification.

→