AI Evals for Product DevelopmentL1.2 · 01

The evaluation
surface map

Where an AI feature can fail, what could be exploited, and what you cannot see yet
Where we left offL1.2 · 02

You measured a 0.64 gap. Could you tell where the failure happened?

What you measured in 1.1
On the v0 traces, pass@5 = 1.0 and reliable@5 = 0.36. In the made-up Q4 example, run four filtered on Q3.
What you could not answer
Was the SQL wrong, was the retrieval bad, or did the summary hallucinate?
All you had was a final output and the sparse v0 logs.
The scenarioL1.2 · 03

Six stages, and a failure in any of them looks the same from the outside.

Query parsingIntent and parameters
RetrievalContext documents
SQL generationBuild the query
SQL executionRun it, get rows
ChartsVisualize
NarrativeSummarize
"Works great for simple lookups but fails on anything complex."
"Sometimes it hallucinates metrics that don't exist."
"SQL queries occasionally fail silently."
Two questionsL1.2 · 04

Your PM will ask two questions at the launch readiness review.

What are all the ways this can fail if we roll it out to 200 users? And what haven't we tested yet?
Without a map
You can't prioritize what to test, can't say how complete your coverage is, and can't name the risk for anyone else.
With a map
You know what can fail, what someone could exploit, and what you are not testing.
Functional surfaceL1.2 · 05

Each stage fails in its own way.

Query parsing
Ambiguous intent, multi-intent queries, unsupported query types
Context retrieval
Wrong documents retrieved, missing context, low relevance scores
SQL generation
Syntax errors, wrong joins, missing filters, non-existent columns
Narrative synthesis
Hallucinated numbers, claims not in the results, incoherent summaries
SQL execution and chart rendering are covered in the exercise.
The methodL1.2 · 06

Predict from the architecture first. Then check the traces.

Architecture-driven
Look at each stage's inputs, outputs and dependencies and predict what can go wrong. No data needed.
Trace-driven
Inspect the v0 traces to confirm the predictions, find sub-types, and find what you missed.
Prediction gives you a head start. Traces refine it.
Adversarial surfaceL1.2 · 07

Adversarial inputs belong on the map from day one.

Prompt injectionCrafted inputs that change what the system does
JailbreaksBypassing safety constraints
Policy violationsGetting prohibited content generated
Private data leaksExtracting personal information from context
Resource exhaustionOverloading the system with expensive queries
In the course's adversarial file, 22 of 100 attacks got through the guardrails.
Coverage gapsL1.2 · 08

A map with no gaps listed is overconfident.

Unmeasured failure modes
Failures you suspect exist but have not seen in traces. "Narrative hallucination rate unknown."
Rare scenarios
Edge cases not in your sample. "No multi-turn sessions longer than 4 turns."
Logging blind spots
Behavior you cannot observe with current logging. "v0 doesn't log retrieved context."
Evidence sufficiencyL1.2 · 09

Split the map: what you can evaluate now, and what needs better logging first.

Can evaluate now
SQL syntax errors, the sql_error field exists. Execution timeouts, logged in sql_error. Query parse outcome, the intent label is logged.
Cannot evaluate yet
Narrative hallucination rate, no context logged. Multi-turn coherence, one trace per session. Retrieval relevance, the score is 0 on every trace.
The "cannot evaluate" list becomes your logging priorities in Week 2.
Predict before the demoL1.2 · 10

Which stage will show the most failures in the v0 traces?

Query parsing
Retrieval
SQL generation
SQL execution
Charts
Narrative
Pick one stage. Write one sentence with your reason before we advance.
The demo: one stage, worked throughL1.2 · 11

SQL generation, predicted from the architecture.

InputsIntent classification plus retrieved context
OutputAn executable SQL query
Depends onSchema catalog, SQL execution engine
Predicted failure types
Syntax errors
Wrong table joins
Missing filters
Non-existent columns
Then we read the first 20 v0 traces, and then all 500, looking at the generated_sql and sql_error fields.
The demo: what the v0 traces showL1.2 · 12

The failures v0 can show you are in SQL and the narrative. Retrieval never appears.

What v0 shows
Narratives that never came back. SQL that failed with no error message. A few logged SQL errors and timeouts.
What v0 cannot show
Retrieval, query parsing and chart failures. Nothing in the logs places a failure in those stages.
A 2025 study of a clinical text-to-SQL system found generation failures outnumbered retrieval failures about four to one.
PracticeL1.2 · 13

Build your evaluation surface map.

Base version, everyone
Predict at least two failure modes each for SQL execution and chart rendering, working from the architecture. Match each attack type on slide 7 to the stages it hits. List at least three coverage gaps. Split everything into can-evaluate-now and cannot-evaluate-yet, using the v0 fields on slide 9.
Going further
Take the v0 failure counts from the demo and place each group in a stage where the fields allow it. For the groups you can't place, write the field v0 would need. Say which attack types your functional tests would already catch.
Common mistakesL1.2 · 14

Three ways a surface map goes wrong.

Mistake
What happens
Do this instead
Annotate without predicting
You only map what you've seen. A rare catastrophic failure, like a PII leak at 0.1%, never appears in 30 traces.
Predict from the architecture before you read a single trace.
Treat adversarial as a later security review
"Later" arrives after a user exploits the system.
Map the adversarial surface alongside the functional one from the start.
List no coverage gaps
False confidence. You think you've tested everything.
Every map lists its blind spots and its logging limits.
Knowledge checkL1.2 · 15

Three judgment calls. Write your answer before you read on.

01A PM says "we tested 100 queries and 92% passed, let's ship." All 100 are simple lookups: no joins, no multi-turn. What does this tell you about coverage, and what do you recommend?
02Your map has six functional failure modes and two adversarial categories. A security researcher shows that instructions planted inside a retrieved document get followed. Why didn't your map catch it?
03You mapped four SQL failure types from 50 traces. How would you know whether you've found all the types that matter?
Next lessonL1.2 · 16

A complete map has three layers. Next, you find the failures that are actually happening.

FunctionalQuery parsing, retrieval, SQL generation, SQL execution, charts, narrative. Each with its own failure modes.
AdversarialPrompt injection, jailbreaks, policy violations, private data leaks, resource exhaustion.
Coverage gapsUnmeasured failure modes, rare scenarios, logging blind spots.
AI ANALYST LAB · aianalystlab.ai