AI Evals for Product DevelopmentL4.2 · 01

Metric design
patterns

Matching a measurement pattern to each part of an AI feature
Where we left offL4.2 · 02

In 4.1 you gave every metric a job. How does each job affect the release decision?

Blocking metric
What happens to the release when it misses?
Optimization metric
What happens to the release when it misses?
Answer from memory before we move on.
The scenarioL4.2 · 03

You have a score for every component, and you still can't answer "are we ready?"

Retrievalrecall@5
SQL generationSQL correctness
Narrativegrounding score
Each team built its metric on its own. No one decided how the three add up to a release. The scores in this lesson are made up.
The scenarioL4.2 · 04

Metrics built one component at a time leave three questions open.

01Which metric can hold the release?
02Which metric is tracked for improvement?
03At what level do we measure: one response, one question, one session, or the user's whole workflow?
Measurement archetypesL4.2 · 05

An archetype is a reusable measurement plan for one type of AI feature.

UnitOutput, task, session or workflow. It should match how users experience the feature.
Core dimensionsWhich kinds of quality matter for this feature type.
Blocking metricsThe must-pass thresholds.
Optimization metricsTracked and improved over time.
AggregationHow scores roll up: worst case, average, or something else.
Six archetypesL4.2 · 06

Six archetypes cover most AI features.

Archetype
Feature type
Default unit
Blocking examples
Drafting
Generated content the user edits
Output plus edit signal
Grammar errors, safety violations
Summarization
Long content condensed to key points
Output
Hallucinated facts, missing key information
Extraction
Structured data from unstructured input
Output
Format errors, missing required fields
RAG
Retrieval plus generation
Component
Retrieval misses, unsupported claims
Agents
Multi-step tool use
Session
Task failure, wrong tool
Decision support
Recommendations or guidance
Outcome
Confident wrong recommendations
The AI Data AnalystL4.2 · 07

The AI Data Analyst is a hybrid. Give each component its own archetype.

RetrievalRAG archetype. Measured separately from what gets written.recall@5
SQL generationTool call, measured like extraction: run it and compare the result with a known answer.SQL correctness
NarrativeSummarization archetype. Does the text match the query result, and is it complete?grounding
Predict before the demoL4.2 · 08

Which unit of measurement fits the narrative component?

OutputOne narrative per query, judged on its own.
TaskWhether the user's analytical question got answered.
SessionCoherence across a multi-turn conversation.
WorkflowWhether the user makes a good decision from the analysis.
Pick one and write one or two sentences on why, before we advance.
The answerL4.2 · 09

Measure each narrative, then roll it up to the question the user asked.

Where you measure
Output level. Each narrative gets its own grounding and completeness scores.
What you report
Task level. Did the user's question get a correct, usable answer?
The aggregation strategy is how you get from the first to the second.
The demo: component-level evaluationL4.2 · 10

One end-to-end score can't tell you which component to fix.

End-to-end quality
One combined score for retrieval and narrative. A bad retrieval and a bad write-up look the same.
Component level
Retrieval recall@5 and narrative grounding scored separately. Different problems, different fixes.
The demo: blocking or optimizationL4.2 · 11

SQL correctness blocks the release. Narrative conciseness gets tracked.

SQL correctness, blocking at 90% or above
A wrong number looks the same as a right one. The user puts it in a deck and presents it.
Narrative conciseness, target 0.85
A wordy narrative is annoying to read. The numbers in it are still right.
PracticeL4.2 · 12

Apply the archetypes to the AI Data Analyst.

Base version, everyone
Pick the entity in the question, like the metric or the time period, whose extraction errors would hurt most downstream, and say why. Justify a 95% precision bar for refusal detection in one sentence. With SQL at 89% per query, work out a ten-question session's pass rate two ways: worst case and average. Fill the archetype template for one component.
If you finish early
Fill the template for all three components: retrieval, SQL generation and narrative.
The templateL4.2 · 13

The archetype template for the AI Data Analyst.

Component
Metric type
Metric
Threshold or target
SQL generation
Blocking
SQL correctness
>= 90%
Retrieval
Blocking
Recall@5
>= 0.80
Narrative
Optimization
Conciseness
target 0.85
Add the current score and a pass or fail for each blocking row, then make the release call.
Common mistakesL4.2 · 14

Five ways metric design goes wrong.

Mistake
What happens
Do this instead
Measuring at the wrong level
Per-query scores look fine while sessions contain wrong answers.
Match the unit to how users experience the feature.
Everything is blocking
The team holds a release over wordy text while SQL errors keep going out.
Classify first, then prioritize the blocking metrics.
Averaging away failures
One wrong answer in ten becomes a 90% session score.
Use worst-case aggregation where one failure breaks trust.
Wrong archetype for the feature
Extraction metrics check completeness and miss made-up claims.
Pick the archetype from what the user does with the output.
Thresholds without failure cost
95% sounds strict until you count the failures at volume.
Multiply the failure rate by daily queries before you agree.
Knowledge checkL4.2 · 15

Three judgment calls. Write your answer before you read on.

01The narrative has two candidate metrics: an LLM judge for grounding, and a rule-based check that all key numbers appear. You have room for one blocking metric. Which one, and why?
02A drafting feature, like an email composer, scores tone 0.88, grammar 0.97 and length 0.72. Which metric should block the release?
03An agent scores tool call correctness 0.94, task success 0.78 and workflow completion 0.65. Which level should drive the release decision?
Next lessonL4.2 · 16

Start from what the user does with the output. Next: what all this measurement costs.

1. Match the feature to an archetype
Edits it → draftingReads it for facts → summarizationValidates the data → extractionAsks follow-ups → RAGHas a multi-step goal → agentMakes a decision → decision support
2. Place each metric
Quality
Performance
Blocking
SQL correctness, grounding
Latency cap
Optimization
Conciseness, tone
Token efficiency
AI ANALYST LAB · aianalystlab.ai