AI Evals for Product DevelopmentL6.2 · 01

From signals
to actions

What to change when users and your metrics disagree, and how you'll know the change helped
Where we left offL6.2 · 02

Think back to your decision in 6.1.

Which piece of evidence was weakest or missing, the one that would have made you more confident?
Name the exact piece. "More data" doesn't count.
Two days after launchL6.2 · 03

Every tracked metric is green. Support tickets tripled.

Monitoring dashboard
SQL success rate79%, steady
Judge score0.81, steady
Latencywithin SLA
Hallucination ratestable
What users are saying
Support tickets3x
"System takes forever now."
"Why does it keep asking me to clarify obvious queries?"
Signal-metric divergenceL6.2 · 04

When users and metrics disagree, believe the users first. Then find out why the metrics missed it.

First, fix the measurement
Add the behavioral signal as a tracked metric, recalibrate the judge, or add the missing rubric dimension.
Then, fix the system
Once you can measure the problem, you can tell whether a change to the system fixed it.
Behavioral signalsL6.2 · 05

Four user behaviors tell you what your quality metrics miss.

Edit rateShare of outputs users modify before they use them
Reformulation rateShare of queries followed by a clarifying or correcting follow-up
Abandonment rateShare of sessions that end without the task done
Escalation rateShare of interactions where the user asks for human help
Log these as tracked metrics from day one, alongside the automated scores.
The diagnosticL6.2 · 06

Put behavior and metrics on one grid.

Eval metrics pass
Eval metrics fail
Behavioral signals high
DivergenceMetrics are lagging. Believe the users, find the metric gap.
System brokenBoth agree something failed.
Behavioral signals low
System OKBoth agree it works.
Metrics sensitiveThe metric flags something users don't feel yet.
PredictL6.2 · 07

48 hours after v2.1 went out:

Moved
Edit rate35% → 51%
Abandonment rate12% → 19%
Held
SQL success ratesteady at 79%
Judge scoreunchanged at 0.81
What quality dimension is the judge probably not scoring? Write one sentence before we go on.
The demo: cross-tabL6.2 · 08

Plot every trace: how much the user edited against what the judge scored.

The data
200 traces. For each one, the share of the narrative the user changed, and the judge's composite score for faithfulness and completeness.
What you're looking for
A cluster of traces the user edited heavily (above 0.5) that the judge passed (above 0.7).
This per-trace edit share is a different number from the dashboard edit rate, which counts outputs that got any edit at all.
The demo: one traceL6.2 · 09

One divergent trace, from both sides.

What the judge scored
Facts accurate
All relevant data points included
SQL and chart correct
What the user read
"The mobile signup funnel exhibited a conversion rate trajectory with progressive attrition across sequential engagement touchpoints, culminating in a terminal conversion rate of 12.3%..."
What they wanted: "Mobile signup conversion was 12.3% last month."
Action surfacesL6.2 · 10

Six places you can change an AI system. Prompts are one.

Surface
What changes
AI Data Analyst example
Prompts
System prompt, few-shot examples
Add conciseness constraints to the narrative prompt
Retrieval
How context is fetched: query reformulation, ranking
Adjust the ranking threshold to cut latency
Model config
Model choice, temperature, max tokens
Use a smaller model for simple queries
UX constraints
Input validation, output filtering
A "Show SQL" toggle instead of long automatic explanations
Guardrails
Refusal logic, confidence thresholds
Ask for clarification only below a confidence threshold
Data quality
Rubric calibration, oracle expansion
Add a readability dimension to the narrative judge
Intervention hypothesisL6.2 · 11

Every change gets a hypothesis before anyone builds it.

If we [change], we expect [metric] to move by [amount] because [mechanism], validated via [method].
No hypothesis
"Rewrite the prompt."
Hypothesis
If we add conciseness constraints to the narrative prompt, we expect edit rate to drop 10 to 15 percent because users will have less to rewrite, validated offline and then in a limited experiment.
PrioritizeL6.2 · 12

Rank the candidates with one score.

Priority = (Impact × Confidence) / Effort
Impact, 1 to 55 removes a root cause. 1 treats a symptom.
Confidence, 1 to 55 has strong evidence behind it. 1 is a guess.
Effort, 1 to 51 is a config change you can deploy in an hour. 5 is re-architecting over several sprints.
Tie? Pick the one that teaches you more, even if it fails.
Validate in stagesL6.2 · 13

Each stage has a pass bar. Fail any stage and you roll back.

Test set evalTarget metric improves. No other metric regresses by more than 0.05.
Shadow modeRuns on live traffic, users don't see it. No latency spike over 5%, no new failure modes.
Limited experiment10% exposure. Guardrails hold, primary metric improves, no segment regresses.
Full rolloutImprovement holds for 7 days, no delayed effects.
The demo: one full interventionL6.2 · 14

The verbosity finding, written up as a change you can test.

Failure mode
Narrative verbosity: facts are right, the text is long and full of jargon, users rewrite it
Action surfaces
Prompts, plus data quality for the judge
Hypothesis
If we add conciseness constraints to the narrative prompt, edit rate drops 10 to 15 percent because users have less to rewrite
Validation
Test set: shorter narratives, judge score within 0.05 · Shadow: no latency spike · 10% experiment: edit rate 10 to 15 percent lower than control · Rollout: holds 7 days
Priority
Impact × Confidence / Effort
PracticeL6.2 · 15

Build a findings-to-actions plan from the v2.1 scenario.

Base version, everyone
Use the numbers and complaints from this lesson. Write one sentence on what you found. Take the two user complaints, slow answers and needless clarifying questions, and write a one-sentence hypothesis for each. Propose at least 3 interventions with scores and validation plans. Rank them. Propose at least 1 change to your metric suite.
Extended version, DS and engineering
Write out how you'd measure per-trace edit share from the original and edited narrative, and how it differs from the dashboard edit rate. Plan the split of the 36 divergent traces by user role and query complexity: what you'd expect, and what would change your ranking.
Knowledge checkL6.2 · 16

Two judgment calls. Write your answer before you read on.

01SQL success is steady at 79% and judge scores at 0.81. Edit rate jumped from 38% to 54% and support tickets doubled. A colleague says, "The metrics are fine, users are just complaining more." What's wrong with that, and what do you do?
02A teammate says, "The narratives feel off, let's rewrite the prompt." What do you ask them before anyone touches it?
Next lessonL6.2 · 17

You can turn a finding into a tested fix. Next, which fixes go first across the whole backlog.

DetectBehavior against metrics on the grid and the scatter plot
MapEach failure mode across the six action surfaces
HypothesizeChange, metric, amount, mechanism, method
ValidateTest set, shadow, 10% experiment, rollout
AI ANALYST LAB · aianalystlab.ai