Mistake
What happens
Do this instead
Measuring at the wrong level
Per-query scores look fine while sessions contain wrong answers.
Match the unit to how users experience the feature.
Everything is blocking
The team holds a release over wordy text while SQL errors keep going out.
Classify first, then prioritize the blocking metrics.
Averaging away failures
One wrong answer in ten becomes a 90% session score.
Use worst-case aggregation where one failure breaks trust.
Wrong archetype for the feature
Extraction metrics check completeness and miss made-up claims.
Pick the archetype from what the user does with the output.
Thresholds without failure cost
95% sounds strict until you count the failures at volume.
Multiply the failure rate by daily queries before you agree.