Error
What the team does
What goes wrong
Optimistic single-sample estimate
"62% pass rate. Ship it."
The next batch comes back at 55%. The estimate never had error bars.
Capability confused with reliability
"pass@5 = 0.90, works great."
reliable@5 = 0.60. Nearly half of users hit at least one failure.
Variance source confused
"The system is too random. Lower the temperature."
85% of the variance was evaluator noise. A sprint spent fixing the wrong thing.
Made-up numbers, to show the pattern.