Measuring everything that's availableA model change moves half the metrics up and half down, and nobody can decide. Assign roles first.
Using similarity when you could run itScoring SQL by text overlap when you could execute it and compare results.
Reporting a point estimate alone"Correctness is 76%" with no interval doesn't say how much to trust it.
Mixing up the two variancesBlaming the system for scores that move because the judge is noisy, or the reverse.
Gating on an unproven signalBlocking releases on a metric that hasn't been shown to be stable.