Mistake
What happens
Do this instead
Ship on pass@k alone
Say pass@5 = 0.95 and reliable@5 = 0.3. About 65% of queries pass on some runs and fail on others.
Always compute both. If the gap is above 0.2, add mitigation before shipping.
Assume temperature 0 is deterministic
One test per query passes. In production, re-runs still differ because of hardware differences.
Measure variance by running multiple trials.
Use a benchmark as the ship decision
Say 92% on a standard benchmark, then failures on real user queries.
Benchmarks screen. Product evaluation decides.