Mistake
What happens
Do this instead
One fixed number for everything
"We always use 100." A system near 50% gets under-sampled; a system near 95% gets over-sampled.
Calculate n for each decision and each success rate.
Ignoring the evaluator
A kappa 0.6 judge with no adjustment gives you about 40% less evidence than you think.
Adjust n for evaluator reliability.
Precision you don't need
Collecting 500 samples for ±3 when the PM would accept ±8 for a ramp.
Ask what decision you're making before sizing.