Mistake
What happens
Do this instead
Tuning on the full dataset
Kappa looks like 0.85 on the data you tuned on and drops to 0.68 on fresh data.
Split first. Report on data the judge never influenced.
Running the holdout every iteration
After ten peeks, the holdout is a second dev set.
One use per release cycle, then rotate it into dev.
Not versioning
A 5% improvement can't be reproduced after 30 examples were added and 10 labels changed.
Version every change. Experiments cite a version.
Never rotating
The suite measures last year's questions.
Check drift quarterly. Rotate when p is below 0.05.