Judge modelWhich model scores the outputs
Judge prompt versionSo a prompt change forces a re-check
Calibration setName, size and splits
TPR, with intervalFrom the test split
TNR, with intervalFrom the test split
Cohen's KappaAgreement with human labels
Observed pass rateWhat the judge reports
Corrected pass rateRogan-Gladen, with interval
Position biasScore and conclusion
Verbosity biasScore and conclusion
Recommended actionsFor example, run both orders
LimitationsWhat the card doesn't cover