Mistake
What happens
Do this instead
Results live in Slack
"Has SQL correctness improved?" No one can say, because nothing was stored.
Every run writes a row to eval_runs.
Incomplete metadata
87 to 83 could be a regression or a new judge. You can't tell which.
Record dataset, judge, sample and code version on every run.
No lineage
How the evaluation changed gets rebuilt from memory.
Record which run each new run refines.
Same depth everywhere
The full suite on every production request adds latency and cost.
Use the environment ladder.
Stages that aren't connected
Aggregation runs on last week's judge scores after the judge step fails.
Explicit stages with defined inputs and outputs.