Mistake
What happens
Do this instead
Thresholds too tight
Dozens of alerts in the first week. The team mutes the channel and a real regression gets lost.
Set lines from normal variation and require the metric to stay bad for a few hours.
Retiring the scripts
The platform goes down during a deployment and there is no monitoring at all.
Keep the 5.5 scripts running as the fallback.
Uniform 1% sampling
A small, high-stakes group produces almost no sampled traces, so its failures stay invisible.
Sample high-risk groups at 100% and the rest at 1%.