The experiment
When a correction policy backfires
A correction policy that helps one model can be a wash on another.
The same offline guardrail lifts Qwen 3B by 25 points (23 fixes against 1 regression) but nets to zero on Qwen 7B (one fix, one regression), because the larger model was already past the boundary the rule targets. A single “final” score would hide that entirely.