The experiment

When a correction policy backfires

A correction policy that helps one model can be a wash on another.

40%60%80%100%Raw modelAfter correction policyQwen 7B +1/−1Qwen 3B +23/−1
The two Qwen systems (in colour) run a correction policy; the four hosted systems (grey) had none, so their raw and final scores sit on top of each other. The policy lifts 3B by 25 points but leaves 7B a net wash.

The same offline guardrail lifts Qwen 3B by 25 points (23 fixes against 1 regression) but nets to zero on Qwen 7B (one fix, one regression), because the larger model was already past the boundary the rule targets. A single “final” score would hide that entirely.