The experiment

The three question sets

The 88 tasks are split three ways, each answering a different question about a system.

Development · 20 tasks
The tuning set. These questions informed the prompt and the guardrail, so a system scoring well here has partly been fitted to them — descriptive, not proof of generalization.
Held-out · 60 tasks
The generalization test: questions no system was tuned on. This is the honest read on real quality.
Adversarial (stress set) · 8 tasks
A small diagnostic probe of deliberately misleading phrasing. With only 8 questions, one task moves the score by 12.5 points — it exposes failure modes but is far too small to rank models.