The experiment
The three question sets
The 88 tasks are split three ways, each answering a different question about a system.
- Development · 20 tasks
- The tuning set. These questions informed the prompt and the guardrail, so a system scoring well here has partly been fitted to them — descriptive, not proof of generalization.
- Held-out · 60 tasks
- The generalization test: questions no system was tuned on. This is the honest read on real quality.
- Adversarial (stress set) · 8 tasks
- A small diagnostic probe of deliberately misleading phrasing. With only 8 questions, one task moves the score by 12.5 points — it exposes failure modes but is far too small to rank models.