The experiment

Reading the leaderboard

Read the leaderboard as three questions, not one score.

0%25%50%75%100%Claude Haiku 4.595%100%50%OpenAI Strong95%98%63%Claude Sonnet 4.694%98%50%OpenAI Small93%95%63%Qwen 7B + guardrail84%87%25%Qwen 3B + guardrail70%68%13%
Three bars per system, dark to light: overall, held-out, stress set. Every system holds up on held-out questions, then drops sharply on the eight-question stress set.

Overall is the share of all 88 tasks correct. Held-out is the generalization number that matters most. Stressis the diagnostic probe. Two systems tie for the highest overall accuracy, so “best” is a judgement call, not a ranking.

— click any row there to see exactly which questions that system missed.