Public experiment · run complete

Evaluate complete LLM systems — not model names.

6 configurations. The same 88 tasks. Every result inspectable.

A complete system is the model, prompt, schema, validation, and correction policy you deploy together. Compare all 6 on identical tasks across accuracy, latency, cost, and the individual failures behind each number.

1 · SystemA full setup, not just a model

Model, prompt, schema, validation and correction policy — scored together, because they ship together.

2 · QuestionOne task with a known right answer

A business question paired with the exact structured query it should produce.

3 · ResultCorrect only if it matches exactly

A green check means the query matched field-for-field. Anything off is a red miss.

Built for engineers and technical leads choosing a production configuration — and for anyone learning how this kind of evaluation is done. The four parts of this site are one argument, read in this order:

Scenario evaluation is planned. Evaltude does not currently accept user data.

Featured experiment

A1.5 · Structured intent model comparison

Complete
88 tasks 6 systems 528 trials Jul 24, 2026

Six complete LLM system configurations evaluated on the same development, held-out, and adversarial semantic tasks.

20 development60 holdout8 adversarial
Best observed accuracy95.5%2 systems tied · 84 / 88 correct
Best default candidateClaude Haiku 4.5

Joint-highest 95.5% overall · 60/60 held-out

Lowest observed p50 OpenAI Small

1.34s · directional legacy run

Lowest marginal estimate Qwen 3B + guardrail

$0.80 per 1,000 sequential

System leaderboard

How the 6 systems compare

Each row is one complete system, ranked by overall accuracy. The two Qwen rows show accuracy after their offline correction policy, with the raw model score underneath; the four hosted rows had no policy applied. Click any row to see the exact questions it got wrong.

SystemOverallHeld-outStress set*Speed*Cost / 1K*

What the evidence says

A winner is not the same as a solved task.

Claude Haiku 4.5 and OpenAI Strong tie at 95.5% overall, but even the best system still misses the hardest cases. The honest takeaway is provisional: pick the best system you can measure today, then keep expanding the test.

Best score on the stress set62.5%only 5 of 8 right