Joint-highest 95.5% overall · 60/60 held-out
Evaluate complete LLM systems — not model names.
6 configurations. The same 88 tasks. Every result inspectable.
A complete system is the model, prompt, schema, validation, and correction policy you deploy together. Compare all 6 on identical tasks across accuracy, latency, cost, and the individual failures behind each number.
Model, prompt, schema, validation and correction policy — scored together, because they ship together.
A business question paired with the exact structured query it should produce.
A green check means the query matched field-for-field. Anything off is a red miss.
Built for engineers and technical leads choosing a production configuration — and for anyone learning how this kind of evaluation is done. The four parts of this site are one argument, read in this order:
Scenario evaluation is planned. Evaltude does not currently accept user data.
Featured experiment
A1.5 · Structured intent model comparison
Six complete LLM system configurations evaluated on the same development, held-out, and adversarial semantic tasks.
1.34s · directional legacy run
$0.80 per 1,000 sequential
System leaderboard
How the 6 systems compare
Each row is one complete system, ranked by overall accuracy. The two Qwen rows show accuracy after their offline correction policy, with the raw model score underneath; the four hosted rows had no policy applied. Click any row to see the exact questions it got wrong.