Start here

What Evaltude is

Evaltude compares complete LLM systems — not just models — on one fixed, shared workload, and connects every number back to the individual attempts behind it.

The experiment on this site asks a narrow, concrete question: given a plain-English business question, can a system produce the correct structured database query? It is scored by exact match against a known-correct answer, so “almost right” counts as wrong.

System

The model plus its prompt and any correction policy — the whole thing that answers.

Question

One task with a known-correct structured query.

Result

Correct only if the produced query matches field-for-field.

The four parts of this site are one argument, not four destinations. is the claim — which complete configurations cleared the bar, and by how much. is the evidence under it: every individual attempt, so an aggregate number can always be traced back to the answers that produced it. Docs — these pages — is the method: how the dataset, scoring, and correction policy were built, and what the result does not establish. is the general form: the evaluation engineering this experiment is one worked application of.

Reading them in that order — claim, evidence, method, method-in-general — is the intended path, and each stands on its own if you only want one. Evaluating your own scenario is a planned addition; it is not available yet, and nothing on this site accepts your data today.