Start here
What Evaltude is
Evaltude is a field guide and evidence lab for evaluating complete AI systems. It combines a reusable curriculum, documented evaluations, engineering analysis and inspectable trial-level evidence.
The current measured evaluation asks one narrow question: given a plain-English business question, can a complete system produce the correct structured query intent? It is scored by exact match against a known-correct answer, so “almost right” counts as wrong.
The model plus its prompt and any correction policy — the whole thing that answers.
One task with a known-correct structured query.
Correct only if the produced query matches field-for-field.
How the site is organized
Explains Evaltude and routes visitors to the reusable method or a documented evaluation.
Contains the curriculum, worked evaluations, supporting evidence, engineering write-ups and reference material.
Lead with a practical decision, apply the method, state a recommendation and expose the evidence behind it.
Start with the homepage, browse the unified Documentation, or choose a worked evaluation. Each measured evaluation links to its trial-level evidence after explaining the scenario, comparison and decision.