Start here
How the evaluation works
Every headline number is computed from deterministic, trial-level evidence — nothing is hand-entered.
- Task — a question, its schema context, the expected query, and which split it belongs to.
- System — provider, model, prompt, output schema, and correction policy.
- Trial — one system’s attempt at one task: raw output, validated intent, policy action, timing, tokens.
- Verifier — exact object equality, plus a field-level list of what did not match.
- Decision — accuracy, paired failures, latency, and marginal cost, aggregated across trials.
What this run proves: the relative behaviour of six frozen configurations on the same 88 tasks, raw and after policy. What it does not prove: production-traffic accuracy, multi-attempt stability, SQL execution correctness, or anything outside this task set.