Start here

How the evaluation works

Every headline number is computed from deterministic, trial-level evidence — nothing is hand-entered.

  1. Task — a question, its schema context, the expected query, and which split it belongs to.
  2. System — provider, model, prompt, output schema, and correction policy.
  3. Trial — one system’s attempt at one task: raw output, validated intent, policy action, timing, tokens.
  4. Verifier — exact object equality, plus a field-level list of what did not match.
  5. Decision — accuracy, paired failures, latency, and marginal cost, aggregated across trials.
What this run proves: the relative behaviour of six frozen configurations on the same 88 tasks, raw and after policy. What it does not prove: production-traffic accuracy, multi-attempt stability, SQL execution correctness, or anything outside this task set.