Reference
Reproducibility & sources
Everything here is computed live from one public dataset.
- the raw evaluation data (a15.json) — every trial, intent, and metric this page renders.
- Run
a15-model-comparison-2026-07-24· verifierexact-query-intent@1· promptbaseline-intent-prompt@1· one attempt per task. - — exact model identifiers, decoding settings, versions, limitations, and what this record still cannot pin down.
Repeated trials are not part of this run. Every score above comes from a single attempt per task, so none of them carry a run-to-run error bar. Repeatability has been measured separately, on a different and much smaller workload — — and that result does not transfer to the systems compared here.
Next extension: pinning the run manifest to a tagged code revision and a captured price list — it is written down but not yet immutable — then repeated trials on this task set, before executable-SQL fixtures or online scenario execution. Until then, treat every score as provisional evidence, not a verdict.