Reference

Reproducibility & sources

Everything here is computed live from one public dataset.

Repeated trials are not part of this run. Every score above comes from a single attempt per task, so none of them carry a run-to-run error bar. Repeatability has been measured separately, on a different and much smaller workload — — and that result does not transfer to the systems compared here.
Next extension: pinning the run manifest to a tagged code revision and a captured price list — it is written down but not yet immutable — then repeated trials on this task set, before executable-SQL fixtures or online scenario execution. Until then, treat every score as provisional evidence, not a verdict.