Evaltudepreview

Reference

Reproducibility & sources

Everything here is computed live from one public dataset.

This dataset is synthetic. The 88 business-style questions, expected intents, and local database records were created for this evaluation. They are not sampled from customers, production traffic, or proprietary client datasets.
Repeated trials are not part of this run. Every score above comes from a single attempt per task, so none of them carry a run-to-run error bar. Repeatability has been measured separately, on a different and much smaller workload — — and that result does not transfer to the systems compared here.
Next extension: pinning the run manifest to a tagged code revision and a captured price list — it is written down but not yet immutable — then repeated trials on this task set, before executable-SQL fixtures or online scenario execution. Until then, treat every score as provisional evidence, not a verdict.