The experiment
Cost and latency, honestly
Cost and latency are inputs to a routing decision — not clean model rankings.
Cost here is a marginal dedicated-GPU estimate (summed request time); it excludes startup, idle, and scale-down, and the hosted systems are token-priced — so it is not a like-for-like production bill.
Latency is deliberately not used to rank models: the legacy harness did not control per-request client reuse, so those numbers are directional only. The production takeaway is a routing policy — send most traffic to a strong-and-cheap system, escalate hard cases — not a single champion.