Network location is part of elapsed time. Local results are not a latency guarantee for every region.
Still needed for a like-for-like comparison
For this edition, we have not reconstructed the exact peak-ratio model pair, run-level timings, input sizes or a matched sample denominator. These are gaps in our verification, not a claim that the provider has never published them.
A correct format is not a correct decision
Keep three questions separate: was the output well formed, did it match a reference answer, and did the whole task succeed?
The official workflow reads a conversation and account information, then combines model judgments with application rules. Its published examples include disagreements and a case where all compared models miss the reference.
A selected example is not an error-rate sample. It helps explain the failure mode, not how often it happens.
Not necessarily. Count input preparation, network time, downstream actions, retries and failed decisions. Compare completion quality before comparing elapsed time.
Is a cost ratio the same as a price discount?
No. A workflow’s cost depends on its inputs, number of calls, output contract and retries. Published unit pricing and the cost of a successful task answer different questions.
Does type safety prove correctness?
A value can fit the allowed format and still be the wrong choice. Evaluate answer quality against labeled examples, then examine what happens when the application acts on a mistake.
What should a fair comparison disclose?
The same task and inputs, exact model versions, reasoning settings, location, concurrency, timing boundaries and success criteria. Report failures and variation, not just the most favorable run.
Follow the evidence
“Source checked” records a publication check, not a successful reproduction. Undated sources keep their publication date unspecified. All check times below are UTC.