Skip to content
Jev Tracker
BenchmarksExperimentsAccessLimitationsLatest
Explore Jev Tracker
BenchmarksExperimentsAccessLimitationsLatest
Home/Benchmarks

Jev benchmarks, in context

How fast is Jev, really?

Read the headline numbers, inspect the test conditions, and separate a provider claim from an independent result.

Read the claims Check the conditions
Published claimProvider evaluation

Headline result
Reported test setup

compare only when
Comparable evidenceMatched conditions

Same task and input
Same baseline and rules

Reading guide only. This is not a benchmark chart or an independent reproduction.

Read this as a claims guide, not a leaderboard. We checked the publications, not the model’s performance. No independent reproduction is included in this edition.

The numbers, with the fine print

Official headline ratios. Not Jev Tracker measurements.

Speed / Official claim

Up to193.6×

Faster, in the provider’s evaluation

TypeSafe’s workflow evaluation headline, not a result for every task.

The provider describes these gains as being at the high end of expected real-world gains. Not our measurement.

Which baseline produced this ratio?

Exact model/configuration behind this headline ratio: not independently reconstructed by Jev Tracker.

Official

Published Sep 15, 2026
Source checked Sep 19, 2026, 04:42 UTC

View source

Cost / Official claim

Up to444.6×

Cheaper, in the provider’s evaluation

TypeSafe’s workflow evaluation headline, not a universal price discount.

A workflow comparison does not predict your bill. Jev Tracker has not reproduced this ratio.

Which baseline produced this ratio?

Exact model/configuration behind this headline ratio: not independently reconstructed by Jev Tracker.

Official

Published Sep 15, 2026
Source checked Sep 19, 2026, 04:42 UTC

View source

What the evaluation actually measures

Published conditions on the left. Our interpretation on the right.

On narrow screens, scroll the table horizontally to read every field.

Published setup and what it means for comparison
ConditionWhat the source saysHow to read it
Tasks

Security incidents, agent trace observability, invoice processing and customer service.

View source
Four bounded workflows, not a broad test of chat, writing or general intelligence.
Aggregation

Each plotted model configuration averages accuracy, cost and time equally across the four workflows.

View source
An aggregate can hide differences between tasks. It is not a median visitor experience.
Reference answers

Consensus labels average GPT-6 Astra and Claude Fable 5.1 at high thinking; other models use default reasoning.

View source
Accuracy here means agreement with the reference. Unequal reasoning settings matter when interpreting speed.
Decision harness

LLM comparisons use a wrapper that returns structured decisions and probabilities.

View source
That output contract can add work. Do not treat it as an ordinary chat-response comparison.
Timing context

The launch post describes runs from a West Coast laptop, generally near service locations.

View source
Network location is part of elapsed time. Local results are not a latency guarantee for every region.

Still needed for a like-for-like comparison

For this edition, we have not reconstructed the exact peak-ratio model pair, run-level timings, input sizes or a matched sample denominator. These are gaps in our verification, not a claim that the provider has never published them.

A correct format is not a correct decision

Keep three questions separate: was the output well formed, did it match a reference answer, and did the whole task succeed?

See decisions inside real examples

A concrete example / Customer service

From account records to a support action

The official workflow reads a conversation and account information, then combines model judgments with application rules. Its published examples include disagreements and a case where all compared models miss the reference.

A selected example is not an error-rate sample. It helps explain the failure mode, not how often it happens.

Official

Source checked Sep 19, 2026, 04:42 UTC

View source

Who is saying what?

A repeated claim does not become an independent test.

Official evaluation
Useful for understanding the provider’s setup and claim. Not independent verification.
Integration coverage

Vercel attributes the same headline ratios to TypeSafe. A second publisher repeating a claim is not a second measurement.

Integration

Published Sep 16, 2026
Source checked Sep 19, 2026, 04:42 UTC

View source
Independent tests
No independently reproduced result has been qualified for inclusion here. That does not mean none exists.
Community observations
No community average, median or sample count is published in this edition. Original methods and comparable conditions must be checked first.

Before you compare two models

How to compare results fairly.

Check current access and pricing
Does a faster decision mean a faster task?

Not necessarily. Count input preparation, network time, downstream actions, retries and failed decisions. Compare completion quality before comparing elapsed time.

Is a cost ratio the same as a price discount?

No. A workflow’s cost depends on its inputs, number of calls, output contract and retries. Published unit pricing and the cost of a successful task answer different questions.

Does type safety prove correctness?

A value can fit the allowed format and still be the wrong choice. Evaluate answer quality against labeled examples, then examine what happens when the application acts on a mistake.

What should a fair comparison disclose?

The same task and inputs, exact model versions, reasoning settings, location, concurrency, timing boundaries and success criteria. Report failures and variation, not just the most favorable run.

Follow the evidence

“Source checked” records a publication check, not a successful reproduction. Undated sources keep their publication date unspecified. All check times below are UTC.

  • Introducing System One Models & Jev
    Official

    Published Sep 15, 2026
    Source checked Sep 19, 2026, 04:42 UTC

    View source
  • TypeSafe workflow evaluations
    Official

    Source checked Sep 19, 2026, 04:42 UTC

    View source
  • Customer service workflow evaluation
    Official

    Source checked Sep 19, 2026, 04:42 UTC

    View source
  • TypeSafe AI Jev now available on AI Gateway
    Integration

    Published Sep 16, 2026
    Source checked Sep 19, 2026, 04:42 UTC

    View source
Read our correction principles
Jev Tracker

Independent coverage of Jev. No affiliation, partnership or endorsement by TypeSafe AI.

Evidence before conclusions.

Sources checked. Uncertainty stated. A source label is not a seal of approval.

About and editorial method

Keep the record clear.

We explain what changed and why.

Corrections and updates
ContactPrivacy

Jev Tracker. Editorial illustrations are concepts, not product screenshots.