Frontier LLMs are up to 25 points more accurate than Jev, but cost over 100x more

Every step up the Pareto frontier buys accuracy at a much higher price. Run a cheap judge on every trace, and save the frontier models for the cases where the extra points matter.

This benchmark was created using LangWatch. Sign up to create your own · Read the method

LangWatch Instant Evals runs on Jev. Check it out

Task
Set

LLM-as-a-judge, core

Accuracy, % · Higher is better

Pick: GPT-6.1 Sol, tied with Claude Opus 5.5 at about half its cost.

  1. , interval 89.1 to 95.8, tied for the lead
  2. , interval 83.9 to 93.9, tied for the lead
  3. , interval 78.4 to 88.5, tier 2
  4. , interval 77.5 to 87.8, tier 2
  5. , interval 76.2 to 85.5, tier 2
  6. , interval 66.1 to 77.2, tier 3
  7. Jev61.3
    , interval 54.8 to 67.9, tier 4

Ahead aloneTied for the leadThe rest95% interval

One job at a time

An evaluator applies its criterion to a trace or a thread: pass or fail. A balanced sample of labelled cases from six demo agents, half true and half false.

LLM-as-a-judge, core: Accuracy, higher is better

  1. GPT-6.1 Sol92.6%, interval 92.6% [89.1, 95.8], tier 1
  2. Claude Opus 5.589.1%, interval 89.1% [83.9, 93.9], tier 1
  3. Claude Sonnet 5.583.5%, interval 83.5% [78.4, 88.5], tier 2
  4. GPT-6 Luna82.6%, interval 82.6% [77.5, 87.8], tier 2
  5. Gemini 3.8 Flash80.9%, interval 80.9% [76.2, 85.5], tier 2
  6. OpenAI Decisions API71.7%, interval 71.7% [66.1, 77.2], tier 3
  7. Jev61.3%, interval 61.3% [54.8, 67.9], tier 4
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 50%, not 0.
Plot

LLM-as-a-judge, core: accuracy (higher is better) vs cost

↖ Better
  • Jev
  • GPT-6.1 Sol
  • GPT-6 Luna
  • OpenAI Decisions API
← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑

Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).

Every model, every job

Every visible model on every task, sorted by LLM-as-a-judge, core. Value with half its interval; bands are tie tiers.
Model
Tier 1 on LLM-as-a-judge, core
GPT-6.1 Sol87.8±2.8 · T187.8% [84.8, 90.4], tier 192.7±2.8 · T192.7% [89.7, 95.4], tier 196.8±1.5 · T196.8% [95.2, 98.2], tier 192.6±3.4 · T192.6% [89.1, 95.8], tier 196.9±3.5 · T196.9% [93.0, 100.0], tier 198.5±1.7 · T198.5% [96.6, 100.0], tier 181.3±6.6 · T281.3% [74.1, 87.3], tier 286.5±6.4 · T286.5% [79.8, 92.5], tier 296.0±3.3 · T196.0% [92.4, 99.0], tier 189.6±3.8 · T189.6% [85.6, 93.1], tier 194.8±4.5 · T194.8% [90.0, 98.9], tier 196.0±2.6 · T196.0% [93.3, 98.5], tier 1
Claude Opus 5.589.2±2.9 · T189.2% [86.0, 91.8], tier 189.2±3.8 · T289.2% [85.2, 92.8], tier 297.8±1.2 · T197.8% [96.5, 98.8], tier 189.1±5.0 · T189.1% [83.9, 93.9], tier 192.7±6.0 · T192.7% [85.9, 97.9], tier 199.0±1.5 · T199.0% [97.0, 100.0], tier 186.8±6.1 · T186.8% [80.2, 92.3], tier 182.3±8.2 · T282.3% [74.0, 90.4], tier 297.0±2.2 · T197.0% [94.6, 99.0], tier 191.5±3.6 · T191.5% [87.6, 94.8], tier 192.7±4.8 · T192.7% [87.5, 97.1], tier 197.5±2.2 · T197.5% [95.0, 99.5], tier 1
Tier 2
Claude Sonnet 5.584.6±3.0 · T284.6% [81.5, 87.5], tier 288.9±3.1 · T288.9% [85.5, 91.8], tier 295.0±1.7 · T295.0% [93.2, 96.6], tier 283.5±5.0 · T283.5% [78.4, 88.5], tier 295.8±3.8 · T195.8% [91.3, 99.0], tier 199.0±1.3 · T199.0% [97.5, 100.0], tier 181.3±6.7 · T281.3% [74.1, 87.4], tier 277.1±7.1 · T377.1% [69.8, 84.0], tier 389.5±4.1 · T289.5% [85.2, 93.4], tier 289.1±3.2 · T189.1% [85.8, 92.1], tier 193.8±4.8 · T193.8% [88.3, 97.9], tier 196.5±2.5 · T196.5% [93.9, 99.0], tier 1
GPT-6 Luna79.2±3.2 · T379.2% [76.0, 82.3], tier 386.8±3.2 · T286.8% [83.5, 89.8], tier 295.0±1.7 · T295.0% [93.3, 96.6], tier 282.6±5.1 · T282.6% [77.5, 87.8], tier 293.8±4.7 · T193.8% [88.5, 97.9], tier 197.5±2.0 · T297.5% [95.5, 99.5], tier 279.9±6.7 · T279.9% [72.8, 86.2], tier 272.9±6.9 · T372.9% [66.0, 79.8], tier 393.5±3.3 · T293.5% [90.1, 96.7], tier 275.2±4.4 · T375.2% [70.7, 79.6], tier 393.8±4.6 · T193.8% [88.8, 97.9], tier 194.0±3.3 · T294.0% [90.5, 97.1], tier 2
Gemini 3.8 Flash83.1±2.8 · T283.1% [80.1, 85.8], tier 294.8±2.3 · T194.8% [92.2, 96.8], tier 197.3±1.2 · T197.3% [96.0, 98.4], tier 180.9±4.6 · T280.9% [76.2, 85.5], tier 295.8±3.7 · T195.8% [91.5, 99.0], tier 199.5±0.8 · T199.5% [98.4, 100.0], tier 183.3±6.2 · T183.3% [76.8, 89.1], tier 195.8±3.6 · T195.8% [91.8, 99.0], tier 195.5±2.6 · T195.5% [92.8, 97.9], tier 185.0±3.5 · T285.0% [81.3, 88.4], tier 292.7±4.7 · T192.7% [87.5, 97.0], tier 197.0±2.3 · T197.0% [94.4, 99.1], tier 1
Tier 3
OpenAI Decisions API75.2±3.3 · T475.2% [71.8, 78.4], tier 470.8±4.5 · T370.8% [66.3, 75.3], tier 375.0±3.4 · T375.0% [71.5, 78.3], tier 371.7±5.6 · T371.7% [66.1, 77.2], tier 371.9±8.7 · T271.9% [62.8, 80.2], tier 277.0±5.5 · T377.0% [71.3, 82.2], tier 383.3±6.9 · T183.3% [76.1, 89.9], tier 162.5±7.6 · T462.5% [55.1, 70.2], tier 473.5±6.4 · T373.5% [67.0, 79.9], tier 370.4±4.4 · T370.4% [66.0, 74.8], tier 378.1±6.9 · T278.1% [71.3, 85.1], tier 274.5±5.9 · T374.5% [68.4, 80.2], tier 3
Tier 4
Jev63.7±4.1 · T563.7% [59.6, 67.7], tier 572.6±4.8 · T372.6% [67.7, 77.3], tier 372.7±3.5 · T372.7% [69.1, 76.1], tier 361.3±6.6 · T461.3% [54.8, 67.9], tier 471.9±7.9 · T271.9% [63.8, 79.6], tier 272.5±5.8 · T372.5% [66.7, 78.3], tier 366.7±9.0 · T366.7% [57.3, 75.3], tier 372.9±7.9 · T372.9% [65.0, 80.8], tier 376.5±5.8 · T376.5% [70.4, 82.1], tier 363.3±4.9 · T463.3% [58.2, 68.1], tier 472.9±9.0 · T272.9% [63.7, 81.6], tier 269.0±6.6 · T369.0% [62.3, 75.5], tier 3

Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.

How we measured

Scores carry a 95% bootstrap interval (2,000 resamples, the versions of one trace or scenario resampled together). Two models are tied when the 95% interval of their paired difference on the same cases holds zero.

Tie tiers

Models in one tier can't be separated by the tie test (the 95% interval of the paired difference on the same cases holds zero). Tiers are computed in the export over every model; filters on this page never change them.

Contamination

The cases are labelled traces from LangWatch’s own demo agents. They were never published, so no model can have trained on them. The page shows aggregates only: no case and no case text.

Latency and hardware

  • Hosted API: Jev, Gemini 3.8 Flash, GPT-6 Luna and OpenAI Decisions API, one API call per judgment, several at a time
  • Vendor CLI: GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5, through their vendors' command-line tools; each call's time includes the tool's start-up

Latencies are only compared within one hardware tier.

Datasets

  • All three jobs, core (834 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • All three jobs, long conversations (288 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • All three jobs, hard cases (600 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • LLM-as-a-judge, core (230 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • LLM-as-a-judge, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • LLM-as-a-judge, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Scenario judge, core (144 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Scenario judge, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Scenario judge, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Search, core (460 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Search, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Search, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • A refusal, or a reply with no verdict in it, counts as wrong, as it would for a customer.
  • Cost and latency were measured on the core set only, one call per judgment.
  • All three jobs: the mean of the three job scores on the set, each job counting the same; its interval and the tie test combine the jobs’ intervals, and its cost is the mean of the three job costs.
  • Jev, Gemini 3.8 Flash, GPT-6 Luna and OpenAI Decisions API: cost recorded on the run. GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5 ran on subscriptions, so their cost is an estimate (~) at the API list price; priced on the tokens the CLI reported, it would be up to $188 per 1,000 judgments.
  • GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5 ran through their vendors' command-line tools, so each call's time includes the tool's start-up. Their latency is shown apart from the API calls (Jev, Gemini 3.8 Flash, GPT-6 Luna and OpenAI Decisions API) and never compared with them.
  • GPT-6.1 Sol: Run through the Codex CLI on a ChatGPT subscription, reasoning effort low. Codex adds its own system prompt in front of ours, and the forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price.
  • Claude Opus 5.5: Run through Claude Code on a Claude subscription. The forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price. A refusal by the CLI’s safeguards counts as wrong; the API may refuse less often.
  • Claude Sonnet 5.5: Run through Claude Code on a Claude subscription. The forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price. A refusal by the CLI’s safeguards counts as wrong; the API may refuse less often.
  • OpenAI Decisions API: Called through OpenAI’s Decisions API, in public beta, which runs GPT-6 Luna. It gets the typed question Jev gets, over the same text Jev reads, and answers with a probability for each answer instead of text. Billed on input tokens only.
  • Gemini 3.8 Flash: Metered at the introductory price, which runs until 31 December 2026; the list price roughly doubles after that.
  • We also re-ran every judge with the richest input that fits its context window; no LLM moved by more than 8.3 points, but a pick changed.

Release production-2026-10-07, integrity sha256-leaves-v1.