No single model judges best

GPT-6.1 Sol, Gemini 3.8 Flash and Claude Sonnet 5.5 each take one job, each tied with Claude Opus 5.5 at about half its cost or less. Jev trails by 17 to 31 points. A refusal counts as wrong.

This benchmark was created using LangWatch. Sign up to create your own · Read the method

LangWatch Instant Evals runs on Jev. Check it out

Task
Set

LLM-as-a-judge, core

Accuracy, % · Higher is better

Pick: GPT-6.1 Sol, tied with Claude Opus 5.5 at about half its cost.

  1. , interval 89.1 to 95.8, tied for the lead
  2. , interval 83.9 to 93.9, tied for the lead
  3. , interval 78.4 to 88.5, tier 2
  4. , interval 76.2 to 85.5, tier 2
  5. Jev61.3
    , interval 54.8 to 67.9, tier 3

Ahead aloneTied for the leadThe rest95% interval

One job at a time

An evaluator applies its criterion to a trace or a thread: pass or fail. A balanced sample of labelled cases from six demo agents, half true and half false.

LLM-as-a-judge, core: Accuracy, higher is better

  1. GPT-6.1 Sol92.6%, interval 92.6% [89.1, 95.8], tier 1
  2. Claude Opus 5.589.1%, interval 89.1% [83.9, 93.9], tier 1
  3. Claude Sonnet 5.583.5%, interval 83.5% [78.4, 88.5], tier 2
  4. Gemini 3.8 Flash80.9%, interval 80.9% [76.2, 85.5], tier 2
  5. Jev61.3%, interval 61.3% [54.8, 67.9], tier 3
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 50%, not 0.
Plot

LLM-as-a-judge, core: accuracy (higher is better) vs cost

↖ Better
  • Jev
  • Gemini 3.8 Flash
  • GPT-6.1 Sol
← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑

Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).

Every model, every job

Every visible model on every task, sorted by LLM-as-a-judge, core. Value with half its interval; bands are tie tiers.
Model
Tier 1 on LLM-as-a-judge, core
GPT-6.1 Sol92.6±3.4 · T192.6% [89.1, 95.8], tier 196.9±3.5 · T196.9% [93.0, 100.0], tier 198.5±1.7 · T198.5% [96.6, 100.0], tier 181.3±6.6 · T281.3% [74.1, 87.3], tier 286.5±6.4 · T286.5% [79.8, 92.5], tier 296.0±3.3 · T196.0% [92.4, 99.0], tier 189.6±3.8 · T189.6% [85.6, 93.1], tier 194.8±4.5 · T194.8% [90.0, 98.9], tier 196.0±2.6 · T196.0% [93.3, 98.5], tier 1
Claude Opus 5.589.1±5.0 · T189.1% [83.9, 93.9], tier 192.7±6.0 · T192.7% [85.9, 97.9], tier 199.0±1.5 · T199.0% [97.0, 100.0], tier 186.8±6.1 · T186.8% [80.2, 92.3], tier 182.3±8.2 · T282.3% [74.0, 90.4], tier 297.0±2.2 · T197.0% [94.6, 99.0], tier 191.5±3.6 · T191.5% [87.6, 94.8], tier 192.7±4.8 · T192.7% [87.5, 97.1], tier 197.5±2.2 · T197.5% [95.0, 99.5], tier 1
Tier 2
Claude Sonnet 5.583.5±5.0 · T283.5% [78.4, 88.5], tier 295.8±3.8 · T195.8% [91.3, 99.0], tier 199.0±1.3 · T199.0% [97.5, 100.0], tier 181.3±6.7 · T281.3% [74.1, 87.4], tier 277.1±7.1 · T377.1% [69.8, 84.0], tier 389.5±4.1 · T289.5% [85.2, 93.4], tier 289.1±3.2 · T189.1% [85.8, 92.1], tier 193.8±4.8 · T193.8% [88.3, 97.9], tier 196.5±2.5 · T196.5% [93.9, 99.0], tier 1
Gemini 3.8 Flash80.9±4.6 · T280.9% [76.2, 85.5], tier 295.8±3.7 · T195.8% [91.5, 99.0], tier 199.5±0.8 · T199.5% [98.4, 100.0], tier 183.3±6.2 · T183.3% [76.8, 89.1], tier 195.8±3.6 · T195.8% [91.8, 99.0], tier 195.5±2.6 · T195.5% [92.8, 97.9], tier 185.0±3.5 · T285.0% [81.3, 88.4], tier 292.7±4.7 · T192.7% [87.5, 97.0], tier 197.0±2.3 · T197.0% [94.4, 99.1], tier 1
Tier 3
Jev61.3±6.6 · T361.3% [54.8, 67.9], tier 371.9±7.9 · T271.9% [63.8, 79.6], tier 272.5±5.8 · T272.5% [66.7, 78.3], tier 266.7±9.0 · T366.7% [57.3, 75.3], tier 372.9±7.9 · T372.9% [65.0, 80.8], tier 376.5±5.8 · T376.5% [70.4, 82.1], tier 363.3±4.9 · T363.3% [58.2, 68.1], tier 372.9±9.0 · T272.9% [63.7, 81.6], tier 269.0±6.6 · T269.0% [62.3, 75.5], tier 2

Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.

How we measured

Scores carry a 95% bootstrap interval (2,000 resamples, the versions of one trace or scenario resampled together). Two models are tied when the 95% interval of their paired difference on the same cases holds zero.

Tie tiers

Models in one tier can't be separated by the tie test (the 95% interval of the paired difference on the same cases holds zero). Tiers are computed in the export over every model; filters on this page never change them.

Contamination

The cases are labelled traces from LangWatch’s own demo agents. They were never published, so no model can have trained on them. The page shows aggregates only: no case and no case text.

Latency and hardware

  • Hosted API: Jev and Gemini 3.8 Flash, one API call per judgment, several at a time
  • Vendor CLI: GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5, through their vendors' command-line tools; each call's time includes the tool's start-up

Latencies are only compared within one hardware tier.

Datasets

  • LLM-as-a-judge, core (230 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • LLM-as-a-judge, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • LLM-as-a-judge, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Scenario judge, core (144 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Scenario judge, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Scenario judge, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Search, core (460 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Search, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • Search, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
  • A refusal, or a reply with no verdict in it, counts as wrong, as it would for a customer.
  • Cost and latency were measured on the core set only, one call per judgment.
  • Jev and Gemini 3.8 Flash: cost recorded on the run. GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5 ran on subscriptions, so their cost is an estimate (~) at the API list price; priced on the tokens the CLI reported, it would be up to $188 per 1,000 judgments.
  • GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5 ran through their vendors' command-line tools, so each call's time includes the tool's start-up. Their latency is shown apart from the API calls (Jev and Gemini 3.8 Flash) and never compared with them.
  • GPT-6.1 Sol: Run through the Codex CLI on a ChatGPT subscription, reasoning effort low. Codex adds its own system prompt in front of ours, and the forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price.
  • Claude Opus 5.5: Run through Claude Code on a Claude subscription. The forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price. A refusal by the CLI’s safeguards counts as wrong; the API may refuse less often.
  • Claude Sonnet 5.5: Run through Claude Code on a Claude subscription. The forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price. A refusal by the CLI’s safeguards counts as wrong; the API may refuse less often.
  • Gemini 3.8 Flash: Metered at the introductory price, which runs until 31 December 2026; the list price roughly doubles after that.
  • We also re-ran every judge with the richest input that fits its context window; no LLM moved by more than 2.1 points and no pick changed.

Release production-2026-10-04, integrity sha256-leaves-v1.