Frontier LLMs are up to 25 points more accurate than Jev, but cost over 100x more
Every step up the Pareto frontier buys accuracy at a much higher price. Run a cheap judge on every trace, and save the frontier models for the cases where the extra points matter.
This benchmark was created using LangWatch. Sign up to create your own · Read the method
LangWatch Instant Evals runs on Jev. Check it out
LLM-as-a-judge, core
Accuracy, % · Higher is better
Pick: GPT-6.1 Sol, tied with Claude Opus 5.5 at about half its cost.
- GPT-6.1 Sol92.6, interval 89.1 to 95.8, tied for the lead
- Claude Opus 5.589.1, interval 83.9 to 93.9, tied for the lead
- , interval 78.4 to 88.5, tier 2
- GPT-6 Luna82.6, interval 77.5 to 87.8, tier 2
- Gemini 3.8 Flash80.9, interval 76.2 to 85.5, tier 2
- , interval 66.1 to 77.2, tier 3
- Jev61.3, interval 54.8 to 67.9, tier 4
Ahead aloneTied for the leadThe rest95% interval
One job at a time
An evaluator applies its criterion to a trace or a thread: pass or fail. A balanced sample of labelled cases from six demo agents, half true and half false.
LLM-as-a-judge, core: Accuracy, higher is better
- GPT-6.1 Sol92.6%, interval 92.6% [89.1, 95.8], tier 1
- Claude Opus 5.589.1%, interval 89.1% [83.9, 93.9], tier 1
- Claude Sonnet 5.583.5%, interval 83.5% [78.4, 88.5], tier 2
- GPT-6 Luna82.6%, interval 82.6% [77.5, 87.8], tier 2
- Gemini 3.8 Flash80.9%, interval 80.9% [76.2, 85.5], tier 2
- OpenAI Decisions API71.7%, interval 71.7% [66.1, 77.2], tier 3
- Jev61.3%, interval 61.3% [54.8, 67.9], tier 4
LLM-as-a-judge, core: accuracy (higher is better) vs cost
- Jev
- GPT-6.1 Sol
- GPT-6 Luna
- OpenAI Decisions API
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Every model, every job
| Model | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Tier 1 on LLM-as-a-judge, core | ||||||||||||
| GPT-6.1 Sol | 87.8±2.8 · T187.8% [84.8, 90.4], tier 1 | 92.7±2.8 · T192.7% [89.7, 95.4], tier 1 | 96.8±1.5 · T196.8% [95.2, 98.2], tier 1 | 92.6±3.4 · T192.6% [89.1, 95.8], tier 1 | 96.9±3.5 · T196.9% [93.0, 100.0], tier 1 | 98.5±1.7 · T198.5% [96.6, 100.0], tier 1 | 81.3±6.6 · T281.3% [74.1, 87.3], tier 2 | 86.5±6.4 · T286.5% [79.8, 92.5], tier 2 | 96.0±3.3 · T196.0% [92.4, 99.0], tier 1 | 89.6±3.8 · T189.6% [85.6, 93.1], tier 1 | 94.8±4.5 · T194.8% [90.0, 98.9], tier 1 | 96.0±2.6 · T196.0% [93.3, 98.5], tier 1 |
| Claude Opus 5.5 | 89.2±2.9 · T189.2% [86.0, 91.8], tier 1 | 89.2±3.8 · T289.2% [85.2, 92.8], tier 2 | 97.8±1.2 · T197.8% [96.5, 98.8], tier 1 | 89.1±5.0 · T189.1% [83.9, 93.9], tier 1 | 92.7±6.0 · T192.7% [85.9, 97.9], tier 1 | 99.0±1.5 · T199.0% [97.0, 100.0], tier 1 | 86.8±6.1 · T186.8% [80.2, 92.3], tier 1 | 82.3±8.2 · T282.3% [74.0, 90.4], tier 2 | 97.0±2.2 · T197.0% [94.6, 99.0], tier 1 | 91.5±3.6 · T191.5% [87.6, 94.8], tier 1 | 92.7±4.8 · T192.7% [87.5, 97.1], tier 1 | 97.5±2.2 · T197.5% [95.0, 99.5], tier 1 |
| Tier 2 | ||||||||||||
| Claude Sonnet 5.5 | 84.6±3.0 · T284.6% [81.5, 87.5], tier 2 | 88.9±3.1 · T288.9% [85.5, 91.8], tier 2 | 95.0±1.7 · T295.0% [93.2, 96.6], tier 2 | 83.5±5.0 · T283.5% [78.4, 88.5], tier 2 | 95.8±3.8 · T195.8% [91.3, 99.0], tier 1 | 99.0±1.3 · T199.0% [97.5, 100.0], tier 1 | 81.3±6.7 · T281.3% [74.1, 87.4], tier 2 | 77.1±7.1 · T377.1% [69.8, 84.0], tier 3 | 89.5±4.1 · T289.5% [85.2, 93.4], tier 2 | 89.1±3.2 · T189.1% [85.8, 92.1], tier 1 | 93.8±4.8 · T193.8% [88.3, 97.9], tier 1 | 96.5±2.5 · T196.5% [93.9, 99.0], tier 1 |
| GPT-6 Luna | 79.2±3.2 · T379.2% [76.0, 82.3], tier 3 | 86.8±3.2 · T286.8% [83.5, 89.8], tier 2 | 95.0±1.7 · T295.0% [93.3, 96.6], tier 2 | 82.6±5.1 · T282.6% [77.5, 87.8], tier 2 | 93.8±4.7 · T193.8% [88.5, 97.9], tier 1 | 97.5±2.0 · T297.5% [95.5, 99.5], tier 2 | 79.9±6.7 · T279.9% [72.8, 86.2], tier 2 | 72.9±6.9 · T372.9% [66.0, 79.8], tier 3 | 93.5±3.3 · T293.5% [90.1, 96.7], tier 2 | 75.2±4.4 · T375.2% [70.7, 79.6], tier 3 | 93.8±4.6 · T193.8% [88.8, 97.9], tier 1 | 94.0±3.3 · T294.0% [90.5, 97.1], tier 2 |
| Gemini 3.8 Flash | 83.1±2.8 · T283.1% [80.1, 85.8], tier 2 | 94.8±2.3 · T194.8% [92.2, 96.8], tier 1 | 97.3±1.2 · T197.3% [96.0, 98.4], tier 1 | 80.9±4.6 · T280.9% [76.2, 85.5], tier 2 | 95.8±3.7 · T195.8% [91.5, 99.0], tier 1 | 99.5±0.8 · T199.5% [98.4, 100.0], tier 1 | 83.3±6.2 · T183.3% [76.8, 89.1], tier 1 | 95.8±3.6 · T195.8% [91.8, 99.0], tier 1 | 95.5±2.6 · T195.5% [92.8, 97.9], tier 1 | 85.0±3.5 · T285.0% [81.3, 88.4], tier 2 | 92.7±4.7 · T192.7% [87.5, 97.0], tier 1 | 97.0±2.3 · T197.0% [94.4, 99.1], tier 1 |
| Tier 3 | ||||||||||||
| OpenAI Decisions API | 75.2±3.3 · T475.2% [71.8, 78.4], tier 4 | 70.8±4.5 · T370.8% [66.3, 75.3], tier 3 | 75.0±3.4 · T375.0% [71.5, 78.3], tier 3 | 71.7±5.6 · T371.7% [66.1, 77.2], tier 3 | 71.9±8.7 · T271.9% [62.8, 80.2], tier 2 | 77.0±5.5 · T377.0% [71.3, 82.2], tier 3 | 83.3±6.9 · T183.3% [76.1, 89.9], tier 1 | 62.5±7.6 · T462.5% [55.1, 70.2], tier 4 | 73.5±6.4 · T373.5% [67.0, 79.9], tier 3 | 70.4±4.4 · T370.4% [66.0, 74.8], tier 3 | 78.1±6.9 · T278.1% [71.3, 85.1], tier 2 | 74.5±5.9 · T374.5% [68.4, 80.2], tier 3 |
| Tier 4 | ||||||||||||
| Jev | 63.7±4.1 · T563.7% [59.6, 67.7], tier 5 | 72.6±4.8 · T372.6% [67.7, 77.3], tier 3 | 72.7±3.5 · T372.7% [69.1, 76.1], tier 3 | 61.3±6.6 · T461.3% [54.8, 67.9], tier 4 | 71.9±7.9 · T271.9% [63.8, 79.6], tier 2 | 72.5±5.8 · T372.5% [66.7, 78.3], tier 3 | 66.7±9.0 · T366.7% [57.3, 75.3], tier 3 | 72.9±7.9 · T372.9% [65.0, 80.8], tier 3 | 76.5±5.8 · T376.5% [70.4, 82.1], tier 3 | 63.3±4.9 · T463.3% [58.2, 68.1], tier 4 | 72.9±9.0 · T272.9% [63.7, 81.6], tier 2 | 69.0±6.6 · T369.0% [62.3, 75.5], tier 3 |
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
How we measured
Scores carry a 95% bootstrap interval (2,000 resamples, the versions of one trace or scenario resampled together). Two models are tied when the 95% interval of their paired difference on the same cases holds zero.
Tie tiers
Models in one tier can't be separated by the tie test (the 95% interval of the paired difference on the same cases holds zero). Tiers are computed in the export over every model; filters on this page never change them.
Contamination
The cases are labelled traces from LangWatch’s own demo agents. They were never published, so no model can have trained on them. The page shows aggregates only: no case and no case text.
Latency and hardware
- Hosted API: Jev, Gemini 3.8 Flash, GPT-6 Luna and OpenAI Decisions API, one API call per judgment, several at a time
- Vendor CLI: GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5, through their vendors' command-line tools; each call's time includes the tool's start-up
Latencies are only compared within one hardware tier.
Datasets
- All three jobs, core (834 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- All three jobs, long conversations (288 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- All three jobs, hard cases (600 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- LLM-as-a-judge, core (230 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- LLM-as-a-judge, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- LLM-as-a-judge, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- Scenario judge, core (144 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- Scenario judge, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- Scenario judge, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- Search, core (460 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- Search, long conversations (96 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- Search, hard cases (200 test items): labelled traces from LangWatch demo agents (private, aggregates only)
- A refusal, or a reply with no verdict in it, counts as wrong, as it would for a customer.
- Cost and latency were measured on the core set only, one call per judgment.
- All three jobs: the mean of the three job scores on the set, each job counting the same; its interval and the tie test combine the jobs’ intervals, and its cost is the mean of the three job costs.
- Jev, Gemini 3.8 Flash, GPT-6 Luna and OpenAI Decisions API: cost recorded on the run. GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5 ran on subscriptions, so their cost is an estimate (~) at the API list price; priced on the tokens the CLI reported, it would be up to $188 per 1,000 judgments.
- GPT-6.1 Sol, Claude Opus 5.5 and Claude Sonnet 5.5 ran through their vendors' command-line tools, so each call's time includes the tool's start-up. Their latency is shown apart from the API calls (Jev, Gemini 3.8 Flash, GPT-6 Luna and OpenAI Decisions API) and never compared with them.
- GPT-6.1 Sol: Run through the Codex CLI on a ChatGPT subscription, reasoning effort low. Codex adds its own system prompt in front of ours, and the forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price.
- Claude Opus 5.5: Run through Claude Code on a Claude subscription. The forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price. A refusal by the CLI’s safeguards counts as wrong; the API may refuse less often.
- Claude Sonnet 5.5: Run through Claude Code on a Claude subscription. The forced function call becomes a JSON reply. Latency includes the CLI start‑up; cost is an estimate at the API list price. A refusal by the CLI’s safeguards counts as wrong; the API may refuse less often.
- OpenAI Decisions API: Called through OpenAI’s Decisions API, in public beta, which runs GPT-6 Luna. It gets the typed question Jev gets, over the same text Jev reads, and answers with a probability for each answer instead of text. Billed on input tokens only.
- Gemini 3.8 Flash: Metered at the introductory price, which runs until 31 December 2026; the list price roughly doubles after that.
- We also re-ran every judge with the richest input that fits its context window; no LLM moved by more than 8.3 points, but a pick changed.
Release production-2026-10-07, integrity sha256-leaves-v1.
