Open Models now match Jev accuracy, but Jev is still cheaper

LangWatch LLM-as-a-Judge Bench evaluates model capabilities in doing evals, agent simulations and trace search, our specialty. The examples are real-world based and never public to avoid benchmaxing and measure true capability for our customers.

This benchmark was created using LangWatch. Sign up to create your own · Read the method

LangWatch Instant Evals runs on Jev. Check it out

Task

All three jobs

Accuracy, % · Higher is better

Decision 2.0 Vega 27B leads Jev by 3.2 points, too little to be sure across 20 comparisons with Jev: the 95% interval of the gap runs from 0.2 to 6.1 points.

  1. , interval 72.5 to 79.4, tied for the lead
  2. , interval 71.7 to 78.5, tied for the lead
  3. Clef74.5
    , interval 70.9 to 78.2, tied for the lead
  4. GLiDE74.3
    , interval 70.9 to 78.0, tied for the lead
  5. , interval 69.7 to 77.2, tied for the lead
  6. , interval 69.4 to 76.9, tied for the lead
  7. Jev72.7
    , interval 69.2 to 76.3, tied for the lead
  8. , interval 68.5 to 75.4, tied for the lead

Ahead aloneTied for the leadThe rest95% interval

One job at a time

All three jobs: Accuracy, higher is better

  1. Decision 2.0 Vega 27B75.8%, interval 75.8% [72.5, 79.4], tier 1
  2. OpenAI Decisions API75.0%, interval 75.0% [71.7, 78.5], tier 1
  3. Clef74.5%, interval 74.5% [70.9, 78.2], tier 1
  4. GLiDE74.3%, interval 74.3% [70.9, 78.0], tier 1
  5. Jev-Omni73.3%, interval 73.3% [69.7, 77.2], tier 1
  6. Winnow73.0%, interval 73.0% [69.4, 76.9], tier 1
  7. Jev72.7%, interval 72.7% [69.2, 76.3], tier 1
  8. Kev-27B72.0%, interval 72.0% [68.5, 75.4], tier 1
  9. DiffusionGemma-Jev69.5%, interval 69.5% [65.7, 73.3], tier 2
  10. Cygnet65.8%, interval 65.8% [62.0, 69.8], tier 2
  11. Hopper 12B65.8%, interval 65.8% [62.0, 69.8], tier 2
  12. Decision 2.0 Lux 9B61.0%, interval 61.0% [57.2, 65.0], tier 3
  13. Clef Flash58.7%, interval 58.7% [54.7, 62.9], tier 3
  14. decider-4b v258.3%, interval 58.3% [54.5, 62.0], tier 3
  15. JevK557.7%, interval 57.7% [54.1, 61.4], tier 3
  16. Decision 2.0 Nox 4B56.3%, interval 56.3% [52.6, 60.3], tier 4
  17. Lev 4B52.7%, interval 52.7% [49.0, 56.3], tier 5
  18. Julia-150.0%, interval 50.0% [46.4, 53.5], tier 5
  19. Decision 2.0 Eos 0.8B49.7%, interval 49.7% [45.9, 53.4], tier 5
  20. Decision 2.0 Sol 2B48.5%, interval 48.5% [44.7, 52.3], tier 6
  21. Decision 2.0 Kai 0.6B48.3%, interval 48.3% [44.3, 52.3], tier 6
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 40%, not 0.
Plot

All three jobs: accuracy (higher is better) vs cost

↖ Better
  • Jev
  • OpenAI Decisions API
← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑

Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Cygnet, decider-4b v2, Decision 2.0 Eos 0.8B, Decision 2.0 Kai 0.6B, Decision 2.0 Lux 9B, Decision 2.0 Nox 4B, Decision 2.0 Sol 2B, Decision 2.0 Vega 27B, DiffusionGemma-Jev, Hopper 12B, Jev-Omni, JevK5, Julia-1, Kev-27B, Lev 4B, Winnow.

Jev alternatives, every job

Every visible model on every task, sorted by All three jobs. Value with half its interval; bands are tie tiers.
Model
Tier 1 on All three jobs
Decision 2.0 Vega 27B75.8±3.5 · T175.8% [72.5, 79.4], tier 175.5±5.7 · T175.5% [69.8, 81.3], tier 171.0±6.2 · T171.0% [64.5, 77.0], tier 181.0±5.3 · T181.0% [75.5, 86.1], tier 1
OpenAI Decisions API75.0±3.4 · T175.0% [71.7, 78.5], tier 177.0±5.5 · T177.0% [71.3, 82.2], tier 174.5±5.9 · T174.5% [68.4, 80.2], tier 173.5±6.4 · T273.5% [67.0, 79.9], tier 2
Clef74.5±3.6 · T174.5% [70.9, 78.2], tier 173.0±6.8 · T173.0% [66.2, 79.7], tier 167.5±5.9 · T167.5% [61.3, 73.1], tier 183.0±5.5 · T183.0% [77.4, 88.4], tier 1
GLiDE74.3±3.5 · T174.3% [70.9, 78.0], tier 172.0±6.2 · T172.0% [65.7, 78.1], tier 171.5±6.7 · T171.5% [64.4, 77.8], tier 179.5±5.1 · T179.5% [74.4, 84.7], tier 1
Jev-Omni73.3±3.8 · T173.3% [69.7, 77.2], tier 172.0±6.5 · T172.0% [65.5, 78.5], tier 172.5±6.5 · T172.5% [65.7, 78.7], tier 175.5±6.2 · T175.5% [69.3, 81.7], tier 1
Winnow73.0±3.8 · T173.0% [69.4, 76.9], tier 170.5±6.2 · T170.5% [64.0, 76.3], tier 169.5±6.2 · T169.5% [63.2, 75.6], tier 179.0±5.9 · T179.0% [73.1, 85.0], tier 1
Jev72.7±3.5 · T172.7% [69.2, 76.3], tier 172.5±5.8 · T172.5% [66.7, 78.3], tier 169.0±6.6 · T169.0% [62.3, 75.5], tier 176.5±5.8 · T176.5% [70.4, 82.1], tier 1
Kev-27B72.0±3.4 · T172.0% [68.5, 75.4], tier 172.0±6.4 · T172.0% [65.6, 78.4], tier 166.0±6.0 · T166.0% [59.7, 71.8], tier 178.0±6.2 · T178.0% [71.6, 84.0], tier 1
Tier 2
DiffusionGemma-Jev69.5±3.8 · T269.5% [65.7, 73.3], tier 270.5±6.8 · T170.5% [63.8, 77.4], tier 163.0±6.6 · T263.0% [56.3, 69.5], tier 275.0±5.6 · T275.0% [69.5, 80.6], tier 2
Cygnet65.8±3.9 · T265.8% [62.0, 69.8], tier 260.0±6.9 · T260.0% [52.8, 66.7], tier 262.5±6.0 · T262.5% [56.5, 68.6], tier 275.0±6.5 · T275.0% [68.5, 81.3], tier 2
Hopper 12B65.8±3.9 · T265.8% [62.0, 69.8], tier 265.0±6.9 · T265.0% [57.8, 71.6], tier 257.0±6.3 · T257.0% [50.5, 63.2], tier 275.5±6.0 · T175.5% [69.5, 81.5], tier 1
Tier 3
Decision 2.0 Lux 9B61.0±3.9 · T361.0% [57.2, 65.0], tier 359.0±6.9 · T259.0% [52.0, 65.8], tier 259.5±6.6 · T259.5% [53.1, 66.2], tier 264.5±6.8 · T364.5% [57.4, 71.0], tier 3
Clef Flash58.7±4.1 · T358.7% [54.7, 62.9], tier 359.5±7.1 · T259.5% [52.3, 66.5], tier 256.5±6.3 · T356.5% [50.3, 62.9], tier 360.0±6.9 · T360.0% [53.0, 66.8], tier 3
decider-4b v258.3±3.8 · T358.3% [54.5, 62.0], tier 359.5±6.9 · T259.5% [52.7, 66.5], tier 254.5±6.5 · T354.5% [48.0, 60.9], tier 361.0±6.1 · T361.0% [54.7, 67.0], tier 3
JevK557.7±3.6 · T357.7% [54.1, 61.4], tier 355.0±7.0 · T355.0% [47.7, 61.8], tier 354.5±5.9 · T354.5% [48.5, 60.2], tier 363.5±6.0 · T363.5% [57.6, 69.7], tier 3
Tier 4
Decision 2.0 Nox 4B56.3±3.8 · T456.3% [52.6, 60.3], tier 458.5±6.8 · T258.5% [51.7, 65.4], tier 254.5±6.3 · T354.5% [48.3, 61.0], tier 356.0±6.4 · T456.0% [49.5, 62.4], tier 4
Tier 5
Lev 4B52.7±3.6 · T552.7% [49.0, 56.3], tier 550.5±6.3 · T350.5% [44.3, 56.9], tier 353.0±6.2 · T353.0% [46.9, 59.3], tier 354.5±6.4 · T454.5% [48.0, 60.9], tier 4
Julia-150.0±3.6 · T550.0% [46.4, 53.5], tier 553.0±5.6 · T353.0% [47.7, 58.8], tier 349.5±6.3 · T349.5% [43.4, 56.1], tier 347.5±6.9 · T547.5% [40.5, 54.3], tier 5
Decision 2.0 Eos 0.8B49.7±3.7 · T549.7% [45.9, 53.4], tier 551.5±6.4 · T351.5% [45.1, 57.8], tier 348.0±6.5 · T348.0% [41.5, 54.5], tier 349.5±6.1 · T549.5% [43.4, 55.6], tier 5
Tier 6
Decision 2.0 Sol 2B48.5±3.8 · T648.5% [44.7, 52.3], tier 651.0±6.5 · T351.0% [44.5, 57.5], tier 346.5±6.9 · T446.5% [39.8, 53.6], tier 448.0±6.4 · T548.0% [41.7, 54.5], tier 5
Decision 2.0 Kai 0.6B48.3±4.0 · T648.3% [44.3, 52.3], tier 650.0±7.0 · T350.0% [43.1, 57.1], tier 345.0±6.6 · T445.0% [38.5, 51.6], tier 450.0±7.1 · T450.0% [42.9, 57.2], tier 4

Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. All three jobs: every check of the three jobs pooled, one accuracy over 600 checks.

Jev

Provider
TypeSafe
Weights
Closed (API)
Size
Not disclosed
Licence
Proprietary
TaskScore [95% interval]TierLatency p50Cost / 1k
All three jobs72.7% [69.2, 76.3]1, tied269 ms (Hosted API)$0.086
LLM-as-a-judge72.5% [66.7, 78.3]1, tied258 ms (Hosted API)$0.11
Search69.0% [62.3, 75.5]1, tied256 ms (Hosted API)$0.11
Scenario judge76.5% [70.4, 82.1]1, tied325 ms (Hosted API)$0.040
    How we measured

    Scores carry a 95% bootstrap interval (2,000 resamples, the checks built from one scenario resampled together). Two models are tied when the 95% interval of their paired difference on the same checks holds zero.

    Tie tiers

    Models in one tier can't be separated by the tie test (the 95% interval of the paired difference on the same cases holds zero). Tiers are computed in the export over every model; filters on this page never change them.

    Contamination

    The checks are synthetic, generated for this benchmark from made-up fixture data, with every label computed from that data. They were never published, so no model can have trained on them. The page shows aggregates only: no check and no check text.

    Latency and hardware

    • Hosted API: Jev, Clef, Clef Flash, GLiDE and OpenAI Decisions API, called as hosted APIs from a laptop, one call per check; the time includes the network
    • H100: Cygnet, Decision 2.0 Lux 9B, Decision 2.0 Vega 27B, DiffusionGemma-Jev, Hopper 12B, Jev-Omni and Kev-27B, open weights on one rented H100; the server’s own time per check, 30 checks one at a time
    • L4: decider-4b v2, Decision 2.0 Eos 0.8B, Decision 2.0 Kai 0.6B, Decision 2.0 Nox 4B, Decision 2.0 Sol 2B, JevK5, Lev 4B and Winnow, open weights on one rented L4; the server’s own time per check, 30 checks one at a time
    • Server CPU: Julia-1, served on the GPU box’s CPU; the server’s own time per check, 30 checks one at a time

    Latencies are only compared within one hardware tier.

    Datasets

    • All three jobs (600 test items): synthetic lookup and arithmetic checks, labels computed from the fixture data (private, aggregates only)
    • LLM-as-a-judge (200 test items): synthetic lookup and arithmetic checks, labels computed from the fixture data (private, aggregates only)
    • Search (200 test items): synthetic lookup and arithmetic checks, labels computed from the fixture data (private, aggregates only)
    • Scenario judge (200 test items): synthetic lookup and arithmetic checks, labels computed from the fixture data (private, aggregates only)
    • An error, a reply with no verdict, or an inconclusive Scenario verdict counts as wrong.
    • All three jobs: every check of the three jobs pooled, one accuracy over 600 checks.
    • Latency is the median time per check, compared only within one hardware tier. Hosted APIs were timed from a laptop on every check, so their time includes the network. Open-weights models were timed on the GPU box by the server itself, 30 checks one at a time, one number for every job.
    • The 15 open-weights models on the GPU ran in bf16 on a single rented L4 or H100 (Winnow in Q8_0 GGUF). Julia-1 ran in F32 on the box’s CPU.
    • GLiDE: A hosted “thinking” decision model: an uncertain check can take extra time.
    • Cygnet: A recipe, not a trained model: a letter readout on stock Gemma 4 12B IT. Run without its calibration temperature, which rescales probabilities but cannot flip a yes or no at the 0.5 threshold.
    • decider-4b v2: Served with a smaller CUDA graph token budget than its default, a capacity setting that does not change answers.
    • JevK5: Takes an 8k token state; no check was longer.
    • Julia-1: Served on the box’s CPU, so its latency is CPU time.
    • Lev 4B: Takes an 8k token state; no check was longer.
    • Winnow: Run from its Q8_0 GGUF weights, the precision its authors recommend.
    • Jev, Clef, Clef Flash, GLiDE and OpenAI Decisions API: cost per 1,000 checks as recorded on the run, at the provider’s price. Cygnet, decider-4b v2, Decision 2.0 Eos 0.8B, Decision 2.0 Kai 0.6B, Decision 2.0 Lux 9B, Decision 2.0 Nox 4B, Decision 2.0 Sol 2B, Decision 2.0 Vega 27B, DiffusionGemma-Jev, Hopper 12B, Jev-Omni, JevK5, Julia-1, Kev-27B, Lev 4B and Winnow ran on a rented GPU, so no per-call price.
    • Cost per 1,000 checks in the leading group: Jev $0.11, OpenAI Decisions API $0.21, GLiDE $0.34 and Clef $0.54 at the hosted price (the median over the three jobs); Jev-Omni €0.20, Kev-27B €0.25, Winnow €0.30 and Decision 2.0 Vega 27B €0.58 at the GPU rent (L4 €0.79 and H100 €2.87 per hour) times the box’s median time per check, one check at a time. Euros are counted one to one as dollars, which can only make a rented GPU look cheaper than it is.

    Release hard-bench-2026-10-07, integrity sha256-leaves-v1.