Nobody clearly beats Jev yet

Decision 2.0 Vega 27B scores 75.8% and Jev 72.7% on 600 checks. With 19 models compared against Jev, no lead over Jev is large enough to be sure. Leading group: Decision 2.0 Vega 27B, Clef, GLiDE, Jev-Omni, Winnow, Jev and Kev-27B.

This benchmark was created using LangWatch. Sign up to create your own · Read the method

LangWatch Instant Evals runs on Jev. Check it out

Task

All three jobs

Accuracy, % · Higher is better

Decision 2.0 Vega 27B leads Jev by 3.2 points, too little to be sure across 19 comparisons with Jev: the 95% interval of the gap runs from 0.2 to 6.1 points.

  1. , interval 72.5 to 79.4, tied for the lead
  2. Clef74.5
    , interval 70.9 to 78.2, tied for the lead
  3. GLiDE74.3
    , interval 70.9 to 78.0, tied for the lead
  4. , interval 69.7 to 77.2, tied for the lead
  5. , interval 69.4 to 76.9, tied for the lead
  6. Jev72.7
    , interval 69.2 to 76.3, tied for the lead
  7. , interval 68.5 to 75.4, tied for the lead
  8. , interval 65.7 to 73.3, tier 2

Ahead aloneTied for the leadThe rest95% interval

One job at a time

The three judge jobs together. 600 checks, 200 per job. Each check gives the judge an agent’s answer and the tool or retrieval result it came from. To verify it, the judge redoes one lookup plus one step: filter rows, count, do arithmetic or date math, convert units, or apply a tiered policy. Half the answers are right; the rest are subtly wrong, such as a neighbouring table row, a superseded record, a total before discount, calendar instead of working days, a tier boundary off by one, per serving instead of per 100 g, or one wrong VAT digit.

All three jobs: Accuracy, higher is better

  1. Decision 2.0 Vega 27B75.8%, interval 75.8% [72.5, 79.4], tier 1
  2. Clef74.5%, interval 74.5% [70.9, 78.2], tier 1
  3. GLiDE74.3%, interval 74.3% [70.9, 78.0], tier 1
  4. Jev-Omni73.3%, interval 73.3% [69.7, 77.2], tier 1
  5. Winnow73.0%, interval 73.0% [69.4, 76.9], tier 1
  6. Jev72.7%, interval 72.7% [69.2, 76.3], tier 1
  7. Kev-27B72.0%, interval 72.0% [68.5, 75.4], tier 1
  8. DiffusionGemma-Jev69.5%, interval 69.5% [65.7, 73.3], tier 2
  9. Cygnet65.8%, interval 65.8% [62.0, 69.8], tier 2
  10. Hopper 12B65.8%, interval 65.8% [62.0, 69.8], tier 2
  11. Decision 2.0 Lux 9B61.0%, interval 61.0% [57.2, 65.0], tier 3
  12. Clef Flash58.7%, interval 58.7% [54.7, 62.9], tier 3
  13. decider-4b v258.3%, interval 58.3% [54.5, 62.0], tier 3
  14. JevK557.7%, interval 57.7% [54.1, 61.4], tier 3
  15. Decision 2.0 Nox 4B56.3%, interval 56.3% [52.6, 60.3], tier 4
  16. Lev 4B52.7%, interval 52.7% [49.0, 56.3], tier 5
  17. Julia-150.0%, interval 50.0% [46.4, 53.5], tier 5
  18. Decision 2.0 Eos 0.8B49.7%, interval 49.7% [45.9, 53.4], tier 5
  19. Decision 2.0 Sol 2B48.5%, interval 48.5% [44.7, 52.3], tier 6
  20. Decision 2.0 Kai 0.6B48.3%, interval 48.3% [44.3, 52.3], tier 6
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 40%, not 0.
Plot

All three jobs: accuracy (higher is better) vs cost

↖ Better
  • Jev
  • Clef
  • GLiDE
← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑

Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Cygnet, decider-4b v2, Decision 2.0 Eos 0.8B, Decision 2.0 Kai 0.6B, Decision 2.0 Lux 9B, Decision 2.0 Nox 4B, Decision 2.0 Sol 2B, Decision 2.0 Vega 27B, DiffusionGemma-Jev, Hopper 12B, Jev-Omni, JevK5, Julia-1, Kev-27B, Lev 4B, Winnow.

Every model, every job

Every visible model on every task, sorted by All three jobs. Value with half its interval; bands are tie tiers.
Model
Tier 1 on All three jobs
Decision 2.0 Vega 27B75.8±3.5 · T175.8% [72.5, 79.4], tier 175.5±5.7 · T175.5% [69.8, 81.3], tier 171.0±6.2 · T171.0% [64.5, 77.0], tier 181.0±5.3 · T181.0% [75.5, 86.1], tier 1
Clef74.5±3.6 · T174.5% [70.9, 78.2], tier 173.0±6.8 · T173.0% [66.2, 79.7], tier 167.5±5.9 · T167.5% [61.3, 73.1], tier 183.0±5.5 · T183.0% [77.4, 88.4], tier 1
GLiDE74.3±3.5 · T174.3% [70.9, 78.0], tier 172.0±6.2 · T172.0% [65.7, 78.1], tier 171.5±6.7 · T171.5% [64.4, 77.8], tier 179.5±5.1 · T179.5% [74.4, 84.7], tier 1
Jev-Omni73.3±3.8 · T173.3% [69.7, 77.2], tier 172.0±6.5 · T172.0% [65.5, 78.5], tier 172.5±6.5 · T172.5% [65.7, 78.7], tier 175.5±6.2 · T175.5% [69.3, 81.7], tier 1
Winnow73.0±3.8 · T173.0% [69.4, 76.9], tier 170.5±6.2 · T170.5% [64.0, 76.3], tier 169.5±6.2 · T169.5% [63.2, 75.6], tier 179.0±5.9 · T179.0% [73.1, 85.0], tier 1
Jev72.7±3.5 · T172.7% [69.2, 76.3], tier 172.5±5.8 · T172.5% [66.7, 78.3], tier 169.0±6.6 · T169.0% [62.3, 75.5], tier 176.5±5.8 · T176.5% [70.4, 82.1], tier 1
Kev-27B72.0±3.4 · T172.0% [68.5, 75.4], tier 172.0±6.4 · T172.0% [65.6, 78.4], tier 166.0±6.0 · T166.0% [59.7, 71.8], tier 178.0±6.2 · T178.0% [71.6, 84.0], tier 1
Tier 2
DiffusionGemma-Jev69.5±3.8 · T269.5% [65.7, 73.3], tier 270.5±6.8 · T170.5% [63.8, 77.4], tier 163.0±6.6 · T263.0% [56.3, 69.5], tier 275.0±5.6 · T275.0% [69.5, 80.6], tier 2
Cygnet65.8±3.9 · T265.8% [62.0, 69.8], tier 260.0±6.9 · T260.0% [52.8, 66.7], tier 262.5±6.0 · T262.5% [56.5, 68.6], tier 275.0±6.5 · T275.0% [68.5, 81.3], tier 2
Hopper 12B65.8±3.9 · T265.8% [62.0, 69.8], tier 265.0±6.9 · T265.0% [57.8, 71.6], tier 257.0±6.3 · T257.0% [50.5, 63.2], tier 275.5±6.0 · T175.5% [69.5, 81.5], tier 1
Tier 3
Decision 2.0 Lux 9B61.0±3.9 · T361.0% [57.2, 65.0], tier 359.0±6.9 · T259.0% [52.0, 65.8], tier 259.5±6.6 · T259.5% [53.1, 66.2], tier 264.5±6.8 · T364.5% [57.4, 71.0], tier 3
Clef Flash58.7±4.1 · T358.7% [54.7, 62.9], tier 359.5±7.1 · T259.5% [52.3, 66.5], tier 256.5±6.3 · T356.5% [50.3, 62.9], tier 360.0±6.9 · T360.0% [53.0, 66.8], tier 3
decider-4b v258.3±3.8 · T358.3% [54.5, 62.0], tier 359.5±6.9 · T259.5% [52.7, 66.5], tier 254.5±6.5 · T354.5% [48.0, 60.9], tier 361.0±6.1 · T361.0% [54.7, 67.0], tier 3
JevK557.7±3.6 · T357.7% [54.1, 61.4], tier 355.0±7.0 · T355.0% [47.7, 61.8], tier 354.5±5.9 · T354.5% [48.5, 60.2], tier 363.5±6.0 · T363.5% [57.6, 69.7], tier 3
Tier 4
Decision 2.0 Nox 4B56.3±3.8 · T456.3% [52.6, 60.3], tier 458.5±6.8 · T258.5% [51.7, 65.4], tier 254.5±6.3 · T354.5% [48.3, 61.0], tier 356.0±6.4 · T456.0% [49.5, 62.4], tier 4
Tier 5
Lev 4B52.7±3.6 · T552.7% [49.0, 56.3], tier 550.5±6.3 · T350.5% [44.3, 56.9], tier 353.0±6.2 · T353.0% [46.9, 59.3], tier 354.5±6.4 · T454.5% [48.0, 60.9], tier 4
Julia-150.0±3.6 · T550.0% [46.4, 53.5], tier 553.0±5.6 · T353.0% [47.7, 58.8], tier 349.5±6.3 · T349.5% [43.4, 56.1], tier 347.5±6.9 · T547.5% [40.5, 54.3], tier 5
Decision 2.0 Eos 0.8B49.7±3.7 · T549.7% [45.9, 53.4], tier 551.5±6.4 · T351.5% [45.1, 57.8], tier 348.0±6.5 · T348.0% [41.5, 54.5], tier 349.5±6.1 · T549.5% [43.4, 55.6], tier 5
Tier 6
Decision 2.0 Sol 2B48.5±3.8 · T648.5% [44.7, 52.3], tier 651.0±6.5 · T351.0% [44.5, 57.5], tier 346.5±6.9 · T446.5% [39.8, 53.6], tier 448.0±6.4 · T548.0% [41.7, 54.5], tier 5
Decision 2.0 Kai 0.6B48.3±4.0 · T648.3% [44.3, 52.3], tier 650.0±7.0 · T350.0% [43.1, 57.1], tier 345.0±6.6 · T445.0% [38.5, 51.6], tier 450.0±7.1 · T450.0% [42.9, 57.2], tier 4

Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. All three jobs: every check of the three jobs pooled, one accuracy over 600 checks.

Jev

Provider
TypeSafe
Weights
Closed (API)
Size
Not disclosed
Licence
Proprietary
TaskScore [95% interval]TierLatency p50Cost / 1k
All three jobs72.7% [69.2, 76.3]1, tied269 ms (Hosted API)$0.086
LLM-as-a-judge72.5% [66.7, 78.3]1, tied258 ms (Hosted API)$0.11
Search69.0% [62.3, 75.5]1, tied256 ms (Hosted API)$0.11
Scenario judge76.5% [70.4, 82.1]1, tied325 ms (Hosted API)$0.040
    How we measured

    Scores carry a 95% bootstrap interval (2,000 resamples, the checks built from one scenario resampled together). Two models are tied when the 95% interval of their paired difference on the same checks holds zero.

    Tie tiers

    Models in one tier can't be separated by the tie test (the 95% interval of the paired difference on the same cases holds zero). Tiers are computed in the export over every model; filters on this page never change them.

    Contamination

    The checks are synthetic, generated for this benchmark from made-up fixture data, with every label computed from that data. They were never published, so no model can have trained on them. The page shows aggregates only: no check and no check text.

    Latency and hardware

    • Hosted API: Jev, Clef, Clef Flash and GLiDE, called as hosted APIs from a laptop, one call per check; the time includes the network
    • H100: Cygnet, Decision 2.0 Lux 9B, Decision 2.0 Vega 27B, DiffusionGemma-Jev, Hopper 12B, Jev-Omni and Kev-27B, open weights on one rented H100; the server’s own time per check, 30 checks one at a time
    • L4: decider-4b v2, Decision 2.0 Eos 0.8B, Decision 2.0 Kai 0.6B, Decision 2.0 Nox 4B, Decision 2.0 Sol 2B, JevK5, Lev 4B and Winnow, open weights on one rented L4; the server’s own time per check, 30 checks one at a time
    • Server CPU: Julia-1, served on the GPU box’s CPU; the server’s own time per check, 30 checks one at a time

    Latencies are only compared within one hardware tier.

    Datasets

    • All three jobs (600 test items): synthetic lookup and arithmetic checks, labels computed from the fixture data (private, aggregates only)
    • LLM-as-a-judge (200 test items): synthetic lookup and arithmetic checks, labels computed from the fixture data (private, aggregates only)
    • Search (200 test items): synthetic lookup and arithmetic checks, labels computed from the fixture data (private, aggregates only)
    • Scenario judge (200 test items): synthetic lookup and arithmetic checks, labels computed from the fixture data (private, aggregates only)
    • An error, a reply with no verdict, or an inconclusive Scenario verdict counts as wrong.
    • All three jobs: every check of the three jobs pooled, one accuracy over 600 checks.
    • Latency is the median time per check, compared only within one hardware tier. Hosted APIs were timed from a laptop on every check, so their time includes the network. Open-weights models were timed on the GPU box by the server itself, 30 checks one at a time, one number for every job.
    • The 15 open-weights models on the GPU ran in bf16 on a single rented L4 or H100 (Winnow in Q8_0 GGUF). Julia-1 ran in F32 on the box’s CPU.
    • GLiDE: A hosted “thinking” decision model: an uncertain check can take extra time.
    • Cygnet: A recipe, not a trained model: a letter readout on stock Gemma 4 12B IT. Run without its calibration temperature, which rescales probabilities but cannot flip a yes or no at the 0.5 threshold.
    • decider-4b v2: Served with a smaller CUDA graph token budget than its default, a capacity setting that does not change answers.
    • JevK5: Takes an 8k token state; no check was longer.
    • Julia-1: Served on the box’s CPU, so its latency is CPU time.
    • Lev 4B: Takes an 8k token state; no check was longer.
    • Winnow: Run from its Q8_0 GGUF weights, the precision its authors recommend.
    • Jev, Clef, Clef Flash and GLiDE: cost per 1,000 checks as recorded on the run, at the provider’s price. Cygnet, decider-4b v2, Decision 2.0 Eos 0.8B, Decision 2.0 Kai 0.6B, Decision 2.0 Lux 9B, Decision 2.0 Nox 4B, Decision 2.0 Sol 2B, Decision 2.0 Vega 27B, DiffusionGemma-Jev, Hopper 12B, Jev-Omni, JevK5, Julia-1, Kev-27B, Lev 4B and Winnow ran on a rented GPU, so no per-call price.

    Release hard-bench-2026-10-05, integrity sha256-leaves-v1.