Open models caught up with Jev

Jev against 15 open models up to 28B on 11 real decision tasks, with 95% intervals and tie tiers. Rows trained on a task's data are grayed and never lead. Pick a task, filter, or click a model.

This benchmark was created using LangWatch. Sign up to create your own · Read the method

Task
Weights
Provider
Size

Showing 16 of 16 models. Filters hide rows; tiers and intervals stay as measured on every model.

Prompt injection

Catch rate at 5% false alarms, % · Higher is better

  1. Jev94.6
    , interval 92.2 to 96.7, tied for the lead
  2. , interval 90.5 to 97.1, tied for the lead
  3. , interval 90.3 to 96.2, tied for the lead
  4. , interval 86.6 to 95.0, tied for the lead
  5. , interval 88.2 to 94.2, tier 2
  6. , interval 88.3 to 94.0, tier 2
  7. , interval 87.8 to 93.5, tier 2
  8. , interval 82.4 to 92.5, tier 2

Ahead aloneTied for the leadThe restTrained on the test data*95% interval

One task at a time

Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.

Prompt injection: Catch rate at 5% false alarms, higher is better

  1. Jev94.6%, interval 94.6% [92.2, 96.7], tier 1
  2. Eikos-27B94.2%, interval 94.2% [90.5, 97.1], tier 1
  3. Shisa DE-194.0%, interval 94.0% [90.3, 96.2], tier 1
  4. AutoJev-27B92.0%, interval 92.0% [86.6, 95.0], tier 1
  5. Kev-4B91.6%, interval 91.6% [88.2, 94.2], tier 2
  6. Eikos-4B90.8%, interval 90.8% [88.3, 94.0], tier 2
  7. SemIf (Qwen3.5-4B)90.6%, interval 90.6% [87.8, 93.5], tier 2
  8. Kev-9B87.6%, interval 87.6% [82.4, 92.5], tier 2
  9. GLiNER2.5-Decide63.2%, interval 63.2% [55.8, 68.9], tier 3
  10. Kev-0.8B56.2%, interval 56.2% [46.3, 62.6], tier 3
  11. Laya-typed49.2%, interval 49.2% [42.6, 58.2], tier 4
  12. Kev-0.6B25.4%, interval 25.4% [16.9, 36.1], tier 5
  13. SimpleJev (Qwen3.5-0.8B)15.2%, interval 15.2% [10.6, 19.4], not ranked
  14. SemIf (Qwen3-0.6B)7.8%, interval 7.8% [5.1, 11.5], not ranked
  15. openJev Verdict 1.4*4.8%, interval 4.8% [2.8, 7.8], trained on this data, not ranked
  16. Laya0.0%, interval 0.0% [0.0, 0.0], not ranked
Dot: catch rate at 5% false alarms. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead.
Plot

Prompt injection: catch rate at 5% false alarms (higher is better) vs cost

↖ Better
  • Laya-typed
  • Kev-0.6B
  • Kev-0.8B
  • Kev-4B
  • Eikos-4B
← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑

Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.

Every model, every task

Every visible model on every task, sorted by Prompt injection. Value with half its interval; bands are tie tiers.
Model
Tier 1 on Prompt injection
Jev94.6±2.3 · T194.6% [92.2, 96.7], tier 190.3±2.2 · T190.3% [87.9, 92.4], tier 190.8±2.9 · T290.8% [88.1, 93.8], tier 280.3±2.8 · T180.3% [77.5, 83.1], tier 193.4±1.6 · T193.4% [91.8, 94.9], tier 189.1±2.0 · T289.1% [87.0, 91.0], tier 279.6±3.5 · T179.6% [76.0, 83.0], tier 178.3±2.5 · T278.3% [75.7, 80.7], tier 278.7±2.5 · T178.7% [76.3, 81.2], tier 168.3±2.9 · T168.3% [65.2, 71.1], tier 157.7±3.1 · T157.7% [54.7, 60.8], tier 173.9±2.1 · T173.9% [71.8, 76.0], tier 170.8±1.9 · T170.8% [68.9, 72.6], tier 162.3±2.8 · T162.3% [59.5, 65.1], tier 185.7±4.6 · T185.7% [81.0, 90.1], tier 1
Eikos-27B94.2±3.3 · T194.2% [90.5, 97.1], tier 190.3±2.2 · T190.3% [87.9, 92.4], tier 195.2±2.5 · T195.2% [92.8, 97.8], tier 182.1±2.8 · T182.1% [79.2, 84.9], tier 190.6*±1.890.6% [88.8, 92.4], not ranked89.6*±1.989.6% [87.7, 91.4], not ranked77.8*±3.777.8% [74.0, 81.4], not ranked79.3±2.5 · T279.3% [76.8, 81.8], tier 275.6±2.7 · T275.6% [73.0, 78.4], tier 266.8±2.9 · T166.8% [63.8, 69.6], tier 156.6±3.2 · T156.6% [53.4, 59.7], tier 173.4±2.2 · T173.4% [71.2, 75.6], tier 168.4±2.0 · T268.4% [66.5, 70.4], tier 259.8±2.8 · T259.8% [56.9, 62.6], tier 291.8*±3.691.8% [87.9, 95.2], not ranked
Shisa DE-194.0±2.9 · T194.0% [90.3, 96.2], tier 188.3±2.6 · T288.3% [85.6, 90.8], tier 284.2±5.9 · T384.2% [77.3, 89.1], tier 379.2±3.0 · T279.2% [76.2, 82.2], tier 288.0±1.9 · T388.0% [86.1, 89.9], tier 388.1±2.1 · T288.1% [85.9, 90.1], tier 20.0±0.00.0% [0.0, 0.0], not ranked81.3±2.3 · T181.3% [78.9, 83.5], tier 178.0±2.6 · T178.0% [75.4, 80.6], tier 165.1±2.9 · T265.1% [62.2, 68.0], tier 250.3±3.3 · T250.3% [47.1, 53.6], tier 275.1±2.2 · T175.1% [72.8, 77.3], tier 165.2±2.0 · T365.2% [63.3, 67.2], tier 338.2±2.738.2% [35.4, 40.9], not ranked86.1±4.7 · T186.1% [81.2, 90.6], tier 1
AutoJev-27B92.0±4.2 · T192.0% [86.6, 95.0], tier 190.6*±2.290.6% [88.4, 92.7], not ranked92.8±2.7 · T192.8% [90.2, 95.5], tier 179.1±2.9 · T279.1% [76.1, 82.0], tier 292.7*±1.792.7% [91.0, 94.4], not ranked89.7±2.0 · T189.7% [87.7, 91.6], tier 180.6±3.5 · T180.6% [77.0, 84.0], tier 180.6±2.4 · T180.6% [78.1, 82.9], tier 175.9±2.8 · T275.9% [73.2, 78.7], tier 266.7±2.9 · T166.7% [63.8, 69.5], tier 152.8±3.2 · T252.8% [49.6, 55.9], tier 273.8±2.1 · T173.8% [71.7, 75.9], tier 168.2±2.0 · T268.2% [66.2, 70.2], tier 263.0±2.8 · T163.0% [60.2, 65.8], tier 186.6±4.5 · T186.6% [81.9, 90.9], tier 1
Tier 2
Kev-4B91.6±3.0 · T291.6% [88.2, 94.2], tier 282.0±3.1 · T382.0% [78.8, 85.0], tier 39.0±2.5 · T59.0% [6.6, 11.7], tier 572.5±3.1 · T372.5% [69.4, 75.5], tier 384.4±2.3 · T484.4% [82.1, 86.7], tier 493.0*±1.593.0% [91.4, 94.5], not ranked85.0*±3.185.0% [81.8, 88.0], not ranked71.4±2.7 · T471.4% [68.6, 74.0], tier 476.2±2.7 · T276.2% [73.6, 79.0], tier 253.7±3.1 · T553.7% [50.5, 56.7], tier 544.8±3.1 · T344.8% [41.7, 48.0], tier 365.8±2.2 · T265.8% [63.6, 68.0], tier 258.0±2.1 · T558.0% [55.9, 60.2], tier 555.8±2.8 · T355.8% [53.0, 58.7], tier 371.4±6.1 · T371.4% [65.2, 77.4], tier 3
Eikos-4B90.8±2.9 · T290.8% [88.3, 94.0], tier 288.8±2.5 · T288.8% [86.3, 91.3], tier 278.6±5.4 · T378.6% [74.5, 85.3], tier 371.0±3.6 · T371.0% [67.4, 74.5], tier 379.6*±2.579.6% [77.1, 82.0], not ranked87.7*±2.187.7% [85.6, 89.7], not ranked74.0*±3.974.0% [70.0, 77.8], not ranked75.5±2.6 · T375.5% [72.8, 78.0], tier 3n/a56.7±3.1 · T456.7% [53.5, 59.7], tier 443.3±3.1 · T343.3% [40.1, 46.3], tier 365.3±2.7 · T265.3% [62.7, 68.0], tier 262.9±2.0 · T462.9% [60.9, 64.8], tier 455.8±2.9 · T355.8% [52.9, 58.7], tier 385.3*±4.785.3% [80.4, 89.8], not ranked
SemIf (Qwen3.5-4B)90.6±2.9 · T290.6% [87.8, 93.5], tier 288.7±2.5 · T288.7% [86.2, 91.2], tier 266.1±6.9 · T466.1% [60.9, 74.6], tier 472.0±3.1 · T372.0% [68.9, 75.2], tier 381.1±2.2 · T581.1% [79.0, 83.4], tier 50.0±0.00.0% [0.0, 0.0], not ranked0.0±0.00.0% [0.0, 0.0], not ranked79.9±2.4 · T179.9% [77.4, 82.2], tier 172.9±2.8 · T372.9% [70.2, 75.7], tier 361.2±3.1 · T361.2% [58.1, 64.2], tier 341.7±3.0 · T441.7% [38.6, 44.7], tier 463.0±2.5 · T363.0% [60.5, 65.4], tier 359.5±2.0 · T559.5% [57.5, 61.5], tier 537.3±2.837.3% [34.4, 40.0], not ranked80.1±5.3 · T280.1% [74.7, 85.3], tier 2
Kev-9B87.6±5.0 · T287.6% [82.4, 92.5], tier 288.5±2.4 · T288.5% [86.0, 90.9], tier 262.5±33.1 · T462.5% [0.2, 66.4], tier 473.1±3.0 · T373.1% [70.0, 76.0], tier 389.6±1.9 · T289.6% [87.7, 91.5], tier 293.0*±1.693.0% [91.4, 94.6], not ranked85.0*±3.285.0% [81.8, 88.2], not ranked72.4±2.7 · T472.4% [69.7, 75.0], tier 475.0±2.7 · T275.0% [72.4, 77.7], tier 255.5±3.0 · T455.5% [52.5, 58.5], tier 443.4±3.1 · T343.4% [40.2, 46.5], tier 3n/an/an/an/a
Tier 3
GLiNER2.5-Decide63.2±6.5 · T363.2% [55.8, 68.9], tier 371.5±4.0 · T571.5% [67.4, 75.4], tier 516.8±6.9 · T516.8% [11.0, 24.8], tier 557.8±3.657.8% [54.2, 61.5], not ranked49.3±0.749.3% [48.5, 50.0], not ranked86.0±2.2 · T386.0% [83.8, 88.2], tier 370.6±4.0 · T270.6% [66.6, 74.6], tier 232.4±2.9 · T732.4% [29.5, 35.3], tier 752.6±3.2 · T552.6% [49.5, 55.8], tier 550.4±3.1 · T650.4% [47.3, 53.5], tier 625.0±2.7 · T625.0% [22.4, 27.7], tier 648.8±2.1 · T448.8% [46.7, 50.8], tier 437.0±2.1 · T837.0% [34.9, 39.1], tier 841.4±2.8 · T541.4% [38.6, 44.2], tier 555.8±6.8 · T455.8% [48.9, 62.4], tier 4
Kev-0.8B56.2±8.2 · T356.2% [46.3, 62.6], tier 376.2±3.6 · T476.2% [72.5, 79.7], tier 422.8±10.1 · T522.8% [6.1, 26.3], tier 555.5±3.2 · T455.5% [52.2, 58.6], tier 476.8±2.6 · T676.8% [74.1, 79.3], tier 691.3*±1.891.3% [89.5, 93.0], not ranked83.0*±3.483.0% [79.6, 86.4], not ranked57.7±2.9 · T557.7% [54.6, 60.5], tier 538.5±3.0 · T838.5% [35.5, 41.6], tier 848.6±3.0 · T648.6% [45.6, 51.7], tier 629.3±2.8 · T529.3% [26.5, 32.2], tier 544.9±2.7 · T544.9% [42.2, 47.7], tier 550.9±2.0 · T750.9% [48.9, 52.9], tier 746.4±2.9 · T446.4% [43.5, 49.3], tier 460.2±6.7 · T460.2% [53.5, 66.8], tier 4
Tier 4
Laya-typed49.2±7.8 · T449.2% [42.6, 58.2], tier 470.0±3.8 · T570.0% [66.0, 73.7], tier 521.0±7.3 · T521.0% [14.6, 29.1], tier 549.4±1.2 · T549.4% [48.2, 50.6], tier 552.6±3.0 · T852.6% [49.6, 55.6], tier 877.0±2.7 · T477.0% [74.2, 79.5], tier 440.0±4.3 · T340.0% [35.8, 44.4], tier 321.5±2.5 · T921.5% [19.0, 24.0], tier 947.1±3.0 · T647.1% [44.1, 50.1], tier 649.7±3.0 · T649.7% [46.8, 52.8], tier 629.8±2.8 · T529.8% [27.0, 32.6], tier 577.4*±2.077.4% [75.4, 79.4], not ranked18.3±1.6 · T1018.3% [16.7, 19.9], tier 1047.8±2.8 · T447.8% [44.9, 50.5], tier 453.7±6.8 · T453.7% [46.9, 60.6], tier 4
Tier 5
Kev-0.6B25.4±9.6 · T525.4% [16.9, 36.1], tier 565.3±4.1 · T665.3% [61.1, 69.2], tier 69.6±5.9 · T59.6% [7.2, 19.1], tier 554.1±2.4 · T454.1% [51.8, 56.6], tier 455.8±2.7 · T755.8% [53.0, 58.5], tier 789.5*±2.089.5% [87.5, 91.5], not ranked78.2*±3.778.2% [74.4, 81.8], not ranked44.0±3.1 · T644.0% [40.8, 47.0], tier 659.3±3.0 · T459.3% [56.3, 62.3], tier 441.6±3.0 · T741.6% [38.6, 44.7], tier 726.2±2.7 · T626.2% [23.4, 28.9], tier 644.8±2.5 · T544.8% [42.3, 47.4], tier 524.8±1.9 · T924.8% [22.9, 26.7], tier 945.3±2.8 · T445.3% [42.5, 48.1], tier 460.6±6.6 · T460.6% [53.8, 67.1], tier 4
Not ranked on this task
SimpleJev (Qwen3.5-0.8B)15.2±4.415.2% [10.6, 19.4], not ranked49.9±4.549.9% [45.4, 54.4], not ranked3.4±1.73.4% [1.8, 5.1], not ranked50.8±3.950.8% [46.8, 54.7], not ranked58.8±3.058.8% [55.9, 61.9], not ranked60.2±3.0 · T560.2% [57.2, 63.2], tier 50.0±0.00.0% [0.0, 0.0], not ranked29.1±2.8 · T829.1% [26.3, 31.9], tier 845.2±3.2 · T645.2% [42.0, 48.4], tier 632.7±3.0 · T832.7% [29.7, 35.7], tier 825.0±2.7 · T725.0% [22.3, 27.7], tier 739.0±1.9 · T639.0% [37.1, 40.8], tier 656.3±1.6 · T656.3% [54.7, 57.8], tier 625.0±2.525.0% [22.5, 27.4], not ranked55.4±7.1 · T455.4% [48.1, 62.3], tier 4
SemIf (Qwen3-0.6B)7.8±3.27.8% [5.1, 11.5], not ranked63.0±4.263.0% [58.7, 67.0], not ranked9.6±4.19.6% [5.1, 13.3], not ranked52.1±3.8 · T552.1% [48.2, 55.9], tier 553.5±3.053.5% [50.5, 56.5], not ranked0.0±0.00.0% [0.0, 0.0], not ranked0.0±0.00.0% [0.0, 0.0], not ranked42.6±3.1 · T642.6% [39.4, 45.7], tier 637.8±2.9 · T837.8% [35.0, 40.9], tier 823.9±2.7 · T923.9% [21.4, 26.7], tier 925.0±2.7 · T725.0% [22.3, 27.7], tier 735.4±2.3 · T735.4% [33.1, 37.7], tier 717.9±1.6 · T1017.9% [16.4, 19.6], tier 1026.2±2.526.2% [23.7, 28.7], not ranked48.9±7.1 · T548.9% [41.8, 55.9], tier 5
openJev Verdict 1.44.8*±2.54.8% [2.8, 7.8], not ranked51.8±4.451.8% [47.5, 56.4], not ranked1.0±1.01.0% [0.2, 2.2], not ranked50.3±3.650.3% [46.8, 54.0], not ranked50.0*±0.050.0% [50.0, 50.0], not ranked78.4*±2.678.4% [75.8, 80.9], not ranked0.0*±0.00.0% [0.0, 0.0], not ranked16.4±2.3 · T1016.4% [14.1, 18.7], tier 1046.3±3.0 · T646.3% [43.2, 49.3], tier 626.5±2.7 · T926.5% [23.8, 29.2], tier 924.8±2.7 · T724.8% [22.1, 27.4], tier 736.4±2.4 · T636.4% [34.0, 38.8], tier 618.5±1.7 · T1018.5% [16.8, 20.3], tier 1022.6±2.422.6% [20.2, 25.0], not ranked57.6*±7.157.6% [50.2, 64.5], not ranked
Laya0.0±0.00.0% [0.0, 0.0], not ranked69.8±3.9 · T569.8% [65.8, 73.6], tier 515.8±11.5 · T515.8% [0.0, 23.0], tier 549.8±1.2 · T549.8% [48.6, 51.0], tier 554.3±2.8 · T854.3% [51.4, 57.1], tier 876.0±2.7 · T476.0% [73.3, 78.6], tier 442.4±4.3 · T342.4% [38.2, 46.8], tier 322.9±2.5 · T922.9% [20.2, 25.3], tier 944.2±2.9 · T744.2% [41.3, 47.2], tier 744.8±2.9 · T744.8% [41.9, 47.8], tier 731.5±2.8 · T531.5% [28.7, 34.3], tier 536.1±2.5 · T636.1% [33.6, 38.6], tier 617.8±1.5 · T1017.8% [16.3, 19.4], tier 1045.9±2.8 · T445.9% [43.1, 48.8], tier 458.9±6.4 · T458.9% [52.4, 65.2], tier 4

Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.

Jev

Provider
TypeSafe
Weights
Closed (API)
Size
Not disclosed
Licence
proprietary
TaskScore [95% interval]TierLatency p50Cost / 1k
Prompt injection94.6% [92.2, 96.7]1, tied~57 ms (Hosted API)n/a
Moderation90.3% [87.9, 92.4]1, tied~65 ms (Hosted API)n/a
PII90.8% [88.1, 93.8]2~65 ms (Hosted API)n/a
RAG faithfulness80.3% [77.5, 83.1]1, tied~65 ms (Hosted API)n/a
Off-topic93.4% [91.8, 94.9]1, ahead~59 ms (Hosted API)n/a
Routing, 20 intents89.1% [87.0, 91.0]2~67 ms (Hosted API)n/a
Routing, 77 intents79.6% [76.0, 83.0]1, tied~81 ms (Hosted API)n/a
Tool routing78.3% [75.7, 80.7]2~70 ms (Hosted API)n/a
Complaint routing78.7% [76.3, 81.2]1, tied~69 ms (Hosted API)n/a
Commit type68.3% [65.2, 71.1]1, tied~73 ms (Hosted API)n/a
Search relevance57.7% [54.7, 60.8]1, tied~63 ms (Hosted API)n/a
Typed decisions73.9% [71.8, 76.0]1, tied~68 ms (Hosted API)n/a
Web-agent actions70.8% [68.9, 72.6]1, ahead~68 ms (Hosted API)n/a
Community sets62.3% [59.5, 65.1]1, tied~65 ms (Hosted API)n/a
JevBench public85.7% [81.0, 90.1]1, tied~77 ms (Hosted API)n/a
  • Training data not disclosed, so contamination can't be ruled out.
  • ~ marks an approximate latency or an estimated cost.
How we measured

Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the metric difference, on the items every model answered.

Tie tiers

Models in one tier can't be separated by the tie test (paired bootstrap of the TPR@5%FPR difference on shared items, 95% CI includes 0; paired bootstrap of the AUROC difference on shared items, 95% CI includes 0; paired bootstrap of the Balanced acc. difference on shared items, 95% CI includes 0; paired McNemar on per-item correctness at the dev threshold, p >= 0.05; paired McNemar on per-item correctness, p >= 0.05; 1965 shared test items; paired McNemar on per-item correctness, p >= 0.05; 2329 shared test items; paired McNemar on per-item correctness, p >= 0.05; 1200 shared test items; paired McNemar on per-item correctness, p >= 0.05; 231 shared test items). Tiers are computed in the export over every model; filters on this page never change them.

Contamination

Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner.

Latency and hardware

  • NVIDIA RTX 3050 Ti Laptop GPU, 4 GB, WSL2 (cost at a cloud-equivalent rate): One model at a time, concurrency 1, on one of three GPUs, recorded per run: an NVIDIA RTX 3050 Ti Laptop GPU (4 GB, WSL2), a cloud NVIDIA L4 (24 GB) or a cloud NVIDIA H100 PCIe (80 GB). Each latency and cost carries the tier its run used; latency is comparable only within one tier.
  • NVIDIA L4, 24 GB (cloud): One model at a time, concurrency 1, on one of three GPUs, recorded per run: an NVIDIA RTX 3050 Ti Laptop GPU (4 GB, WSL2), a cloud NVIDIA L4 (24 GB) or a cloud NVIDIA H100 PCIe (80 GB). Each latency and cost carries the tier its run used; latency is comparable only within one tier.
  • NVIDIA H100 PCIe, 80 GB (cloud): One model at a time, concurrency 1, on one of three GPUs, recorded per run: an NVIDIA RTX 3050 Ti Laptop GPU (4 GB, WSL2), a cloud NVIDIA L4 (24 GB) or a cloud NVIDIA H100 PCIe (80 GB). Each latency and cost carries the tier its run used; latency is comparable only within one tier.
  • Hosted API: the laptop (Denmark), concurrency 1; approximate latency: p50 is the client time minus the network round trip measured in the same window (warm keep-alive requests to the API host); p95 is the API's server-reported time

Latencies are only compared within one hardware tier.

Datasets

  • Prompt injection (1,000 test items): deepset/prompt-injections (Apache-2.0), jackhhao/jailbreak-classification (Apache-2.0), reshabhs/SPML_Chatbot_Prompt_Injection (MIT), djapp18/JailbreaksOverTime (CC-BY-4.0)
  • Moderation (667 test items): mmathys/openai-moderation-api-evaluation (MIT), nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (CC-BY-4.0)
  • PII (1,000 test items): gretelai/gretel-pii-masking-en-v1 (Apache-2.0), beki/privy (MIT), nvidia/Nemotron-PII (CC-BY-4.0)
  • RAG faithfulness (666 test items): pminervini/HaluEval (Apache-2.0)
  • Off-topic (1,000 test items): clinc/clinc_oos (CC-BY-3.0), mteb/amazon_massive_intent (CC-BY-4.0), benayas/snips (CC0-1.0)
  • Routing, 20 intents (1,000 test items): legacy-datasets/banking77 (CC-BY-4.0)
  • Routing, 77 intents (500 test items): legacy-datasets/banking77 (CC-BY-4.0)
  • Tool routing (1,000 test items): gorilla-llm/Berkeley-Function-Calling-Leaderboard (Apache-2.0)
  • Complaint routing (1,000 test items): BEE-spoke-data/consumer-finance-complaints (CC0-1.0)
  • Commit type (1,000 test items): github.com/angular/angular (MIT), github.com/vitejs/vite (MIT)
  • Search relevance (1,000 test items): tasksource/esci (Apache-2.0)
  • Typed decisions (1,965 test items): LocalLLaMA/typed-decisions (Apache-2.0)
  • Web-agent actions (2,329 test items): AndeyTait/JevForge-Mind2Web (CC BY 4.0)
  • Community sets (1,200 test items): Praveenrajus/jev-bench (CC BY 4.0)
  • JevBench public (231 test items): fstandhartinger/jevbench (MIT)
  • gliner2.5-decide: Related specialty (the card names moderation; the release blog demonstrates prompt injection); no jailbreak dataset is named; no overlap with Fastino's published sample (the training set itself is unpublished).
  • laya-typed: Trained on unnamed jailbreak and prompt-injection data; deepset/prompt-injections is declared held out.
  • verdict-1.4: fitted its calibrator on deepset/prompt-injections train split (`use: calibration`); fitted its calibrator on TrustAIRLab in-the-wild jailbreaks (inside djapp18/JailbreaksOverTime) (`use: calibration`): that is 2 of this suite's 4 sources, 50% of its items. Shown for reference, not ranked here.
  • laya: Trained on unnamed jailbreak and prompt-injection data; deepset/prompt-injections is declared held out.
  • Verdict 1.4 returns a near-constant probability here: the middle 90% of its answers span 0.12 around 0.5, and on off-topic it returned one identical value for all 1,000 items, so this AUROC ranks numerical noise rather than measuring the model. Reading it upside down would not give a real score either. Scoring polarity was checked against the author's own reference adapter and is correct (harness investigation, 2026-09-23).
  • GLiNER2.5-Decide ran behind our own shim over the author's gliner2 2.0.0 package, in FP16 on a 4 GB laptop GPU; the checkpoint is FP32, which spills out of GPU memory on long inputs on that card. Against FP32 on CPU, FP16 picked the same answer on 20 of 20 short inputs (max probability difference 0.0002) and on 54 of 54 long ones of 1,962 to 4,053 tokens (max difference 0.0017). Nothing is truncated (gliner2's default), so long inputs run past the 512 positions of the DeBERTa encoder; label names and descriptions count toward that length, and the author's guidance for long documents is chunking, which this run does not do. Requests over 512 encoded tokens: routing-77 100%, routing-20 82%, complaint-routing 81%, x-typed-decisions 75%, tool-routing 58%, x-jevbench-community 29%, rag-faithfulness 22%, prompt-injection 16%, moderation 15%, x-mind2web-actions 14%, search-relevance 8%, pii 7%, commit-type 6%, off-topic 0.2%. Even in FP16, requests above about 2,300 tokens spill out of GPU memory: 2 prompt-injection requests (73% of that task's measured time) and 1 tool-routing request (4%), so those latencies describe our card more than the model. Its latency is laptop-tier and not comparable with the cloud rows. There is no independent reference to reproduce these numbers against: Fastino's own test split is private.
  • autojev-27b: trained on nvidia/Aegis-AI-Content-Safety-Dataset-2.0 test split (`use: weights`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: data.py TRAIN_QUOTAS.
  • gliner2.5-decide: The model card names moderation as a specialty of its unpublished synthetic training data; none of our sources is named; no overlap with Fastino's published sample (the training set itself is unpublished).
  • laya-typed: Trained on unnamed human-labelled toxicity data, a closely related task.
  • laya: Trained on unnamed human-labelled toxicity data, a closely related task.
  • verdict-1.4: Its base model's training lineage includes toxicity and hate-speech datasets, a closely related task.
  • gliner2.5-decide: Fastino's evaluation suite, made by the same generator as the training data, has a contains-PII question; none of our PII sources is named; no overlap with Fastino's published sample (the training set itself is unpublished).
  • This zero-shot readout saturates at its lowest rating bin (every probability between 0.0102 and 0.0143) and answers 'no PII' to all 1,000 documents. The residual ordering runs backwards because our PII negatives are redacted documents whose stand-in phrases name the very categories the questions ask about ('the address on file', 'the customer'), which is a documented property of the dataset. The score is a dataset artefact, not an inverted signal (harness investigation, 2026-09-23).
  • gliner2.5-decide: The model card names yes/no questions over a passage as a capability of its unpublished synthetic training; no NLI or faithfulness dataset is named; no overlap with Fastino's published sample (the training set itself is unpublished).
  • kev-0.8b: Trained on MNLI and BoolQ, closely related entailment and passage-QA tasks (no item overlap found).
  • kev-0.6b: Trained on MNLI and BoolQ, closely related entailment and passage-QA tasks (no item overlap found).
  • verdict-1.4: Its base model's training lineage includes natural-language-inference datasets, a closely related task.
  • laya: Trained on NLI and fact-verification data (unnamed), closely related to faithfulness checking.
  • laya-typed: Trained on NLI and fact-verification data (unnamed), closely related to faithfulness checking.
  • Laya does carry signal on one of this framing's three sub-questions (contradicts: AUROC 61 / 64), but that question also returns the lowest probabilities, so the pre-registered max-combine rule almost always hands the item's score to one of the two uninformative sub-questions instead. The result is a combined score at or slightly below chance for both Laya and Laya-typed. This is a limitation of the max-combine framing on this model, not a polarity error; the single-question 'broad' framing is also at chance for Laya here (harness investigation, 2026-09-23).
  • GLiNER2.5-Decide's scores are marked degenerate here: its probabilities barely vary (the central 90% spans 0.141 on this subset, just under the 0.15 bar; 0.152 on the full test set, just over it). Its balanced accuracy is 57.8%, and its AUROC of 59.3% [55.1, 63.5] shows the compressed scores order the items only slightly better than chance, so read both numbers as fragile.
  • autojev-27b: trained on mteb/amazon_massive_intent test (en-US, de-DE) split (`use: weights`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: data.py TRAIN_QUOTAS.
  • eikos-27b: selected its checkpoint on clinc/clinc_oos (`use: selection`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: NOTICE: evaluation-only, part of the card's 'general battery'.
  • eikos-4b: selected its checkpoint on clinc/clinc_oos (`use: selection`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: NOTICE: evaluation-only, part of the card's 'general battery'.
  • laya: Trained on unnamed intent-classification data, the task family of this category's sources.
  • laya-typed: Trained on unnamed intent-classification data, the task family of this category's sources.
  • verdict-1.4: trained on clinc/clinc_oos oos_train (100 out-of-scope queries only) split (`use: weights`); fitted its calibrator on clinc/clinc_oos test split (`use: calibration`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: training code.
  • gliner2.5-decide: The model card names intent classification as a specialty of its unpublished synthetic training data; CLINC, MASSIVE and SNIPS are not named as training data (MASSIVE and SNIPS were zero-shot evaluation sets in the GLiNER2 paper); no overlap with Fastino's published sample (the training set itself is unpublished).
  • GLiNER2.5-Decide's balanced accuracy here is 49.3%, chance level, and its scores are marked degenerate: its probabilities barely vary (the central 90% spans 0.13). What ordering they have runs the wrong way (AUROC 35.9%): the model gives higher P(yes) to in-scope messages when asked whether a message is outside the listed scope. With a single question its dev AUROC is 0.29 (broad) and 0.26 (detailed); in the chosen decomposed framing the 'unsupported task' question is inverted the same way (test AUROC 0.29) and supplies the max-combined score on 552 of 1,000 items, while 'other subject' alone is correctly oriented (0.71). Scoring polarity is correct: the same yes/no mapping is correctly oriented on every other binary task. This is how the model reads a negated scope question (harness investigation, 2026-09-25).
  • kev-4b: trained on legacy-datasets/banking77 (`use: weights`); selected its checkpoint on legacy-datasets/banking77 test split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: inherited from kev-0.8b; unverified for this checkpoint.
  • kev-9b: trained on legacy-datasets/banking77 (`use: weights`); selected its checkpoint on legacy-datasets/banking77 test split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: inherited from kev-0.8b; unverified for this checkpoint.
  • kev-0.8b: trained on legacy-datasets/banking77 (`use: weights`); selected its checkpoint on legacy-datasets/banking77 test split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: Kev's own development split; also the served temperature.
  • eikos-27b: selected its checkpoint on legacy-datasets/banking77 (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: NOTICE: evaluation-only, part of the card's 'general battery'.
  • kev-0.6b: trained on legacy-datasets/banking77 (`use: weights`); selected its checkpoint on legacy-datasets/banking77 test split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: Kev's own development split.
  • eikos-4b: selected its checkpoint on legacy-datasets/banking77 (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: NOTICE: evaluation-only, part of the card's 'general battery'.
  • gliner2.5-decide: The model card names customer and banking intent as a specialty of its unpublished synthetic training data; Banking77 is not named; no overlap with Fastino's published sample (the training set itself is unpublished).
  • verdict-1.4: trained on legacy-datasets/banking77 train split (`use: weights`); fitted its calibrator on legacy-datasets/banking77 test split (`use: calibration`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: training code.
  • laya-typed: Trained on unnamed banking-intent data; its authors list Banking77 itself as held out.
  • laya: Trained on unnamed banking-intent data; its authors list Banking77 itself as held out.
  • Option-definition bug: all 13 routing-20 items whose gold intent is get_physical_card are PIN questions, which contradicts our written definition of that option ('the customer asks about getting a physical card'). Jev, which reads only the definition, gets 1 of 13; Kev-0.8B, trained on Banking77, gets 13 of 13. beneficiary_not_allowed has the same problem. A memorised dataset convention shows up as accuracy. With the affected items and the adjudicated label errors removed, routing-20's first tier is a three-way tie of Kev-0.8B, Jev and Kev-0.6B, and Jev joins routing-77's first tier (data-quality audit, 2026-09-22). The audit itself warns that dropping flagged items favours Jev by construction, because the flagged items are the ones Jev confidently missed, so read that re-scored tier as a bound, not a result.
  • 27 routing-20 and 18 routing-77 messages also appear verbatim in Kev's own locked test set.
  • GLiNER2.5-Decide never picks none_of_these (0 of 1,000 items), so it scores 0% on the bfcl-live-irrelevance half and 64.8% on bfcl-live-multiple; its overall 32.4% is below always-majority (50%) and every trivial baseline.
  • gliner2.5-decide: The model card names email and ticket routing as a specialty of its unpublished synthetic training data; CFPB complaints are not named; no overlap with Fastino's published sample (the training set itself is unpublished).
  • laya-typed: Trained on support-ticket queue routing (Tobi-Bueck/customer-support-tickets), a closely related task.
  • laya: Trained on support-ticket queue routing (Tobi-Bueck/customer-support-tickets), a closely related task.
  • laya: Trained on MS MARCO query-passage relevance, a closely related task.
  • laya-typed: Trained on MS MARCO query-passage relevance, a closely related task.
  • laya-typed: trained on LocalLLaMA/typed-decisions train split (`use: weights`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: model card.
  • gliner2.5-decide: The model card names typed decisions of the same kind (agent completion, severity, routing) as specialties of its unpublished synthetic training; typed-decisions is not named; no overlap with Fastino's published sample (the training set itself is unpublished).
  • GLiNER2.5-Decide answers this suite's yes/no questions 'yes' 70.5% of the time against 22.6% in the gold labels, so its yes/no accuracy (42.3%) is below always answering 'no' (77.4%).
  • autojev-27b: trained on amazon_massive_intent (inside Praveenrajus/jev-bench) test (en-US, de-DE) split (`use: weights`); trained on google/civil_comments (inside Praveenrajus/jev-bench) (`use: weights`): that is 2 of this suite's 8 sources, 25% of its items. Still ranked: the overlap is below the one-third bar. Declared evidence: data.py TRAIN_QUOTAS.
  • laya-typed: Trained on unnamed toxicity, rubric-rating and fact-checking data whose descriptions match four of this suite's eight configs; possibly the same datasets.
  • kev-0.8b: Trained on MNLI, a task close to this suite's fact-checking config (150 of 1,200 items).
  • laya: Trained on unnamed toxicity, rubric-rating and fact-checking data whose descriptions match four of this suite's eight configs; possibly the same datasets.
  • kev-0.6b: Trained on MNLI, a task close to this suite's fact-checking config (150 of 1,200 items).
  • gliner2.5-decide: Partial: the model card names moderation, ordinal scores and intents as specialties, related to some configs of this suite; no overlap with Fastino's published sample (the training set itself is unpublished).
  • verdict-1.4: Its base model's lineage includes toxicity, hate-speech and NLI data, close to three of this suite's eight configs.
  • GLiNER2.5-Decide answers this suite's yes/no questions 'yes' 3% of the time against 50% in the gold labels (yes/no accuracy 49.0%), and 4 of its 8 configs are at or below random: civil_comments 48.7 and fever_evidence 49.3 (random 50), helpsteer2 helpfulness 18.7 and verbosity 16.0 (random 20).
  • eikos-27b: selected its checkpoint on fstandhartinger/jevbench public split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: NOTICE evaluation-only + 8-gram decontamination in snapshot_final.sh; reported in the card's release table.
  • eikos-4b: selected its checkpoint on fstandhartinger/jevbench public split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: NOTICE evaluation-only + 8-gram decontamination in snapshot_final.sh; reported in the card's release table.
  • kev-0.6b: Trained on template-generated policy-rule cases like this set's policy items (no item overlap found).
  • kev-0.8b: Trained on template-generated policy-rule cases like this set's policy items (no item overlap found).
  • verdict-1.4: tuned its inference settings on fstandhartinger/jevbench 231 public items split (`use: tuning`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: README "measured on the 231 public JevBench tasks".
  • On this suite 24% of GLiNER2.5-Decide's requests exceed 512 encoded tokens, and 25 requests above about 2,300 tokens spill out of GPU memory even in FP16, taking 98% of its measured time: its latency here describes our card, not the model.

Release 2026-09-25.2, benchmark commit c372cd5, 34 input files, integrity sha256-leaves-v1.