Jev beats every tiny open model

Jev against 7 open models under 1B parameters on 15 tasks, with 95% intervals and tie tiers. Rows trained on a task's data are grayed and never lead. Pick a task, filter, or click a model.

This benchmark was created using LangWatch. Sign up to create your own · Read the method

Task
Weights
Provider

Showing 8 of 8 models. Filters hide rows; tiers and intervals stay as measured on every model.

Prompt injection

Catch rate at 5% false alarms, % · Higher is better

  1. Jev94.6
    , interval 92.2 to 96.7, ahead alone
  2. , interval 46.3 to 62.6, tier 2
  3. , interval 42.6 to 58.2, tier 2
  4. , interval 16.9 to 36.1, tier 3
  5. , interval 10.6 to 19.4, not ranked
  6. , interval 5.1 to 11.5, not ranked
  7. , interval 2.8 to 7.8, not ranked
  8. , interval 0.0 to 0.0, not ranked

Ahead aloneTied for the leadThe restTrained on the test data*95% interval

One task at a time

Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.

Prompt injection: Catch rate at 5% false alarms, higher is better

  1. Jev94.6%, interval 94.6% [92.2, 96.7], tier 1
  2. Kev-0.8B56.2%, interval 56.2% [46.3, 62.6], tier 2
  3. Laya-typed49.2%, interval 49.2% [42.6, 58.2], tier 2
  4. Kev-0.6B25.4%, interval 25.4% [16.9, 36.1], tier 3
  5. SimpleJev (Qwen3.5-0.8B)15.2%, interval 15.2% [10.6, 19.4], not ranked
  6. SemIf (Qwen3-0.6B)7.8%, interval 7.8% [5.1, 11.5], not ranked
  7. openJev Verdict 1.44.8%, interval 4.8% [2.8, 7.8], not ranked
  8. Laya0.0%, interval 0.0% [0.0, 0.0], not ranked
Dot: catch rate at 5% false alarms. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead.
Plot

Prompt injection: catch rate at 5% false alarms (higher is better) vs cost

↖ Better
  • Laya-typed
  • Kev-0.6B
  • Kev-0.8B
← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑

Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.

Every model, every task

Every visible model on every task, sorted by Prompt injection. Value with half its interval; bands are tie tiers.
Model
Tier 1 on Prompt injection
Jev94.6±2.3 · T194.6% [92.2, 96.7], tier 190.3±2.2 · T190.3% [87.9, 92.4], tier 190.8±2.9 · T190.8% [88.1, 93.8], tier 180.3±2.8 · T180.3% [77.5, 83.1], tier 193.4±1.6 · T193.4% [91.8, 94.9], tier 189.1±2.0 · T189.1% [87.0, 91.0], tier 179.6±3.5 · T179.6% [76.0, 83.0], tier 178.3±2.5 · T178.3% [75.7, 80.7], tier 178.7±2.5 · T178.7% [76.3, 81.2], tier 168.3±2.9 · T168.3% [65.2, 71.1], tier 157.7±3.1 · T157.7% [54.7, 60.8], tier 173.9±2.1 · T173.9% [71.8, 76.0], tier 170.8±1.9 · T170.8% [68.9, 72.6], tier 162.3±2.8 · T162.3% [59.5, 65.1], tier 185.7±4.6 · T185.7% [81.0, 90.1], tier 1
Tier 2
Kev-0.8B56.2±8.2 · T256.2% [46.3, 62.6], tier 276.2±3.6 · T276.2% [72.5, 79.7], tier 222.8±10.1 · T222.8% [6.1, 26.3], tier 255.5±3.2 · T255.5% [52.2, 58.6], tier 276.8±2.6 · T276.8% [74.1, 79.3], tier 291.3*±1.891.3% [89.5, 93.0], not ranked83.0*±3.483.0% [79.6, 86.4], not ranked57.7±2.9 · T257.7% [54.6, 60.5], tier 238.5±3.0 · T538.5% [35.5, 41.6], tier 548.6±3.0 · T248.6% [45.6, 51.7], tier 229.3±2.8 · T229.3% [26.5, 32.2], tier 244.9±2.7 · T244.9% [42.2, 47.7], tier 250.9±2.0 · T350.9% [48.9, 52.9], tier 346.4±2.9 · T246.4% [43.5, 49.3], tier 260.2±6.7 · T260.2% [53.5, 66.8], tier 2
Laya-typed49.2±7.8 · T249.2% [42.6, 58.2], tier 270.0±3.8 · T370.0% [66.0, 73.7], tier 321.0±7.3 · T221.0% [14.6, 29.1], tier 249.4±1.2 · T349.4% [48.2, 50.6], tier 352.6±3.0 · T452.6% [49.6, 55.6], tier 477.0±2.7 · T277.0% [74.2, 79.5], tier 240.0±4.3 · T240.0% [35.8, 44.4], tier 221.5±2.5 · T521.5% [19.0, 24.0], tier 547.1±3.0 · T347.1% [44.1, 50.1], tier 349.7±3.0 · T249.7% [46.8, 52.8], tier 229.8±2.8 · T229.8% [27.0, 32.6], tier 277.4*±2.077.4% [75.4, 79.4], not ranked18.3±1.6 · T518.3% [16.7, 19.9], tier 547.8±2.8 · T247.8% [44.9, 50.5], tier 253.7±6.8 · T253.7% [46.9, 60.6], tier 2
Tier 3
Kev-0.6B25.4±9.6 · T325.4% [16.9, 36.1], tier 365.3±4.1 · T365.3% [61.1, 69.2], tier 39.6±5.9 · T29.6% [7.2, 19.1], tier 254.1±2.4 · T254.1% [51.8, 56.6], tier 255.8±2.7 · T355.8% [53.0, 58.5], tier 389.5*±2.089.5% [87.5, 91.5], not ranked78.2*±3.778.2% [74.4, 81.8], not ranked44.0±3.1 · T344.0% [40.8, 47.0], tier 359.3±3.0 · T259.3% [56.3, 62.3], tier 241.6±3.0 · T341.6% [38.6, 44.7], tier 326.2±2.7 · T326.2% [23.4, 28.9], tier 344.8±2.5 · T244.8% [42.3, 47.4], tier 224.8±1.9 · T424.8% [22.9, 26.7], tier 445.3±2.8 · T245.3% [42.5, 48.1], tier 260.6±6.6 · T260.6% [53.8, 67.1], tier 2
Not ranked on this task
SimpleJev (Qwen3.5-0.8B)15.2±4.415.2% [10.6, 19.4], not ranked49.9±4.549.9% [45.4, 54.4], not ranked3.4±1.73.4% [1.8, 5.1], not ranked50.8±3.950.8% [46.8, 54.7], not ranked58.8±3.058.8% [55.9, 61.9], not ranked60.2±3.0 · T360.2% [57.2, 63.2], tier 30.0±0.00.0% [0.0, 0.0], not ranked29.1±2.8 · T429.1% [26.3, 31.9], tier 445.2±3.2 · T345.2% [42.0, 48.4], tier 332.7±3.0 · T432.7% [29.7, 35.7], tier 425.0±2.7 · T425.0% [22.3, 27.7], tier 439.0±1.9 · T339.0% [37.1, 40.8], tier 356.3±1.6 · T256.3% [54.7, 57.8], tier 225.0±2.525.0% [22.5, 27.4], not ranked55.4±7.1 · T255.4% [48.1, 62.3], tier 2
SemIf (Qwen3-0.6B)7.8±3.27.8% [5.1, 11.5], not ranked63.0±4.263.0% [58.7, 67.0], not ranked9.6±4.19.6% [5.1, 13.3], not ranked52.1±3.8 · T252.1% [48.2, 55.9], tier 253.5±3.053.5% [50.5, 56.5], not ranked0.0±0.00.0% [0.0, 0.0], not ranked0.0±0.00.0% [0.0, 0.0], not ranked42.6±3.1 · T342.6% [39.4, 45.7], tier 337.8±2.9 · T537.8% [35.0, 40.9], tier 523.9±2.7 · T523.9% [21.4, 26.7], tier 525.0±2.7 · T425.0% [22.3, 27.7], tier 435.4±2.3 · T435.4% [33.1, 37.7], tier 417.9±1.6 · T517.9% [16.4, 19.6], tier 526.2±2.526.2% [23.7, 28.7], not ranked48.9±7.1 · T348.9% [41.8, 55.9], tier 3
openJev Verdict 1.44.8±2.54.8% [2.8, 7.8], not ranked51.8±4.451.8% [47.5, 56.4], not ranked1.0±1.01.0% [0.2, 2.2], not ranked50.3±3.650.3% [46.8, 54.0], not ranked50.0*±0.050.0% [50.0, 50.0], not ranked78.4*±2.678.4% [75.8, 80.9], not ranked0.0*±0.00.0% [0.0, 0.0], not ranked16.4±2.3 · T616.4% [14.1, 18.7], tier 646.3±3.0 · T346.3% [43.2, 49.3], tier 326.5±2.7 · T526.5% [23.8, 29.2], tier 524.8±2.7 · T424.8% [22.1, 27.4], tier 436.4±2.4 · T336.4% [34.0, 38.8], tier 318.5±1.7 · T518.5% [16.8, 20.3], tier 522.6±2.422.6% [20.2, 25.0], not ranked57.6*±7.157.6% [50.2, 64.5], not ranked
Laya0.0±0.00.0% [0.0, 0.0], not ranked69.8±3.9 · T369.8% [65.8, 73.6], tier 315.8±11.5 · T215.8% [0.0, 23.0], tier 249.8±1.2 · T349.8% [48.6, 51.0], tier 354.3±2.8 · T454.3% [51.4, 57.1], tier 476.0±2.7 · T276.0% [73.3, 78.6], tier 242.4±4.3 · T242.4% [38.2, 46.8], tier 222.9±2.5 · T522.9% [20.2, 25.3], tier 544.2±2.9 · T444.2% [41.3, 47.2], tier 444.8±2.9 · T344.8% [41.9, 47.8], tier 331.5±2.8 · T231.5% [28.7, 34.3], tier 236.1±2.5 · T336.1% [33.6, 38.6], tier 317.8±1.5 · T517.8% [16.3, 19.4], tier 545.9±2.8 · T245.9% [43.1, 48.8], tier 258.9±6.4 · T258.9% [52.4, 65.2], tier 2

Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.

Jev

Provider
TypeSafe
Weights
Closed (API)
Size
Not disclosed
Licence
proprietary
TaskScore [95% interval]TierLatency p50Cost / 1k
Prompt injection94.6% [92.2, 96.7]1, ahead323 ms (Hosted API)n/a
Moderation90.3% [87.9, 92.4]1, ahead322 ms (Hosted API)n/a
PII90.8% [88.1, 93.8]1, ahead316 ms (Hosted API)n/a
RAG faithfulness80.3% [77.5, 83.1]1, ahead314 ms (Hosted API)n/a
Off-topic93.4% [91.8, 94.9]1, ahead316 ms (Hosted API)n/a
Routing, 20 intents89.1% [87.0, 91.0]1, ahead316 ms (Hosted API)n/a
Routing, 77 intents79.6% [76.0, 83.0]1, ahead318 ms (Hosted API)n/a
Tool routing78.3% [75.7, 80.7]1, ahead320 ms (Hosted API)n/a
Complaint routing78.7% [76.3, 81.2]1, ahead311 ms (Hosted API)n/a
Commit type68.3% [65.2, 71.1]1, ahead307 ms (Hosted API)n/a
Search relevance57.7% [54.7, 60.8]1, ahead312 ms (Hosted API)n/a
Typed decisions73.9% [71.8, 76.0]1, ahead323 ms (Hosted API)n/a
Web-agent actions70.8% [68.9, 72.6]1, ahead321 ms (Hosted API)n/a
Community sets62.3% [59.5, 65.1]1, ahead320 ms (Hosted API)n/a
JevBench public85.7% [81.0, 90.1]1, ahead326 ms (Hosted API)n/a
  • Training data not disclosed, so contamination can't be ruled out.
How we measured

Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the metric difference, on the items every model answered.

Tie tiers

Models in one tier can't be separated by the tie test (paired bootstrap of the TPR@5%FPR difference on shared items, 95% CI includes 0; paired bootstrap of the AUROC difference on shared items, 95% CI includes 0; paired bootstrap of the Balanced acc. difference on shared items, 95% CI includes 0; paired McNemar on per-item correctness at the dev threshold, p >= 0.05; paired McNemar on per-item correctness, p >= 0.05). Tiers are computed in the export over every model; filters on this page never change them.

Contamination

Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner.

Latency and hardware

  • Laptop GPU: NVIDIA GeForce RTX 3050 Ti Laptop GPU, 4 GB, WSL2, concurrency 1
  • Hosted API: the same laptop, concurrency 1, latency includes the network round trip

Latencies are only compared within one hardware tier.

Datasets

  • Prompt injection (1,000 test items): deepset/prompt-injections (Apache-2.0), jackhhao/jailbreak-classification (Apache-2.0), reshabhs/SPML_Chatbot_Prompt_Injection (MIT), djapp18/JailbreaksOverTime (CC-BY-4.0)
  • Moderation (667 test items): mmathys/openai-moderation-api-evaluation (MIT), nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (CC-BY-4.0)
  • PII (1,000 test items): gretelai/gretel-pii-masking-en-v1 (Apache-2.0), beki/privy (MIT), nvidia/Nemotron-PII (CC-BY-4.0)
  • RAG faithfulness (666 test items): pminervini/HaluEval (Apache-2.0)
  • Off-topic (1,000 test items): clinc/clinc_oos (CC-BY-3.0), mteb/amazon_massive_intent (CC-BY-4.0), benayas/snips (CC0-1.0)
  • Routing, 20 intents (1,000 test items): legacy-datasets/banking77 (CC-BY-4.0)
  • Routing, 77 intents (500 test items): legacy-datasets/banking77 (CC-BY-4.0)
  • Tool routing (1,000 test items): gorilla-llm/Berkeley-Function-Calling-Leaderboard (Apache-2.0)
  • Complaint routing (1,000 test items): BEE-spoke-data/consumer-finance-complaints (CC0-1.0)
  • Commit type (1,000 test items): github.com/angular/angular (MIT), github.com/vitejs/vite (MIT)
  • Search relevance (1,000 test items): tasksource/esci (Apache-2.0)
  • Typed decisions (1,965 test items): LocalLLaMA/typed-decisions (Apache-2.0)
  • Web-agent actions (2,329 test items): AndeyTait/JevForge-Mind2Web (CC BY 4.0)
  • Community sets (1,200 test items): Praveenrajus/jev-bench (CC BY 4.0)
  • JevBench public (231 test items): fstandhartinger/jevbench (MIT)
  • Trained on unnamed jailbreak and prompt-injection data; deepset/prompt-injections is declared held out.
  • Its confidence calibrator (not the model weights) was fitted on prompts from two of this category's sources.
  • Verdict 1.4 returns a near-constant probability here: the middle 90% of its answers span 0.12 around 0.5, and on off-topic it returned one identical value for all 1,000 items, so this AUROC ranks numerical noise rather than measuring the model. Reading it upside down would not give a real score either. Scoring polarity was checked against the author's own reference adapter and is correct (harness investigation, 2026-09-23).
  • Trained on unnamed human-labelled toxicity data, a closely related task.
  • Its base model's training lineage includes toxicity and hate-speech datasets, a closely related task.
  • This zero-shot readout saturates at its lowest rating bin (every probability between 0.0102 and 0.0143) and answers 'no PII' to all 1,000 documents. The residual ordering runs backwards because our PII negatives are redacted documents whose stand-in phrases name the very categories the questions ask about ('the address on file', 'the customer'), which is a documented property of the dataset. The score is a dataset artefact, not an inverted signal (harness investigation, 2026-09-23).
  • Trained on MNLI and BoolQ, closely related entailment and passage-QA tasks (no item overlap found).
  • Its base model's training lineage includes natural-language-inference datasets, a closely related task.
  • Trained on NLI and fact-verification data (unnamed), closely related to faithfulness checking.
  • Laya does carry signal on one of this framing's three sub-questions (contradicts: AUROC 61 / 64), but that question also returns the lowest probabilities, so the pre-registered max-combine rule almost always hands the item's score to one of the two uninformative sub-questions instead. The result is a combined score at or slightly below chance for both Laya and Laya-typed. This is a limitation of the max-combine framing on this model, not a polarity error; the single-question 'broad' framing is also at chance for Laya here (harness investigation, 2026-09-23).
  • Trained on unnamed intent-classification data, the task family of this category's sources.
  • Trained on CLINC150's out-of-scope queries, the source of a third of this category, and calibrated on CLINC test queries; shown for reference, not ranked.
  • Trained on Banking77, the dataset behind this category (train split), and its checkpoint was picked on a set holding 27 of these 1,000 test messages; shown for reference, not ranked.
  • Trained on Banking77, the dataset behind this category (train split), and calibrated on Banking77 test messages; shown for reference, not ranked.
  • Trained on unnamed banking-intent data; its authors list Banking77 itself as held out.
  • Option-definition bug: all 13 routing-20 items whose gold intent is get_physical_card are PIN questions, which contradicts our written definition of that option ('the customer asks about getting a physical card'). Jev, which reads only the definition, gets 1 of 13; Kev-0.8B, trained on Banking77, gets 13 of 13. beneficiary_not_allowed has the same problem. A memorised dataset convention shows up as accuracy. With the affected items and the adjudicated label errors removed, routing-20's first tier is a three-way tie of Kev-0.8B, Jev and Kev-0.6B, and Jev joins routing-77's first tier (data-quality audit, 2026-09-22). The audit itself warns that dropping flagged items favours Jev by construction, because the flagged items are the ones Jev confidently missed, so read that re-scored tier as a bound, not a result.
  • 27 routing-20 and 18 routing-77 messages also appear verbatim in Kev's own locked test set.
  • Trained on Banking77, the dataset behind this category (train split), and its checkpoint was picked on a set holding 16 of these 500 test messages; shown for reference, not ranked.
  • Trained on support-ticket queue routing (Tobi-Bueck/customer-support-tickets), a closely related task.
  • Trained on MS MARCO query-passage relevance, a closely related task.
  • Fine-tuned on this dataset's training split (same generator, workflows and questions); shown for reference, not ranked.
  • Trained on unnamed toxicity, rubric-rating and fact-checking data whose descriptions match four of this suite's eight configs; possibly the same datasets.
  • Trained on MNLI, a task close to this suite's fact-checking config (150 of 1,200 items).
  • Its base model's lineage includes toxicity, hate-speech and NLI data, close to three of this suite's eight configs.
  • Trained on template-generated policy-rule cases like this set's policy items (no item overlap found).
  • openJev Verdict 1.4's inference settings were chosen by measuring on these same 231 public items; shown for reference, not ranked.

Release 2026-09-23.1, benchmark commit 58b54bf, 34 input files, integrity sha256-leaves-v1.