Prompt injection Catch rate at 5% false alarms
- Jev
- 94.6%
- Best open
- Eikos-27B 94.2%
Tied, 0.4 points apart
Out of distribution, open models up to 28B tie Jev on 8 of 11 tasks, lead on 2 and trail on off-topic.
This benchmark was created using LangWatch. Sign up to create your ownLangWatch Instant Evals runs on Jev. Check it out
Tied, 0.4 points apart
Jev against the best open model not trained on the task’s data; tied means the tie test cannot separate them.
Tied, 0.4 points apart
Tied, <0.1 points apart
Eikos-27B ahead by 4.4 points
Tied, 1.8 points apart
Jev ahead by 3.8 points
Tied, 0.6 points apart
Tied, 1.0 points apart
Shisa DE-1 ahead by 3.0 points
Tied, 0.7 points apart
Tied, 1.5 points apart
Tied, 1.1 points apart
Eikos-27B has the highest open score on 6 of 11 tasks.
* Trained on that task’s dataset: shown for reference, never ranked. Models are marked on what their authors disclosed; no declaration is not the same as not trained on it. A model marked * scores higher than Jev on moderation; routing, 20 intents; and routing, 77 intents.
The top 5 models per task, plus Jev. Overlapping lines mean the data cannot separate two models.
Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.
Tied, 0.4 points apart
| Model | Coverage | Cut inputs | Flips | p95 ms (GPU) |
|---|---|---|---|---|
| Jev | 100% | 0% | 2/300 | 841 API |
| Eikos-27B | 100% | 0% | 0/300 | 170 H100 |
| Shisa DE-1 | 100% | n/a | 0/300 | 60 H100 |
| AutoJev-27B | 100% | n/a | 0/300 | 198 H100 |
| Kev-4B | 100% | n/a | 0/300 | 260 L4 |
Flips: answers that changed between two identical runs on the first 300 items.
Without spml (750 items), from the data‑quality audit: SPML is separable by length alone (length-only AUROC 0.998): its benign items are short questions and its injections long persona prompts. Without it, the best open model changes.
Primary score per task and model, in percent. Deeper orange is better; gray is not ranked.
| Task | Jev | Eikos-27B | Shisa DE-1 | AutoJev-27B | SemIf (Qwen3.5-4B) | Eikos-4B | Kev-9B | Kev-4B | GLiNER2.5-Decide | Kev-0.8B | Laya-typed | Laya | Kev-0.6B | SimpleJev (Qwen3.5-0.8B) | SemIf (Qwen3-0.6B) | openJev Verdict 1.4 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Prompt injectionCatch rate at 5% false alarms, tied | 95: 94.6%, tied with the leader | 94: 94.2%, tied with the leader | 94: 94.0%, tied with the leader | 92: 92.0%, tied with the leader | 91: 90.6% | 91: 90.8% | 88: 87.6% | 92: 91.6% | 63: 63.2% | 56: 56.2% | 49: 49.2% | 0‡: 0.0%, threshold unreachable | 25: 25.4% | 15‡: 15.2%, degenerate score | 8‡: 7.8%, degenerate score | 5*‡: 4.8%, trained on this dataset, degenerate score |
| ModerationAUROC, tied | 90: 90.3%, tied with the leader | 90: 90.3%, tied with the leader | 88: 88.3% | 91*: 90.6%, trained on this dataset | 89: 88.7%, tied with the leader | 89: 88.8%, tied with the leader | 89: 88.5% | 82: 82.0% | 72: 71.5% | 76: 76.2% | 70: 70.0% | 70: 69.8% | 65: 65.3% | 50‡: 49.9%, degenerate score | 63‡: 63.0%, degenerate score | 52‡: 51.8%, degenerate score |
| PIICatch rate at 5% false alarms | 91: 90.8% | 95: 95.2%, ahead | 84: 84.2% | 93: 92.8%, tied with the leader | 66: 66.1% | 79: 78.6% | 63: 62.5% | 9: 9.0% | 17: 16.8% | 23: 22.8% | 21: 21.0% | 16: 15.8% | 10: 9.6% | 3‡: 3.4%, degenerate score | 10‡: 9.6%, degenerate score | 1‡: 1.0%, degenerate score |
| RAG faithfulnessBalanced accuracy, tied | 80: 80.3%, tied with the leader | 82: 82.1%, tied with the leader | 79: 79.2% | 79: 79.1% | 72: 72.0% | 71: 71.0% | 73: 73.1% | 72: 72.5% | 58‡: 57.8%, degenerate score | 55: 55.5% | 49: 49.4% | 50: 49.8% | 54: 54.1% | 51‡: 50.8%, degenerate score | 52: 52.1% | 50‡: 50.3%, degenerate score |
| Off-topicBalanced accuracy | 93: 93.4%, ahead | 91*: 90.6%, trained on this dataset | 88: 88.0% | 93*: 92.7%, trained on this dataset | 81: 81.1% | 80*: 79.6%, trained on this dataset | 90: 89.6% | 84: 84.4% | 49‡: 49.3%, degenerate score | 77: 76.8% | 53: 52.6% | 54: 54.3% | 56: 55.8% | 59‡: 58.8%, degenerate score | 53‡: 53.5%, degenerate score | =*: constant output |
| Routing, 20 intentsAccuracy, tied | 89: 89.1%, tied with the leader | 90*: 89.6%, trained on this dataset | 88: 88.1% | 90: 89.7%, tied with the leader | –: not scored | 88*: 87.7%, trained on this dataset | 93*: 93.0%, trained on this dataset | 93*: 93.0%, trained on this dataset | 86: 86.0% | 91*: 91.3%, trained on this dataset | 77: 77.0% | 76: 76.0% | 90*: 89.5%, trained on this dataset | 60: 60.2% | –: not scored | 78*: 78.4%, trained on this dataset |
| Routing, 77 intentsAccuracy, tied | 80: 79.6%, tied with the leader | 78*: 77.8%, trained on this dataset | –: not scored | 81: 80.6%, tied with the leader | –: not scored | 74*: 74.0%, trained on this dataset | 85*: 85.0%, trained on this dataset | 85*: 85.0%, trained on this dataset | 71: 70.6% | 83*: 83.0%, trained on this dataset | 40: 40.0% | 42: 42.4% | 78*: 78.2%, trained on this dataset | –: not scored | –: not scored | –*: not scored |
| Tool routingAccuracy | 78: 78.3% | 79: 79.3% | 81: 81.3%, ahead | 81: 80.6%, tied with the leader | 80: 79.9%, tied with the leader | 76: 75.5% | 72: 72.4% | 71: 71.4% | 32: 32.4% | 58: 57.7% | 22: 21.5% | 23: 22.9% | 44: 44.0% | 29: 29.1% | 43: 42.6% | 16: 16.4% |
| Complaint routingAccuracy, tied | 79: 78.7%, tied with the leader | 76: 75.6% | 78: 78.0%, tied with the leader | 76: 75.9% | 73: 72.9% | –: no runno run | 75: 75.0% | 76: 76.2% | 53: 52.6% | 39: 38.5% | 47: 47.1% | 44: 44.2% | 59: 59.3% | 45: 45.2% | 38: 37.8% | 46: 46.3% |
| Commit typeAccuracy, tied | 68: 68.3%, tied with the leader | 67: 66.8%, tied with the leader | 65: 65.1% | 67: 66.7%, tied with the leader | 61: 61.2% | 57: 56.7% | 56: 55.5% | 54: 53.7% | 50: 50.4% | 49: 48.6% | 50: 49.7% | 45: 44.8% | 42: 41.6% | 33: 32.7% | 24: 23.9% | 27: 26.5% |
| Search relevanceAccuracy, tied | 58: 57.7%, tied with the leader | 57: 56.6%, tied with the leader | 50: 50.3% | 53: 52.8% | 42: 41.7% | 43: 43.3% | 43: 43.4% | 45: 44.8% | 25: 25.0% | 29: 29.3% | 30: 29.8% | 32: 31.5% | 26: 26.2% | 25: 25.0% | 25: 25.0% | 25: 24.8% |
| General suites, not counted in the title | ||||||||||||||||
| Typed decisionsAccuracy, tied | 74: 73.9%, tied with the leader | 73: 73.4%, tied with the leader | 75: 75.1%, tied with the leader | 74: 73.8%, tied with the leader | 63: 63.0% | 65: 65.3% | –: no runno run | 66: 65.8% | 49: 48.8% | 45: 44.9% | 77*: 77.4%, trained on this dataset | 36: 36.1% | 45: 44.8% | 39: 39.0% | 35: 35.4% | 36: 36.4% |
| Web-agent actionsAccuracy | 71: 70.8%, ahead | 68: 68.4% | 65: 65.2% | 68: 68.2% | 60: 59.5% | 63: 62.9% | –: no runno run | 58: 58.0% | 37: 37.0% | 51: 50.9% | 18: 18.3% | 18: 17.8% | 25: 24.8% | 56: 56.3% | 18: 17.9% | 19: 18.5% |
| Community setsAccuracy, tied | 62: 62.3%, tied with the leader | 60: 59.8% | 38: 38.2% | 63: 63.0%, tied with the leader | 37: 37.3% | 56: 55.8% | –: no runno run | 56: 55.8% | 41: 41.4% | 46: 46.4% | 48: 47.8% | 46: 45.9% | 45: 45.3% | 25: 25.0% | 26: 26.2% | 23: 22.6% |
Frozen test sets, 95% bootstrap intervals, a paired tie test and a contamination audit.
Each use case has a 60‑item dev split and a test split of up to 1,000 items, frozen with checksums before any model ran.
Every model got each task in two or three wordings; the wording and the decision threshold were picked on the dev split, once, then frozen.
Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The general suites use 10,000 resamples, clustered so the variants of one item move together. The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the metric difference, on the items every model answered.
Open models ran one at a time on a laptop GPU, an L4 or an H100, so latency compares only on the same GPU; every p95 names its GPU. Jev’s latency is measured from a client and includes the network. Cost is not shown on this release.
Inputs over each set’s length limit (2,500 to 12,000 characters) were removed for every model alike before sampling. Inputs a model cuts at its own token window stay in the score.
Each model and task pair was audited against the model’s published training, selection and calibration data. Grayed rows trained on the test items or on the same dataset; footnoted rows trained on a related task. Models are marked on what their authors disclosed; no declaration is not the same as not trained on it.
Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner. Grayed rows are shown for reference, never ranked.
Jev’s training data is not disclosed, and every open model starts from a pretrained base whose corpus is not itemised. An unmarked row means no known contamination, not proof that there is none.
Shisa DE-1 declares no training data and is grayed on no task; openJev Verdict 1.4 disclosed its data and is grayed on 4 of 14.
Where a grayed model has the best open score, counted:
No trivial baseline beats the best open model on any of the 11 tasks that have one. The notes below say what a better set would change.
Tasks from public datasets, each with its licence. Two tasks are scored without a source licensed for non‑commercial use, and two internal test sets are not shown.
| Task | Items | Sources | Licence |
|---|---|---|---|
| Prompt injection | 1,000 | deepset/prompt-injections, jackhhao/jailbreak-classification, reshabhs/SPML_Chatbot_Prompt_Injection, djapp18/JailbreaksOverTime | Apache-2.0, MIT, CC-BY-4.0 |
| Moderationwithout toxicchat | 667 | mmathys/openai-moderation-api-evaluation, nvidia/Aegis-AI-Content-Safety-Dataset-2.0 | MIT, CC-BY-4.0 |
| PII | 1,000 | gretelai/gretel-pii-masking-en-v1, beki/privy, nvidia/Nemotron-PII | Apache-2.0, MIT, CC-BY-4.0 |
| RAG faithfulnesswithout halueval-dialogue | 666 | pminervini/HaluEval | Apache-2.0 |
| Off-topic | 1,000 | clinc/clinc_oos, mteb/amazon_massive_intent, benayas/snips | CC-BY-3.0, CC-BY-4.0, CC0-1.0 |
| Routing, 20 intents | 1,000 | legacy-datasets/banking77 | CC-BY-4.0 |
| Routing, 77 intents | 500 | legacy-datasets/banking77 | CC-BY-4.0 |
| Tool routing | 1,000 | gorilla-llm/Berkeley-Function-Calling-Leaderboard | Apache-2.0 |
| Complaint routing | 1,000 | BEE-spoke-data/consumer-finance-complaints | CC0-1.0 |
| Commit type | 1,000 | github.com/angular/angular, github.com/vitejs/vite | MIT |
| Search relevance | 1,000 | tasksource/esci | Apache-2.0 |
| Typed decisions | 1,965 | LocalLLaMA/typed-decisions | Apache-2.0 |
| Web-agent actions | 2,329 | AndeyTait/JevForge-Mind2Web | CC BY 4.0 |
| Community sets | 1,200 | Praveenrajus/jev-bench | CC BY 4.0 |
Release 2026-09-25.1, benchmark commit c372cd5. All 235 runs behind these numbers were made from a working tree with uncommitted harness changes, so the exact code of those runs is not reproducible from a commit alone.
29 cloud‑GPU result rows were completed from the GPU box’s own reports after a copy error cut the local records short; on the 17,553 items both copies hold, they agree on every one.
Verified against the model cards and the dataset licences on 2026-09-25. Spot something outdated? Tell us.
Instant Evals runs on Jev, so we wanted to know how far open models are from it before recommending either. We made real judgment calls along the way (the contamination rule, the framings, the decision thresholds, which caveats to show), and every one of them is written down on this page so you can check it. What we did not touch by hand is the ranking itself: the tie test decides who is ahead.
On each task the tie test decides: one side is ahead, or they are tied when the data cannot separate them. Grayed rows were trained on the task’s data; they are shown for reference and never counted.
The model was trained on the test items or on another split of the same dataset, so its score measures memory of that dataset more than the skill. The contamination audit graded every model and task pair; the rule and the evidence are in the method section.
Models are marked on what their authors disclosed. A model that declares nothing has nothing to mark, and that is not proof it never saw the data.
Two sources are licensed for non-commercial use only, so the moderation and RAG numbers here are recomputed without them. Internal test sets are not shown at all. The audit also names sources that a trivial classifier solves; those stay in the headline, with the number without them next to it.
Only on the same GPU. Open models ran on a laptop GPU, an L4 or an H100, and each latency names its GPU. Jev’s includes the network round trip to its API.
Partly. Instant Evals runs the same hosted Jev judge over your own production history from the LangWatch CLI, so that half is the same judge. The open models are published weights you can serve yourself, but not with this benchmark’s exact setup: the harness, the framings and the decision thresholds used here are not public yet.
Instant Evals asks one question of your whole production history from the CLI.