Prompt injection Catch rate at 5% false alarms
- Jev
- 94.6%
- Best open
- Kev-0.8B 56.2%
- Gap
- leads by 38.4 pts
- No-model baseline
- Bag of words (TF-IDF) 72.6%: beats every open model
On 15 tasks, out of distribution, Jev leads the best open model under 1B by 12 to 68 points.
LangWatch Instant Evals runs on Jev. Check it outThis benchmark was created using LangWatch. Sign up to create your own
Jev’s lead over the best open model not trained on the task’s data, per task.
* Trained on that task’s dataset: shown for reference, never ranked. A model marked * scores higher than Jev on Routing, 20 intents; Routing, 77 intents; and Typed decisions.
No-model baseline: a trivial classifier with no language model (bag of words, string overlap or a lexical matcher), never ranked. It beats every open model on 9 of 11 tasks that have one.
Overlapping lines mean the data cannot separate two models.
Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.
| Model | Coverage | Cut inputs | Flips | p95 ms | $ / 1k |
|---|---|---|---|---|---|
| Jev | 100% | 0% | 2/300 | 841 | not shown |
| Kev-0.8B | 100% | n/a | 0/300 | 230 | $0.010* |
| Kev-0.6B | 100% | n/a | 0/300 | 171 | $0.0073* |
| Laya-typed | 100% | 5% | 0/300 | 202 | $0.0080* |
| Laya | 100% | 16% | 0/300 | 154 | $0.0086* |
| SimpleJev (Qwen3.5-0.8B) | 100% | n/a | 0/300 | 298 | $0.019* |
| openJev Verdict 1.4 | 100% | n/a | 0/300 | 57 | $0.0043* |
| SemIf (Qwen3-0.6B) | 100% | n/a | 0/300 | 141 | $0.0061* |
Flips: answers that changed between two identical runs on the first 300 items.
Without spml (750 items), from the data‑quality audit: SPML is separable by length alone (length-only AUROC 0.998): its benign items are short questions and its injections long persona prompts. Without it, the best open model changes.
Primary score per task, in percent. Darker is better.
| Task | Jev | Kev-0.8B | Kev-0.6B | Laya-typed | Laya | SimpleJev (Qwen3.5-0.8B) | openJev Verdict 1.4 | SemIf (Qwen3-0.6B) |
|---|---|---|---|---|---|---|---|---|
| Prompt injectionCatch rate at 5% false alarms | 95: 94.6%, top tier | 56: 56.2% | 25: 25.4% | 49: 49.2% | 0: 0.0% | 15: 15.2% | 5: 4.8% | 8: 7.8% |
| ModerationAUROC | 90: 90.3%, top tier | 76: 76.2% | 65: 65.3% | 70: 70.0% | 70: 69.8% | 50: 49.9% | 52: 51.8% | 63: 63.0% |
| PIICatch rate at 5% false alarms | 91: 90.8%, top tier | 23: 22.8% | 10: 9.6% | 21: 21.0% | 16: 15.8% | 3: 3.4% | 1: 1.0% | 10: 9.6% |
| RAG faithfulnessBalanced accuracy | 80: 80.3%, top tier | 55: 55.5% | 54: 54.1% | 49: 49.4% | 50: 49.8% | 51: 50.8% | 50: 50.3% | 52: 52.1% |
| Off-topicBalanced accuracy | 93: 93.4%, top tier | 77: 76.8% | 56: 55.8% | 53: 52.6% | 54: 54.3% | 59: 58.8% | =*: constant output | 53: 53.5% |
| Routing, 20 intentsAccuracy | 89: 89.1%, top tier | 91*: 91.3%, trained on this dataset | 90*: 89.5%, trained on this dataset | 77: 77.0% | 76: 76.0% | 60: 60.2% | 78*: 78.4%, trained on this dataset | –: not scored |
| Routing, 77 intentsAccuracy | 80: 79.6%, top tier | 83*: 83.0%, trained on this dataset | 78*: 78.2%, trained on this dataset | 40: 40.0% | 42: 42.4% | –: not scored | –*: not scored | –: not scored |
| Tool routingAccuracy | 78: 78.3%, top tier | 58: 57.7% | 44: 44.0% | 22: 21.5% | 23: 22.9% | 29: 29.1% | 16: 16.4% | 43: 42.6% |
| Complaint routingAccuracy | 79: 78.7%, top tier | 39: 38.5% | 59: 59.3% | 47: 47.1% | 44: 44.2% | 45: 45.2% | 46: 46.3% | 38: 37.8% |
| Commit typeAccuracy | 68: 68.3%, top tier | 49: 48.6% | 42: 41.6% | 50: 49.7% | 45: 44.8% | 33: 32.7% | 27: 26.5% | 24: 23.9% |
| Search relevanceAccuracy | 58: 57.7%, top tier | 29: 29.3% | 26: 26.2% | 30: 29.8% | 32: 31.5% | 25: 25.0% | 25: 24.8% | 25: 25.0% |
| Typed decisionsAccuracy | 74: 73.9%, top tier | 45: 44.9% | 45: 44.8% | 77*: 77.4%, trained on this dataset | 36: 36.1% | 39: 39.0% | 36: 36.4% | 35: 35.4% |
| Web-agent actionsAccuracy | 71: 70.8%, top tier | 51: 50.9% | 25: 24.8% | 18: 18.3% | 18: 17.8% | 56: 56.3% | 19: 18.5% | 18: 17.9% |
| Community setsAccuracy | 62: 62.3%, top tier | 46: 46.4% | 45: 45.3% | 48: 47.8% | 46: 45.9% | 25: 25.0% | 23: 22.6% | 26: 26.2% |
| JevBench publicAccuracy | 86: 85.7%, top tier | 60: 60.2% | 61: 60.6% | 54: 53.7% | 59: 58.9% | 55: 55.4% | 58*: 57.6%, trained on this dataset | 49: 48.9% |
Frozen test sets, 95% bootstrap intervals, a paired tie test and a contamination audit.
Each use case has a 60‑item dev split and a 1,000‑item test split (500 for 77‑way routing), frozen with checksums before any model ran.
Every model got each task in two or three wordings; the wording and the decision threshold were picked on the dev split, once, then frozen.
Scores carry a 95% percentile bootstrap interval over items (2,000 resamples). Tiers come from a paired McNemar test on accuracy, or a paired bootstrap of the metric difference on the items every model shared.
Open models ran one at a time on a laptop GPU (NVIDIA RTX 3050 Ti, 4 GB), so their latency is laptop tier. Jev’s latency is measured from the same laptop and includes the network. Local cost is $0.35 per GPU‑hour times measured time. Jev has the highest p95 on 10 of 11 real tasks.
Inputs over each set’s length limit (2,500 to 12,000 characters) were removed for every model alike before sampling. Inputs a model cuts at its own token window stay in the score; the share of cut inputs is shown next to coverage.
Each model and task pair was audited against the model’s published training, selection and calibration data. Grayed rows trained on the test items or on the same dataset; footnoted rows trained on a related task.
Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner. Grayed rows are shown for reference, never ranked.
Jev’s training data is not disclosed, and every open model starts from a pretrained base whose corpus is not itemised. An unmarked row means no known contamination, not proof that there is none.
A trivial baseline fitted on the test items beats every open model on 9 of 11 tasks that have one. The notes below say what a better set would change.
15 tasks from public datasets, each with its licence. Two sources licensed for non‑commercial use are left out of the published numbers, and two internal test sets are not shown.
| Task | Items | Sources | Licence |
|---|---|---|---|
| Prompt injection | 1,000 | deepset/prompt-injections, jackhhao/jailbreak-classification, reshabhs/SPML_Chatbot_Prompt_Injection, djapp18/JailbreaksOverTime | Apache-2.0, MIT, CC-BY-4.0 |
| Moderationwithout toxicchat | 667 | mmathys/openai-moderation-api-evaluation, nvidia/Aegis-AI-Content-Safety-Dataset-2.0 | MIT, CC-BY-4.0 |
| PII | 1,000 | gretelai/gretel-pii-masking-en-v1, beki/privy, nvidia/Nemotron-PII | Apache-2.0, MIT, CC-BY-4.0 |
| RAG faithfulnesswithout halueval-dialogue | 666 | pminervini/HaluEval | Apache-2.0 |
| Off-topic | 1,000 | clinc/clinc_oos, mteb/amazon_massive_intent, benayas/snips | CC-BY-3.0, CC-BY-4.0, CC0-1.0 |
| Routing, 20 intents | 1,000 | legacy-datasets/banking77 | CC-BY-4.0 |
| Routing, 77 intents | 500 | legacy-datasets/banking77 | CC-BY-4.0 |
| Tool routing | 1,000 | gorilla-llm/Berkeley-Function-Calling-Leaderboard | Apache-2.0 |
| Complaint routing | 1,000 | BEE-spoke-data/consumer-finance-complaints | CC0-1.0 |
| Commit type | 1,000 | github.com/angular/angular, github.com/vitejs/vite | MIT |
| Search relevance | 1,000 | tasksource/esci | Apache-2.0 |
| Typed decisions | 1,965 | LocalLLaMA/typed-decisions | Apache-2.0 |
| Web-agent actions | 2,329 | AndeyTait/JevForge-Mind2Web | CC BY 4.0 |
| Community sets | 1,200 | Praveenrajus/jev-bench | CC BY 4.0 |
| JevBench public | 231 | fstandhartinger/jevbench | MIT |
Release 2026-09-23.1, benchmark commit 58b54bf. 120 of the runs behind these numbers were made from a working tree with uncommitted harness changes, so the exact code of those runs is not reproducible from a commit alone.
Jev’s cost per 1,000 decisions is not shown on this release.
Verified against the model cards, the dataset licences and the public list price on 2026-09-22. Spot something outdated? Tell us.
Instant Evals runs on Jev, so we wanted to know how far the open Jev-class models are from it before recommending either. We made real judgment calls along the way (the contamination rule, the framings, the decision thresholds, which caveats to show), and every one of them is written down on this page so you can check it. What we did not touch by hand is the ranking itself: the tie test decides who leads on the numbers.
Each task shows its own leader, and two models in the same tier cannot be separated by the data. Grayed rows were trained on the task’s dataset; they are shown for reference and are not counted as the best open model.
The model was trained on the test items or on another split of the same dataset, so its score measures memory of that dataset more than the skill. The contamination audit graded every model and task pair; the rule and the evidence are in the method section.
Two sources are licensed for non-commercial use only, so the moderation and RAG numbers here are recomputed without them. Two internal test sets are not shown at all. The audit also names sources that a trivial classifier solves; those stay in the headline, with the number without them next to it.
Only within its kind. Open models ran one at a time on a laptop GPU, so their latency is laptop tier and a data-centre GPU is faster. Jev’s latency is measured from the same laptop and includes the network round trip to its API.
For the open models it is an assumption: a cloud-equivalent GPU rate times the measured time at concurrency 1, marked as assumed wherever it appears. Jev’s cost is not shown on this release: it is confidential under the vendor’s customer agreement until legal clears it for a public page.
Partly. Instant Evals runs the same hosted Jev judge over your own production history from the LangWatch CLI, so that half is the same judge. The open models are published weights you can serve yourself, but not with this benchmark’s exact setup: the harness, the framings and the decision thresholds used here are not public yet.
Instant Evals asks one question of your whole production history from the CLI.