Prompt injection Balanced accuracy
- Jev
- 95.2%
- Best LLM
- Claude Haiku 4.5 92.7%
Jev ahead by 2.5 points
Per prompt: Jev 0.0026¢, Claude Haiku 4.5 0.061¢ (24 times)
Jev ties on one task and leads on one. The LLM ahead costs a median 34 times more per prompt.
This benchmark was created using LangWatch. Sign up to create your ownLangWatch Instant Evals runs on Jev. Check it outEvery frontier LLM call ran through the LangWatch AI Gateway: one key for all five models, cost and latency logged per call. See the gateway
Jev against the best frontier LLM on each task; tied means the tie test cannot separate them. Cost is per prompt: one decision, one call.
Jev ahead by 2.5 points
Per prompt: Jev 0.0026¢, Claude Haiku 4.5 0.061¢ (24 times)
Gemini 3.8 Flash ahead by 3.4 points
Per prompt: Jev 0.0020¢, Gemini 3.8 Flash 0.045¢ (22 times)
Claude Opus 5.5 ahead by 2.5 points
Per prompt: Jev 0.0023¢, Claude Opus 5.5 0.32¢ (141 times)
Claude Opus 5.5 ahead by 4.7 points
Per prompt: Jev 0.0029¢, Claude Opus 5.5 0.48¢ (169 times)
Claude Opus 5.5 ahead by 2.6 points
Per prompt: Jev 0.0036¢, Claude Opus 5.5 0.12¢ (34 times)
Claude Opus 5.5 ahead by 5.9 points
Per prompt: Jev 0.0034¢, Claude Opus 5.5 0.50¢ (147 times)
Claude Opus 5.5 ahead by 8.4 points
Per prompt: Jev 0.0093¢, Claude Opus 5.5 0.16¢ (17 times)
GPT-6 Luna ahead by 7.2 points
Per prompt: Jev 0.0042¢, GPT-6 Luna 0.0080¢ (1.9 times)
Tied, 0.9 points apart
Per prompt: Jev 0.0041¢, Gemini 3.8 Flash 0.083¢ (20 times)
Gemini 3.8 Flash ahead by 9.7 points
Per prompt: Jev 0.0025¢, Gemini 3.8 Flash 0.059¢ (24 times)
Claude Opus 5.5 ahead by 5.9 points
Per prompt: Jev 0.0027¢, Claude Opus 5.5 0.48¢ (174 times)
On prompt injection, moderation and PII the score is balanced accuracy for every model: the LLMs answer with one label, so AUROC and catch rates cannot be computed for them. Jev is scored at the decision threshold frozen on the dev split.
Every set here is public, and no frontier vendor discloses its training data (nor is Jev’s disclosed), so any model on this page may have seen these datasets. No row is grayed; read no score as proof of generalisation.
Every model on each task, with its latency and its cost per prompt in the same order. Overlapping lines mean the data cannot separate two models.
Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.
Balanced accuracy for every model: the LLMs answer with one label, so catch rate at 5% false alarms cannot be computed for them. Jev at its dev‑frozen threshold.
Jev ahead by 2.5 points
Balanced accuracy, 95% interval
Latency, p50 per prompt
Cost per prompt
p50 per prompt, p95 under it. Frontier LLMs: each call timed from Denmark through the LangWatch AI Gateway, 6 at a time, network included; orange is the fastest of them. Jev (~, hatched): server‑side, approximate, so on another basis and never orange. Cost per prompt: one prompt is one decision, one call to the model, at the recorded token usage and each vendor’s list price, cached input included; orange is the cheapest.
Score per task and model, in percent, shaded by rank within each task: darker is higher.
| Task | Jev | Claude Opus 5.5 | Gemini 3.8 Flash | GPT-5.6 Terra | GPT-6 Luna | Claude Haiku 4.5 |
|---|---|---|---|---|---|---|
| Prompt injectionBalanced accuracy | 95: 95.2%, highest score on this task | 91: 90.9%undefined | 90: 90.4%undefined | 88: 88.2%undefined | 88: 88.0%undefined | 93: 92.7%undefined |
| ModerationBalanced accuracy | 78: 78.0%undefined | 81: 81.1%undefined | 81: 81.4%, highest score on this task | 80: 79.9%undefined | 81: 80.8%undefined | 81: 81.1%undefined |
| PIIBalanced accuracy | 93: 92.9%undefined | 95: 95.4%, highest score on this task | 93: 93.1%undefined | 86: 86.4%undefined | 87: 87.2%undefined | 83: 83.1%undefined |
| RAG faithfulnessBalanced accuracy | 80: 80.3%undefined | 85: 85.0%, highest score on this task | 83: 83.3%undefined | 81: 80.6%undefined | 80: 79.7%undefined | 77: 77.4%undefined |
| Off-topicBalanced accuracy | 93: 93.4%undefined | 96: 96.0%, highest score on this task | 95: 95.0%undefined | 92: 92.3%undefined | 91: 90.6%undefined | 91: 91.5%undefined |
| Routing, 20 intentsAccuracy | 89: 89.1%undefined | 95: 95.0%, highest score on this task | 93: 92.7%undefined | 92: 92.4%undefined | 90: 90.3%undefined | 89: 89.3%undefined |
| Routing, 77 intentsAccuracy | 80: 79.6%undefined | 88: 88.0%, highest score on this task | 85: 84.8%undefined | 83: 83.4%undefined | 82: 82.2%undefined | 80: 80.0%undefined |
| Tool routingAccuracy | 78: 78.3%undefined | 79: 78.7%undefined | 84: 83.6%undefined | 84: 84.2%undefined | 86: 85.5%, highest score on this task | 82: 82.0%undefined |
| Complaint routingAccuracy, tied | 79: 78.7%undefined | 79: 78.9%undefined | 80: 79.6%, highest score on this task | 78: 78.4%undefined | 78: 77.9%undefined | 77: 77.1%undefined |
| Commit typeAccuracy | 68: 68.3%undefined | 77: 77.3%undefined | 78: 78.0%, highest score on this task | 67: 66.5%undefined | 68: 67.7%undefined | 68: 67.9%undefined |
| Search relevanceAccuracy | 58: 57.7%undefined | 64: 63.6%, highest score on this task | 63: 62.5%undefined | 58: 57.6%undefined | 55: 54.7%undefined | 51: 50.6%undefined |
The Jev benchmark’s frozen test sets, one prompt per decision, 95% bootstrap intervals and a paired tie test.
The 11 use cases of the Jev benchmark, with the same frozen dev and test splits (up to 1,000 test items each). The open models on the same sets.
Each LLM got the task and its options in the system message and the item last, and answered with one label through a JSON schema. The wording was frozen on the 60‑item dev split before the test run. Reasoning was off, or at the lowest setting each model accepts; some models still reason briefly at their lowest setting.
The frontier LLMs answer with one label, not a probability, so threshold-free metrics (AUROC, catch rate at a fixed false-alarm rate) cannot be computed for them. On moderation, PII and prompt injection every model is therefore compared on balanced accuracy, computed the same way for all: Jev at the decision threshold frozen on the dev split, each LLM at its label.
Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the difference in balanced accuracy, on the same items for every model; an empty answer counts as wrong.
Frontier LLMs: each call timed from Denmark through the LangWatch AI Gateway, 6 at a time, network and gateway included (p50 1,030 ms to 2,539 ms by task and model). Jev: approximate server‑side time (~, p50 ~59 ms to ~80 ms). The two bases differ, so the page names no fastest model across them.
Recorded token usage of the test run at each vendor’s public list price, cached input included where the provider applied it; Jev at its list price (input tokens only, output free). Shown per prompt; the export also carries it per 1,000 decisions. Jev is the cheapest on nine of 11 tasks; GPT-6 Luna costs less on off-topic and routing, 77 intents.
Every frontier LLM call went through the LangWatch AI Gateway, which logged the cost and latency of each call, with prompt caching on the shared part of long prompts.
A model is ranked when it answered at least 95% of a task’s items; an empty answer counts as wrong.
Tasks from public datasets, each with its licence. Two tasks are scored without a source licensed for non‑commercial use.
| Task | Items | Sources | Licence |
|---|---|---|---|
| Prompt injection | 1,000 | deepset/prompt-injections, jackhhao/jailbreak-classification, reshabhs/SPML_Chatbot_Prompt_Injection, djapp18/JailbreaksOverTime | Apache-2.0, MIT, CC-BY-4.0 |
| Moderationwithout toxicchat | 667 | mmathys/openai-moderation-api-evaluation, nvidia/Aegis-AI-Content-Safety-Dataset-2.0 | MIT, CC-BY-4.0 |
| PII | 1,000 | gretelai/gretel-pii-masking-en-v1, beki/privy, nvidia/Nemotron-PII | Apache-2.0, MIT, CC-BY-4.0 |
| RAG faithfulnesswithout halueval-dialogue | 666 | pminervini/HaluEval | Apache-2.0 |
| Off-topic | 1,000 | clinc/clinc_oos, mteb/amazon_massive_intent, benayas/snips | CC-BY-3.0, CC-BY-4.0, CC0-1.0 |
| Routing, 20 intents | 1,000 | legacy-datasets/banking77 | CC-BY-4.0 |
| Routing, 77 intents | 500 | legacy-datasets/banking77 | CC-BY-4.0 |
| Tool routing | 1,000 | gorilla-llm/Berkeley-Function-Calling-Leaderboard | Apache-2.0 |
| Complaint routing | 1,000 | BEE-spoke-data/consumer-finance-complaints | CC0-1.0 |
| Commit type | 1,000 | github.com/angular/angular, github.com/vitejs/vite | MIT |
| Search relevance | 1,000 | tasksource/esci | Apache-2.0 |
Release 2026-09-26.2, benchmark commit 2386153. All 66 runs behind these numbers were made from a working tree with uncommitted harness changes, so the exact code of those runs is not reproducible from a commit alone.
Spot something outdated? Tell us.
Instant Evals runs on Jev, and the obvious alternative is to ask a frontier LLM the same question. We wanted to know where that is worth its price. The tie test decides who is ahead on each task, and a task where a frontier LLM beats Jev is shown as such.
The frontier LLMs answer with one label, not a probability, so AUROC and catch rates at a fixed false-alarm rate cannot be computed for them. Every model is compared on balanced accuracy there, computed the same way: Jev at the decision threshold frozen on the dev split, each LLM at its label.
Among the frontier LLMs, yes: each call was timed from Denmark through the LangWatch AI Gateway, several at a time, network included. Jev’s latency is an approximate server-side time, so it is on another basis and the page names no fastest model across the two.
One prompt is one decision: one call to the model. The cost is the recorded token usage of the test run at each vendor’s public list price, with cached input priced as the provider applied it.
The paired tie test, run on the same items for every model (an empty answer counts as wrong), cannot separate the two scores: the gap is within what resampling the same items produces. A tied task is neither a win nor a loss for either side.
No. The page publishes aggregates only: scores with intervals, latency and cost per model and task. We read the per-item results in LangWatch Experiments, one experiment per task with every model as a target; you can set up the same kind of comparison on your own data in a LangWatch project.
Instant Evals asks one question of your whole production history from the CLI.