Frontier LLMs beat Jev on 9 of 11 tasks

Jev against 7 frontier LLMs on 11 real decision tasks, with 95% intervals and tie tiers. Pick a task, filter, or click a model.

This benchmark was created using LangWatch. Sign up to create your own · Read the method

Task
Weights
Provider
Size

Showing 8 of 8 models. Filters hide rows; tiers and intervals stay as measured on every model.

Prompt injection

Balanced accuracy, % · Higher is better

  1. Jev95.2
    , interval 93.9 to 96.5, ahead alone
  2. , interval 91.1 to 94.3, tier 2
  3. , interval 89.6 to 92.9, tier 2
  4. , interval 89.2 to 92.5, tier 2
  5. , interval 88.6 to 92.1, tier 3
  6. , interval 86.4 to 90.2, tier 4
  7. , interval 86.3 to 90.0, tier 4
  8. , interval 86.1 to 89.8, tier 4

Ahead aloneTied for the leadThe rest95% interval

One task at a time

Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.

Prompt injection: Balanced accuracy, higher is better

  1. Jev95.2%, interval 95.2% [93.9, 96.5], tier 1
  2. Claude Haiku 4.592.7%, interval 92.7% [91.1, 94.3], tier 2
  3. GPT-6.1 Sol91.3%, interval 91.3% [89.6, 92.9], tier 2
  4. Claude Opus 5.590.9%, interval 90.9% [89.2, 92.5], tier 2
  5. Gemini 3.8 Flash90.4%, interval 90.4% [88.6, 92.1], tier 3
  6. Claude Sonnet 5.588.3%, interval 88.3% [86.4, 90.2], tier 4
  7. GPT-5.6 Terra88.2%, interval 88.2% [86.3, 90.0], tier 4
  8. GPT-6 Luna88.0%, interval 88.0% [86.1, 89.8], tier 4
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 86%, not 0.
Plot

Prompt injection: balanced accuracy (higher is better) vs cost

↖ Better
  • Jev
← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑

Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).

Every model, every task

Every visible model on every task, sorted by Prompt injection. Value with half its interval; bands are tie tiers.
Model
Tier 1 on Prompt injection
Jev95.2±1.3 · T195.2% [93.9, 96.5], tier 178.0±3.0 · T278.0% [74.9, 81.0], tier 292.9±1.5 · T292.9% [91.3, 94.4], tier 280.3±2.8 · T280.3% [77.5, 83.1], tier 293.4±1.6 · T393.4% [91.8, 94.9], tier 389.1±2.0 · T389.1% [87.0, 91.0], tier 379.6±3.5 · T479.6% [76.0, 83.0], tier 478.3±2.5 · T478.3% [75.7, 80.7], tier 478.7±2.5 · T178.7% [76.3, 81.2], tier 168.3±2.9 · T368.3% [65.2, 71.1], tier 357.7±3.1 · T257.7% [54.7, 60.8], tier 2
Tier 2
Claude Haiku 4.592.7±1.6 · T292.7% [91.1, 94.3], tier 281.1±3.0 · T181.1% [78.0, 84.0], tier 183.1±2.2 · T583.1% [80.9, 85.3], tier 577.4±2.9 · T377.4% [74.4, 80.3], tier 391.5±1.7 · T491.5% [89.7, 93.2], tier 489.3±2.0 · T389.3% [87.3, 91.2], tier 380.0±3.4 · T380.0% [76.6, 83.4], tier 382.0±2.3 · T282.0% [79.6, 84.2], tier 277.1±2.6 · T277.1% [74.6, 79.7], tier 267.9±2.9 · T367.9% [64.8, 70.7], tier 350.6±3.2 · T350.6% [47.4, 53.7], tier 3
GPT-6.1 Sol91.3±1.7 · T291.3% [89.6, 92.9], tier 281.4±2.9 · T181.4% [78.5, 84.2], tier 191.3±1.7 · T391.3% [89.7, 93.0], tier 384.2±2.7 · T184.2% [81.5, 86.9], tier 195.3±1.3 · T195.3% [93.9, 96.6], tier 192.8±1.6 · T292.8% [91.2, 94.3], tier 284.4±3.2 · T284.4% [81.2, 87.6], tier 279.9±2.4 · T379.9% [77.4, 82.2], tier 379.0±2.4 · T179.0% [76.5, 81.4], tier 174.5±2.7 · T274.5% [71.6, 77.0], tier 263.0±3.1 · T163.0% [60.0, 66.1], tier 1
Claude Opus 5.590.9±1.7 · T290.9% [89.2, 92.5], tier 281.1±3.0 · T181.1% [78.0, 84.0], tier 195.4±1.3 · T195.4% [94.1, 96.6], tier 185.0±2.6 · T185.0% [82.3, 87.5], tier 196.0±1.2 · T196.0% [94.8, 97.1], tier 195.0±1.3 · T195.0% [93.6, 96.3], tier 188.0±2.8 · T188.0% [85.2, 90.8], tier 178.7±2.5 · T478.7% [76.1, 81.1], tier 478.9±2.5 · T178.9% [76.4, 81.4], tier 177.3±2.6 · T177.3% [74.7, 79.9], tier 163.6±3.0 · T163.6% [60.7, 66.7], tier 1
Tier 3
Gemini 3.8 Flash90.4±1.7 · T390.4% [88.6, 92.1], tier 381.4±2.9 · T181.4% [78.4, 84.1], tier 193.1±1.6 · T293.1% [91.5, 94.6], tier 283.3±2.9 · T183.3% [80.4, 86.2], tier 195.0±1.4 · T295.0% [93.6, 96.3], tier 292.7±1.5 · T292.7% [91.1, 94.2], tier 284.8±3.2 · T284.8% [81.6, 88.0], tier 283.6±2.2 · T283.6% [81.3, 85.7], tier 279.6±2.4 · T179.6% [77.1, 82.0], tier 178.0±2.6 · T178.0% [75.3, 80.4], tier 162.5±3.0 · T162.5% [59.6, 65.6], tier 1
Tier 4
Claude Sonnet 5.588.3±1.9 · T488.3% [86.4, 90.2], tier 481.6±3.0 · T181.6% [78.5, 84.4], tier 194.0±1.5 · T294.0% [92.6, 95.5], tier 283.6±2.8 · T183.6% [80.8, 86.4], tier 195.4±1.3 · T195.4% [94.1, 96.6], tier 191.7±1.8 · T291.7% [89.9, 93.4], tier 284.2±3.1 · T284.2% [81.0, 87.2], tier 279.4±2.4 · T379.4% [76.9, 81.7], tier 378.3±2.6 · T178.3% [75.7, 80.8], tier 173.3±2.7 · T273.3% [70.5, 75.9], tier 261.1±3.0 · T161.1% [58.1, 64.0], tier 1
GPT-5.6 Terra88.2±1.8 · T488.2% [86.3, 90.0], tier 479.9±2.9 · T179.9% [77.0, 82.9], tier 186.4±2.0 · T486.4% [84.5, 88.4], tier 480.6±2.9 · T280.6% [77.7, 83.6], tier 292.3±1.7 · T392.3% [90.6, 93.9], tier 392.4±1.7 · T292.4% [90.7, 94.0], tier 283.4±3.3 · T283.4% [80.0, 86.6], tier 284.2±2.2 · T184.2% [82.0, 86.3], tier 178.4±2.6 · T178.4% [75.9, 81.0], tier 166.5±2.9 · T366.5% [63.4, 69.3], tier 357.6±3.1 · T257.6% [54.3, 60.6], tier 2
GPT-6 Luna88.0±1.9 · T488.0% [86.1, 89.8], tier 480.8±3.0 · T180.8% [77.6, 83.7], tier 187.2±1.9 · T487.2% [85.3, 89.1], tier 479.7±3.0 · T279.7% [76.7, 82.7], tier 290.6±1.8 · T490.6% [88.7, 92.4], tier 490.3±1.9 · T390.3% [88.4, 92.1], tier 382.2±3.4 · T382.2% [78.8, 85.6], tier 385.5±2.2 · T185.5% [83.3, 87.6], tier 177.9±2.5 · T277.9% [75.4, 80.4], tier 267.7±2.9 · T367.7% [64.5, 70.4], tier 354.7±3.1 · T254.7% [51.6, 57.8], tier 2

Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.

Jev

Provider
TypeSafe
Weights
Closed (API)
Size
Not disclosed
Licence
proprietary
TaskScore [95% interval]TierLatency p50Cost / 1k
Prompt injection95.2% [93.9, 96.5]1, ahead~59 ms (Hosted API)$0.026
Moderation78.0% [74.9, 81.0]2~65 ms (Hosted API)$0.020
PII92.9% [91.3, 94.4]2~65 ms (Hosted API)$0.023
RAG faithfulness80.3% [77.5, 83.1]2~65 ms (Hosted API)$0.029
Off-topic93.4% [91.8, 94.9]3~60 ms (Hosted API)$0.036
Routing, 20 intents89.1% [87.0, 91.0]3~67 ms (Hosted API)$0.034
Routing, 77 intents79.6% [76.0, 83.0]4~80 ms (Hosted API)$0.093
Tool routing78.3% [75.7, 80.7]4~71 ms (Hosted API)$0.042
Complaint routing78.7% [76.3, 81.2]1, tied~70 ms (Hosted API)$0.041
Commit type68.3% [65.2, 71.1]3~73 ms (Hosted API)$0.025
Search relevance57.7% [54.7, 60.8]2~63 ms (Hosted API)$0.027
  • Training data not disclosed, so contamination can't be ruled out.
  • ~ marks an approximate latency or an estimated cost.
How we measured

Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the metric difference, on the items every model answered.

Tie tiers

Models in one tier can't be separated by the tie test (paired bootstrap of the Balanced accuracy difference on shared items, 95% CI includes 0; paired bootstrap of the Balanced acc. difference on shared items, 95% CI includes 0; paired McNemar on per-item correctness at the dev threshold, p >= 0.05). Tiers are computed in the export over every model; filters on this page never change them.

Contamination

Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner.

Latency and hardware

  • Hosted API: the laptop (Denmark) through the LangWatch AI Gateway, concurrency 6; latency is the time of each call during the scoring run and includes the network and the gateway hop. the laptop (Denmark), concurrency 1; approximate latency: p50 is the client time minus the network round trip measured in the same window (warm keep-alive requests to the API host); p95 is the API's server-reported time.

Latencies are only compared within one hardware tier.

Datasets

  • Prompt injection (1,000 test items): deepset/prompt-injections (Apache-2.0), jackhhao/jailbreak-classification (Apache-2.0), reshabhs/SPML_Chatbot_Prompt_Injection (MIT), djapp18/JailbreaksOverTime (CC-BY-4.0)
  • Moderation (667 test items): mmathys/openai-moderation-api-evaluation (MIT), nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (CC-BY-4.0)
  • PII (1,000 test items): gretelai/gretel-pii-masking-en-v1 (Apache-2.0), beki/privy (MIT), nvidia/Nemotron-PII (CC-BY-4.0)
  • RAG faithfulness (666 test items): pminervini/HaluEval (Apache-2.0)
  • Off-topic (1,000 test items): clinc/clinc_oos (CC-BY-3.0), mteb/amazon_massive_intent (CC-BY-4.0), benayas/snips (CC0-1.0)
  • Routing, 20 intents (1,000 test items): legacy-datasets/banking77 (CC-BY-4.0)
  • Routing, 77 intents (500 test items): legacy-datasets/banking77 (CC-BY-4.0)
  • Tool routing (1,000 test items): gorilla-llm/Berkeley-Function-Calling-Leaderboard (Apache-2.0)
  • Complaint routing (1,000 test items): BEE-spoke-data/consumer-finance-complaints (CC0-1.0)
  • Commit type (1,000 test items): github.com/angular/angular (MIT), github.com/vitejs/vite (MIT)
  • Search relevance (1,000 test items): tasksource/esci (Apache-2.0)
  • Jev's latency on RAG faithfulness comes from its latency run, whose 300 items are all from HaluEval dialogue, the non-commercial source this task is scored without: only the timing is used, no item or score.

Release 2026-10-01.1, benchmark commit 19491fc, 26 input files, integrity sha256-leaves-v1.