Frontier LLMs beat Jev on 9 of 11 tasks

Jev ties on one task and leads on one. The LLM ahead costs a median 34 times more per prompt.

This benchmark was created using LangWatch. Sign up to create your ownLangWatch Instant Evals runs on Jev. Check it outEvery frontier LLM call ran through the LangWatch AI Gateway: one key for all five models, cost and latency logged per call. See the gateway

All tasks: average score

11 tasks
  • Jev (diamond)
  • Frontier LLM (circle)
  • Blue: Jev
  • Orange: the highest average among the frontier LLMs
  • Grey: the rest
  1. Claude Opus 5.584.5%, 95% interval 83.8 to 85.2, averaged over 11 of 11 tasks
  2. Gemini 3.8 Flash84.0%, 95% interval 83.3 to 84.7, averaged over 11 of 11 tasks
  3. Jev81.0%, 95% interval 80.3 to 81.8, averaged over 11 of 11 tasks
  4. GPT-5.6 Terra80.9%, 95% interval 80.1 to 81.7, averaged over 11 of 11 tasks
  5. GPT-6 Luna80.4%, 95% interval 79.6 to 81.2, averaged over 11 of 11 tasks
  6. Claude Haiku 4.579.3%, 95% interval 78.5 to 80.1, averaged over 11 of 11 tasks
Dot = average, line = combined 95% interval. Axis starts at 78%, not 0.Unweighted average of the task scores (balanced accuracy and accuracy); the interval treats tasks as independent. It is not a ranking: pick a task for the tie test.

Who is ahead on each task

Jev against the best frontier LLM on each task; tied means the tie test cannot separate them. Cost is per prompt: one decision, one call.

Where a small model is enough

  • Enough for prompt injection and complaint routing: Jev is ahead of or tied with the best frontier LLM there, which costs 20 to 24 times more per prompt.
  • Not enough for the other nine: a frontier LLM is ahead by 2.5 to 9.7 points, most on commit type.
  • The highest score among them: Claude Opus 5.5 on six; Gemini 3.8 Flash on moderation and commit type, at 24 times Jev’s cost per prompt; GPT-6 Luna on tool routing, at 1.9 times Jev’s cost per prompt.

Prompt injection Balanced accuracy

Jev
95.2%
Best LLM
Claude Haiku 4.5 92.7%

Jev ahead by 2.5 points

Per prompt: Jev 0.0026¢, Claude Haiku 4.5 0.061¢ (24 times)

Moderation Balanced accuracy

Jev
78.0%
Best LLM
Gemini 3.8 Flash 81.4%

Gemini 3.8 Flash ahead by 3.4 points

Per prompt: Jev 0.0020¢, Gemini 3.8 Flash 0.045¢ (22 times)

PII Balanced accuracy

Jev
92.9%
Best LLM
Claude Opus 5.5 95.4%

Claude Opus 5.5 ahead by 2.5 points

Per prompt: Jev 0.0023¢, Claude Opus 5.5 0.32¢ (141 times)

RAG faithfulness Balanced accuracy

Jev
80.3%
Best LLM
Claude Opus 5.5 85.0%

Claude Opus 5.5 ahead by 4.7 points

Per prompt: Jev 0.0029¢, Claude Opus 5.5 0.48¢ (169 times)

Off-topic Balanced accuracy

Jev
93.4%
Best LLM
Claude Opus 5.5 96.0%

Claude Opus 5.5 ahead by 2.6 points

Per prompt: Jev 0.0036¢, Claude Opus 5.5 0.12¢ (34 times)

Routing, 20 intents Accuracy

Jev
89.1%
Best LLM
Claude Opus 5.5 95.0%

Claude Opus 5.5 ahead by 5.9 points

Per prompt: Jev 0.0034¢, Claude Opus 5.5 0.50¢ (147 times)

Routing, 77 intents Accuracy

Jev
79.6%
Best LLM
Claude Opus 5.5 88.0%

Claude Opus 5.5 ahead by 8.4 points

Per prompt: Jev 0.0093¢, Claude Opus 5.5 0.16¢ (17 times)

Tool routing Accuracy

Jev
78.3%
Best LLM
GPT-6 Luna 85.5%

GPT-6 Luna ahead by 7.2 points

Per prompt: Jev 0.0042¢, GPT-6 Luna 0.0080¢ (1.9 times)

Complaint routing Accuracy

Jev
78.7%
Best LLM
Gemini 3.8 Flash 79.6%

Tied, 0.9 points apart

Per prompt: Jev 0.0041¢, Gemini 3.8 Flash 0.083¢ (20 times)

Commit type Accuracy

Jev
68.3%
Best LLM
Gemini 3.8 Flash 78.0%

Gemini 3.8 Flash ahead by 9.7 points

Per prompt: Jev 0.0025¢, Gemini 3.8 Flash 0.059¢ (24 times)

Search relevance Accuracy

Jev
57.7%
Best LLM
Claude Opus 5.5 63.6%

Claude Opus 5.5 ahead by 5.9 points

Per prompt: Jev 0.0027¢, Claude Opus 5.5 0.48¢ (174 times)

On prompt injection, moderation and PII the score is balanced accuracy for every model: the LLMs answer with one label, so AUROC and catch rates cannot be computed for them. Jev is scored at the decision threshold frozen on the dev split.

Every set here is public, and no frontier vendor discloses its training data (nor is Jev’s disclosed), so any model on this page may have seen these datasets. No row is grayed; read no score as proof of generalisation.

Pick a task

Every model on each task, with its latency and its cost per prompt in the same order. Overlapping lines mean the data cannot separate two models.

Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.

Balanced accuracy for every model: the LLMs answer with one label, so catch rate at 5% false alarms cannot be computed for them. Jev at its dev‑frozen threshold.

  • Jev (diamond)
  • Frontier LLM (circle)
  • Orange: ahead on the tie test
  • Blue: best of the other side

Jev ahead by 2.5 points

Balanced accuracy, 95% interval

  1. Jev95.2%, 95% interval 93.9 to 96.5, ahead, p50 ~59 ms, 0.0026¢ per prompt
  2. Claude Haiku 4.592.7%, 95% interval 91.1 to 94.3, p50 1,337 ms, 0.061¢ per prompt
  3. Claude Opus 5.590.9%, 95% interval 89.2 to 92.5, p50 2,194 ms, 0.38¢ per prompt
  4. Gemini 3.8 Flash90.4%, 95% interval 88.6 to 92.1, p50 2,216 ms, 0.043¢ per prompt
  5. GPT-5.6 Terra88.2%, 95% interval 86.3 to 90.0, p50 1,673 ms, 0.10¢ per prompt
  6. GPT-6 Luna88.0%, 95% interval 86.1 to 89.8, p50 1,108 ms, 0.0050¢ per prompt
Balanced accuracy, %, 1,000 test items. Dot = score, line = 95% bootstrap interval. Axis starts at 86%, not 0. Claude Opus 5.5 answered 97.5% of the items and Gemini 3.8 Flash answered 99.0% of the items; the rest count as wrong. Every task is in the matrix below.

Latency, p50 per prompt

  1. Jev~59 ms§p95 ~97 ms
  2. Claude Haiku 4.51,337 msp95 2,087 ms
  3. Claude Opus 5.52,194 msp95 4,321 ms
  4. Gemini 3.8 Flash2,216 msp95 6,100 ms
  5. GPT-5.6 Terra1,673 msp95 2,460 ms
  6. GPT-6 Luna1,108 msp95 1,740 ms

Cost per prompt

  1. Jev0.0026¢
  2. Claude Haiku 4.50.061¢
  3. Claude Opus 5.50.38¢
  4. Gemini 3.8 Flash0.043¢
  5. GPT-5.6 Terra0.10¢
  6. GPT-6 Luna0.0050¢

p50 per prompt, p95 under it. Frontier LLMs: each call timed from Denmark through the LangWatch AI Gateway, 6 at a time, network included; orange is the fastest of them. Jev (~, hatched): server‑side, approximate, so on another basis and never orange. Cost per prompt: one prompt is one decision, one call to the model, at the recorded token usage and each vendor’s list price, cached input included; orange is the cheapest.

Caveats for this task

Caveats

  • Each model's question wording (and Jev's decision threshold) was picked on a separate 60-item dev split, before the test run; the frontier LLMs got the same wordings as Jev.
  • Inputs over the set's length limit (2,500 to 12,000 characters, per set) were removed before sampling, for every model alike.
  • Frontier LLM latency is the time of each call during the scoring run (6 calls at a time), measured from Denmark through the LangWatch AI Gateway: it includes the network and the gateway hop.
  • Cost is the recorded token usage of the test run at each vendor's public list price, cached-input discounts included where the provider applied them (Jev: input tokens only, output free).
  • Flip rate: answers that changed when the same items were sent again. Measured for Jev on every task but RAG faithfulness (its repeat run holds only the left-out non-commercial items) and for GPT-6 Luna on prompt injection and moderation; the other frontier LLMs were not repeated.
  • Labels are the datasets' own; no human re-check yet.
  • Jev latency is approximate. p50: measured from Denmark minus the network round trip measured in the same window; p95: Jev's server-reported time, because the network's tail can't be subtracted reliably. It is not on the same basis as the frontier LLMs' per-call times.

Every task at a glance

Score per task and model, in percent, shaded by rank within each task: darker is higher.

  • Prompt injection Balanced accuracyJev 95.2%, highest score on this taskClaude Opus 5.5 90.9%undefinedGemini 3.8 Flash 90.4%undefinedGPT-5.6 Terra 88.2%undefinedGPT-6 Luna 88.0%undefinedClaude Haiku 4.5 92.7%undefined
  • Moderation Balanced accuracyJev 78.0%undefinedClaude Opus 5.5 81.1%undefinedGemini 3.8 Flash 81.4%, highest score on this taskGPT-5.6 Terra 79.9%undefinedGPT-6 Luna 80.8%undefinedClaude Haiku 4.5 81.1%undefined
  • PII Balanced accuracyJev 92.9%undefinedClaude Opus 5.5 95.4%, highest score on this taskGemini 3.8 Flash 93.1%undefinedGPT-5.6 Terra 86.4%undefinedGPT-6 Luna 87.2%undefinedClaude Haiku 4.5 83.1%undefined
  • RAG faithfulness Balanced accuracyJev 80.3%undefinedClaude Opus 5.5 85.0%, highest score on this taskGemini 3.8 Flash 83.3%undefinedGPT-5.6 Terra 80.6%undefinedGPT-6 Luna 79.7%undefinedClaude Haiku 4.5 77.4%undefined
  • Off-topic Balanced accuracyJev 93.4%undefinedClaude Opus 5.5 96.0%, highest score on this taskGemini 3.8 Flash 95.0%undefinedGPT-5.6 Terra 92.3%undefinedGPT-6 Luna 90.6%undefinedClaude Haiku 4.5 91.5%undefined
  • Routing, 20 intents AccuracyJev 89.1%undefinedClaude Opus 5.5 95.0%, highest score on this taskGemini 3.8 Flash 92.7%undefinedGPT-5.6 Terra 92.4%undefinedGPT-6 Luna 90.3%undefinedClaude Haiku 4.5 89.3%undefined
  • Routing, 77 intents AccuracyJev 79.6%undefinedClaude Opus 5.5 88.0%, highest score on this taskGemini 3.8 Flash 84.8%undefinedGPT-5.6 Terra 83.4%undefinedGPT-6 Luna 82.2%undefinedClaude Haiku 4.5 80.0%undefined
  • Tool routing AccuracyJev 78.3%undefinedClaude Opus 5.5 78.7%undefinedGemini 3.8 Flash 83.6%undefinedGPT-5.6 Terra 84.2%undefinedGPT-6 Luna 85.5%, highest score on this taskClaude Haiku 4.5 82.0%undefined
  • Complaint routing Accuracy, tiedJev 78.7%undefinedClaude Opus 5.5 78.9%undefinedGemini 3.8 Flash 79.6%, highest score on this taskGPT-5.6 Terra 78.4%undefinedGPT-6 Luna 77.9%undefinedClaude Haiku 4.5 77.1%undefined
  • Commit type AccuracyJev 68.3%undefinedClaude Opus 5.5 77.3%undefinedGemini 3.8 Flash 78.0%, highest score on this taskGPT-5.6 Terra 66.5%undefinedGPT-6 Luna 67.7%undefinedClaude Haiku 4.5 67.9%undefined
  • Search relevance AccuracyJev 57.7%undefinedClaude Opus 5.5 63.6%, highest score on this taskGemini 3.8 Flash 62.5%undefinedGPT-5.6 Terra 57.6%undefinedGPT-6 Luna 54.7%undefinedClaude Haiku 4.5 50.6%undefined
1st, 2nd, 3rd, 4th, 5th and below within the task: darker is higher rank; equal scores share a step. Whether the leader is ahead or tied is in the row header and the cards.grey: not ranked (trained on the data, withheld or not scored)Left to right: Jev, Claude Opus 5.5, Gemini 3.8 Flash, GPT-5.6 Terra, GPT-6 Luna, Claude Haiku 4.5.

How the numbers were made

The Jev benchmark’s frozen test sets, one prompt per decision, 95% bootstrap intervals and a paired tie test.

Full method, datasets and footnotes
  1. 1

    Same tasks, same sets

    The 11 use cases of the Jev benchmark, with the same frozen dev and test splits (up to 1,000 test items each). The open models on the same sets.

  2. 2

    One prompt per decision

    Each LLM got the task and its options in the system message and the item last, and answered with one label through a JSON schema. The wording was frozen on the 60‑item dev split before the test run. Reasoning was off, or at the lowest setting each model accepts; some models still reason briefly at their lowest setting.

  3. 3

    Label‑only scores

    The frontier LLMs answer with one label, not a probability, so threshold-free metrics (AUROC, catch rate at a fixed false-alarm rate) cannot be computed for them. On moderation, PII and prompt injection every model is therefore compared on balanced accuracy, computed the same way for all: Jev at the decision threshold frozen on the dev split, each LLM at its label.

  4. 4

    Intervals and the tie test

    Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the difference in balanced accuracy, on the same items for every model; an empty answer counts as wrong.

  5. 5

    Latency

    Frontier LLMs: each call timed from Denmark through the LangWatch AI Gateway, 6 at a time, network and gateway included (p50 1,030 ms to 2,539 ms by task and model). Jev: approximate server‑side time (~, p50 ~59 ms to ~80 ms). The two bases differ, so the page names no fastest model across them.

  6. 6

    Cost

    Recorded token usage of the test run at each vendor’s public list price, cached input included where the provider applied it; Jev at its list price (input tokens only, output free). Shown per prompt; the export also carries it per 1,000 decisions. Jev is the cheapest on nine of 11 tasks; GPT-6 Luna costs less on off-topic and routing, 77 intents.

  7. 7

    Through the gateway

    Every frontier LLM call went through the LangWatch AI Gateway, which logged the cost and latency of each call, with prompt caching on the shared part of long prompts.

  8. 8

    Coverage

    A model is ranked when it answered at least 95% of a task’s items; an empty answer counts as wrong.

What is left out

  • No flip rate for the frontier LLMs: their test runs were not repeated, apart from GPT-6 Luna on two tasks (in the caveats).

Datasets and provenance

Tasks from public datasets, each with its licence. Two tasks are scored without a source licensed for non‑commercial use.

TaskItemsSourcesLicence
Prompt injection1,000deepset/prompt-injections, jackhhao/jailbreak-classification, reshabhs/SPML_Chatbot_Prompt_Injection, djapp18/JailbreaksOverTimeApache-2.0, MIT, CC-BY-4.0
Moderationwithout toxicchat667mmathys/openai-moderation-api-evaluation, nvidia/Aegis-AI-Content-Safety-Dataset-2.0MIT, CC-BY-4.0
PII1,000gretelai/gretel-pii-masking-en-v1, beki/privy, nvidia/Nemotron-PIIApache-2.0, MIT, CC-BY-4.0
RAG faithfulnesswithout halueval-dialogue666pminervini/HaluEvalApache-2.0
Off-topic1,000clinc/clinc_oos, mteb/amazon_massive_intent, benayas/snipsCC-BY-3.0, CC-BY-4.0, CC0-1.0
Routing, 20 intents1,000legacy-datasets/banking77CC-BY-4.0
Routing, 77 intents500legacy-datasets/banking77CC-BY-4.0
Tool routing1,000gorilla-llm/Berkeley-Function-Calling-LeaderboardApache-2.0
Complaint routing1,000BEE-spoke-data/consumer-finance-complaintsCC0-1.0
Commit type1,000github.com/angular/angular, github.com/vitejs/viteMIT
Search relevance1,000tasksource/esciApache-2.0
  • JailbreaksOverTime: Piet et al. 2025, CC-BY-4.0.
  • Nemotron-PII: NVIDIA, CC-BY-4.0 (attribution requested when results are published).
  • Aegis AI Content Safety 2.0: NVIDIA, CC-BY-4.0.
  • Banking77: Casanueva et al. 2020, CC-BY-4.0.
  • CLINC150: Larson et al. 2019, CC-BY-3.0. MASSIVE: Amazon, CC-BY-4.0. SNIPS: Sonos, CC0-1.0.
  • Shopping Queries Dataset (ESCI): Reddy et al. 2022, Apache-2.0.
  • Berkeley Function Calling Leaderboard v3 live: Apache-2.0.
  • CFPB consumer complaint narratives: US government public domain (BEE-spoke copy, CC0-1.0).
  • HaluEval qa and summarization: Li et al. 2023 (HotpotQA CC-BY-SA-4.0, CNN/DailyMail).
  • Banking77: Casanueva et al. 2020, CC-BY-4.0.

Footnotes

  1. §Jev latency: Jev is a hosted API measured from Denmark. For p50 we measured the network round trip to its servers in the same window (71 requests, p50 208 ms) and subtracted it from our measured p50. The network tail could not be subtracted reliably, so p95 is Jev’s own server‑reported p95 (95 to 114 ms by task). Both are approximate, and neither is on the frontier LLMs’ basis, which includes the network and the gateway.

Provenance

Release 2026-09-26.2, benchmark commit 2386153. All 66 runs behind these numbers were made from a working tree with uncommitted harness changes, so the exact code of those runs is not reproducible from a commit alone.

Spot something outdated? Tell us.

Questions

Why does LangWatch benchmark the model its own product runs on against frontier LLMs?

Instant Evals runs on Jev, and the obvious alternative is to ask a frontier LLM the same question. We wanted to know where that is worth its price. The tie test decides who is ahead on each task, and a task where a frontier LLM beats Jev is shown as such.

Why balanced accuracy on moderation, PII and prompt injection?

The frontier LLMs answer with one label, not a probability, so AUROC and catch rates at a fixed false-alarm rate cannot be computed for them. Every model is compared on balanced accuracy there, computed the same way: Jev at the decision threshold frozen on the dev split, each LLM at its label.

Is the latency comparable?

Among the frontier LLMs, yes: each call was timed from Denmark through the LangWatch AI Gateway, several at a time, network included. Jev’s latency is an approximate server-side time, so it is on another basis and the page names no fastest model across the two.

What does the cost per prompt include?

One prompt is one decision: one call to the model. The cost is the recorded token usage of the test run at each vendor’s public list price, with cached input priced as the provider applied it.

What does tied mean?

The paired tie test, run on the same items for every model (an empty answer counts as wrong), cannot separate the two scores: the gap is within what resampling the same items produces. A tied task is neither a win nor a loss for either side.

Are the per-item results published?

No. The page publishes aggregates only: scores with intervals, latency and cost per model and task. We read the per-item results in LangWatch Experiments, one experiment per task with every model as a target; you can set up the same kind of comparison on your own data in a LangWatch project.

Run the hosted judge on your own data.

Instant Evals asks one question of your whole production history from the CLI.

$npx langwatch instant-eval run