Open models caught up with Jev

Out of distribution, open models up to 28B tie Jev on 8 of 11 tasks, lead on 2 and trail on off-topic.

This benchmark was created using LangWatch. Sign up to create your ownLangWatch Instant Evals runs on Jev. Check it out

Prompt injection: catch rate at 5% false alarms

1,000 test items
  • Jev, hosted API (diamond)
  • Open model (circle)
  • Green: tied with the leader

Tied, 0.4 points apart

  1. Jev94.6%, 95% interval 92.2 to 96.7, tied with the leader, p95 841 ms (API), flips 2/300
  2. Eikos-27B94.2%, 95% interval 90.5 to 97.1, tied with the leader, p95 170 ms (H100), flips 0/300
  3. Shisa DE-194.0%, 95% interval 90.3 to 96.2, tied with the leader, p95 60 ms (H100), flips 0/300
  4. AutoJev-27B92.0%, 95% interval 86.6 to 95.0, tied with the leader, p95 198 ms (H100), flips 0/300
Catch rate at 5% false alarms, %, 1,000 test items. Dot = score, line = 95% bootstrap interval. Axis starts at 86%, not 0.4 of 16 models shown: Jev, the best open model not trained on the task’s data and every model tied with the leader. Every model is in the matrix below.

Who is ahead on each task

Jev against the best open model not trained on the task’s data; tied means the tie test cannot separate them.

Prompt injection Catch rate at 5% false alarms

Jev
94.6%
Best open
Eikos-27B 94.2%

Tied, 0.4 points apart

Moderation AUROC

Jev
90.3%
Best open
Eikos-27B 90.3%

Tied, <0.1 points apart

PII Catch rate at 5% false alarms

Jev
90.8%
Best open
Eikos-27B 95.2%

Eikos-27B ahead by 4.4 points

RAG faithfulness Balanced accuracy

Jev
80.3%
Best open
Eikos-27B 82.1%

Tied, 1.8 points apart

Off-topic Balanced accuracy

Jev
93.4%
Best open
Kev-9B 89.6%

Jev ahead by 3.8 points

Routing, 20 intents Accuracy

Jev
89.1%
Best open
AutoJev-27B 89.7%

Tied, 0.6 points apart

Routing, 77 intents Accuracy

Jev
79.6%
Best open
AutoJev-27B 80.6%

Tied, 1.0 points apart

Tool routing Accuracy

Jev
78.3%
Best open
Shisa DE-1 81.3%

Shisa DE-1 ahead by 3.0 points

Complaint routing Accuracy

Jev
78.7%
Best open
Shisa DE-1 78.0%

Tied, 0.7 points apart

Commit type Accuracy

Jev
68.3%
Best open
Eikos-27B 66.8%

Tied, 1.5 points apart

Search relevance Accuracy

Jev
57.7%
Best open
Eikos-27B 56.6%

Tied, 1.1 points apart

Eikos-27B has the highest open score on 6 of 11 tasks.

* Trained on that task’s dataset: shown for reference, never ranked. Models are marked on what their authors disclosed; no declaration is not the same as not trained on it. A model marked * scores higher than Jev on moderation; routing, 20 intents; and routing, 77 intents.

Pick a task

The top 5 models per task, plus Jev. Overlapping lines mean the data cannot separate two models.

Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.

  • Jev, hosted API (diamond)
  • Open model (circle)
  • Green: tied with the leader

Tied, 0.4 points apart

  1. Jev94.6%, 95% interval 92.2 to 96.7, tied with the leader, p95 841 ms (API), flips 2/300
  2. Eikos-27B94.2%, 95% interval 90.5 to 97.1, tied with the leader, p95 170 ms (H100), flips 0/300
  3. Shisa DE-194.0%, 95% interval 90.3 to 96.2, tied with the leader, p95 60 ms (H100), flips 0/300
  4. AutoJev-27B92.0%, 95% interval 86.6 to 95.0, tied with the leader, p95 198 ms (H100), flips 0/300
  5. Kev-4B91.6%, 95% interval 88.2 to 94.2, p95 260 ms (L4), flips 0/300
Catch rate at 5% false alarms, %, 1,000 test items. Dot = score, line = 95% bootstrap interval. Axis starts at 86%, not 0.5 of 16 models shown: the top 5 by score among models not trained on the task’s data. Every model is in the matrix below.
ModelCoverageCut inputsFlipsp95 ms (GPU)
Jev100%0%2/300841 API
Eikos-27B100%0%0/300170 H100
Shisa DE-1100%n/a0/30060 H100
AutoJev-27B100%n/a0/300198 H100
Kev-4B100%n/a0/300260 L4
Caveats for this task

Flips: answers that changed between two identical runs on the first 300 items.

Without spml (750 items), from the data‑quality audit: SPML is separable by length alone (length-only AUROC 0.998): its benign items are short questions and its injections long persona prompts. Without it, the best open model changes.

Rows not ranked

  • SimpleJev (Qwen3.5-0.8B): Degenerate score. this system's probabilities barely vary (the central 90% spans only 0.0010), so its AUROC 69.0% is a tie-break on numerical noise rather than a measurement
  • SemIf (Qwen3-0.6B): Degenerate score. this system's probabilities barely vary (the central 90% spans only 0.0953), so its AUROC 46.8% is a tie-break on numerical noise rather than a measurement
  • openJev Verdict 1.4: Degenerate score. this system's probabilities barely vary (the central 90% spans only 0.1207), so its AUROC 36.0% is a tie-break on numerical noise rather than a measurement
  • Laya: Threshold unreachable. no threshold reaches the false-alarm budget: a block of benign items scores at the maximum, although AUROC 81.6% says the model ranks the classes apart

Caveats

  • Each model's question wording and decision threshold were picked on a separate 60-item dev split, before the 1,000-item test run.
  • Inputs over the set's length limit (2,500 to 12,000 characters, per set) were removed before sampling, for every model alike.
  • Laya and Verdict read at most 512 tokens (Laya-typed 1,024). Longer inputs are cut by the model's own code. Cut items stay in the score; the share of cut inputs per model is shown next to coverage where the server reports token usage (Laya and Laya-typed), and cannot be measured for Verdict, which reports none.
  • Open models ran one at a time at concurrency 1, each run on one of three GPUs: a laptop NVIDIA RTX 3050 Ti (4 GB), a cloud NVIDIA L4 or a cloud NVIDIA H100 PCIe. Each latency is labelled with the GPU its run used, and latencies are comparable only on the same GPU.
  • Jev latency is measured at our client on the same laptop and includes the network round trip to the hosted API.
  • Flip rate: answers that changed when the same items were sent again.
  • Labels are the datasets' own; no human re-check yet.
  • Reference baselines are trivial classifiers fitted with 5-fold cross-validation on the test items themselves (length only, a bag-of-words TF-IDF logistic regression, a few regexes). They are not contenders: a baseline near a model means the set can be solved without reading it the way a model does.
  • A data-quality audit of the frozen sets (2026-09-22) found sources that a trivial baseline solves; the headline stays on the full frozen set that was run, with the number without that source next to it.

Every task at a glance

Primary score per task and model, in percent. Deeper orange is better; gray is not ranked.

  • Prompt injection Catch rate at 5% false alarms, tiedJev 94.6%, tied with the leaderEikos-27B 94.2%, tied with the leaderShisa DE-1 94.0%, tied with the leaderAutoJev-27B 92.0%, tied with the leaderSemIf (Qwen3.5-4B) 90.6%Eikos-4B 90.8%Kev-9B 87.6%Kev-4B 91.6%GLiNER2.5-Decide 63.2%Kev-0.8B 56.2%Laya-typed 49.2%Laya 0.0%, threshold unreachableKev-0.6B 25.4%SimpleJev (Qwen3.5-0.8B) 15.2%, degenerate scoreSemIf (Qwen3-0.6B) 7.8%, degenerate scoreopenJev Verdict 1.4 4.8%, trained on this dataset, degenerate score
  • Moderation AUROC, tiedJev 90.3%, tied with the leaderEikos-27B 90.3%, tied with the leaderShisa DE-1 88.3%AutoJev-27B 90.6%, trained on this datasetSemIf (Qwen3.5-4B) 88.7%, tied with the leaderEikos-4B 88.8%, tied with the leaderKev-9B 88.5%Kev-4B 82.0%GLiNER2.5-Decide 71.5%Kev-0.8B 76.2%Laya-typed 70.0%Laya 69.8%Kev-0.6B 65.3%SimpleJev (Qwen3.5-0.8B) 49.9%, degenerate scoreSemIf (Qwen3-0.6B) 63.0%, degenerate scoreopenJev Verdict 1.4 51.8%, degenerate score
  • PII Catch rate at 5% false alarmsJev 90.8%Eikos-27B 95.2%, aheadShisa DE-1 84.2%AutoJev-27B 92.8%, tied with the leaderSemIf (Qwen3.5-4B) 66.1%Eikos-4B 78.6%Kev-9B 62.5%Kev-4B 9.0%GLiNER2.5-Decide 16.8%Kev-0.8B 22.8%Laya-typed 21.0%Laya 15.8%Kev-0.6B 9.6%SimpleJev (Qwen3.5-0.8B) 3.4%, degenerate scoreSemIf (Qwen3-0.6B) 9.6%, degenerate scoreopenJev Verdict 1.4 1.0%, degenerate score
  • RAG faithfulness Balanced accuracy, tiedJev 80.3%, tied with the leaderEikos-27B 82.1%, tied with the leaderShisa DE-1 79.2%AutoJev-27B 79.1%SemIf (Qwen3.5-4B) 72.0%Eikos-4B 71.0%Kev-9B 73.1%Kev-4B 72.5%GLiNER2.5-Decide 57.8%, degenerate scoreKev-0.8B 55.5%Laya-typed 49.4%Laya 49.8%Kev-0.6B 54.1%SimpleJev (Qwen3.5-0.8B) 50.8%, degenerate scoreSemIf (Qwen3-0.6B) 52.1%openJev Verdict 1.4 50.3%, degenerate score
  • Off-topic Balanced accuracyJev 93.4%, aheadEikos-27B 90.6%, trained on this datasetShisa DE-1 88.0%AutoJev-27B 92.7%, trained on this datasetSemIf (Qwen3.5-4B) 81.1%Eikos-4B 79.6%, trained on this datasetKev-9B 89.6%Kev-4B 84.4%GLiNER2.5-Decide 49.3%, degenerate scoreKev-0.8B 76.8%Laya-typed 52.6%Laya 54.3%Kev-0.6B 55.8%SimpleJev (Qwen3.5-0.8B) 58.8%, degenerate scoreSemIf (Qwen3-0.6B) 53.5%, degenerate scoreopenJev Verdict 1.4 constant output
  • Routing, 20 intents Accuracy, tiedJev 89.1%, tied with the leaderEikos-27B 89.6%, trained on this datasetShisa DE-1 88.1%AutoJev-27B 89.7%, tied with the leaderSemIf (Qwen3.5-4B) not scoredEikos-4B 87.7%, trained on this datasetKev-9B 93.0%, trained on this datasetKev-4B 93.0%, trained on this datasetGLiNER2.5-Decide 86.0%Kev-0.8B 91.3%, trained on this datasetLaya-typed 77.0%Laya 76.0%Kev-0.6B 89.5%, trained on this datasetSimpleJev (Qwen3.5-0.8B) 60.2%SemIf (Qwen3-0.6B) not scoredopenJev Verdict 1.4 78.4%, trained on this dataset
  • Routing, 77 intents Accuracy, tiedJev 79.6%, tied with the leaderEikos-27B 77.8%, trained on this datasetShisa DE-1 not scoredAutoJev-27B 80.6%, tied with the leaderSemIf (Qwen3.5-4B) not scoredEikos-4B 74.0%, trained on this datasetKev-9B 85.0%, trained on this datasetKev-4B 85.0%, trained on this datasetGLiNER2.5-Decide 70.6%Kev-0.8B 83.0%, trained on this datasetLaya-typed 40.0%Laya 42.4%Kev-0.6B 78.2%, trained on this datasetSimpleJev (Qwen3.5-0.8B) not scoredSemIf (Qwen3-0.6B) not scoredopenJev Verdict 1.4 not scored
  • Tool routing AccuracyJev 78.3%Eikos-27B 79.3%Shisa DE-1 81.3%, aheadAutoJev-27B 80.6%, tied with the leaderSemIf (Qwen3.5-4B) 79.9%, tied with the leaderEikos-4B 75.5%Kev-9B 72.4%Kev-4B 71.4%GLiNER2.5-Decide 32.4%Kev-0.8B 57.7%Laya-typed 21.5%Laya 22.9%Kev-0.6B 44.0%SimpleJev (Qwen3.5-0.8B) 29.1%SemIf (Qwen3-0.6B) 42.6%openJev Verdict 1.4 16.4%
  • Complaint routing Accuracy, tiedJev 78.7%, tied with the leaderEikos-27B 75.6%Shisa DE-1 78.0%, tied with the leaderAutoJev-27B 75.9%SemIf (Qwen3.5-4B) 72.9%Eikos-4B no runKev-9B 75.0%Kev-4B 76.2%GLiNER2.5-Decide 52.6%Kev-0.8B 38.5%Laya-typed 47.1%Laya 44.2%Kev-0.6B 59.3%SimpleJev (Qwen3.5-0.8B) 45.2%SemIf (Qwen3-0.6B) 37.8%openJev Verdict 1.4 46.3%
  • Commit type Accuracy, tiedJev 68.3%, tied with the leaderEikos-27B 66.8%, tied with the leaderShisa DE-1 65.1%AutoJev-27B 66.7%, tied with the leaderSemIf (Qwen3.5-4B) 61.2%Eikos-4B 56.7%Kev-9B 55.5%Kev-4B 53.7%GLiNER2.5-Decide 50.4%Kev-0.8B 48.6%Laya-typed 49.7%Laya 44.8%Kev-0.6B 41.6%SimpleJev (Qwen3.5-0.8B) 32.7%SemIf (Qwen3-0.6B) 23.9%openJev Verdict 1.4 26.5%
  • Search relevance Accuracy, tiedJev 57.7%, tied with the leaderEikos-27B 56.6%, tied with the leaderShisa DE-1 50.3%AutoJev-27B 52.8%SemIf (Qwen3.5-4B) 41.7%Eikos-4B 43.3%Kev-9B 43.4%Kev-4B 44.8%GLiNER2.5-Decide 25.0%Kev-0.8B 29.3%Laya-typed 29.8%Laya 31.5%Kev-0.6B 26.2%SimpleJev (Qwen3.5-0.8B) 25.0%SemIf (Qwen3-0.6B) 25.0%openJev Verdict 1.4 24.8%
  • General suites, not counted in the title
  • Typed decisions Accuracy, tiedJev 73.9%, tied with the leaderEikos-27B 73.4%, tied with the leaderShisa DE-1 75.1%, tied with the leaderAutoJev-27B 73.8%, tied with the leaderSemIf (Qwen3.5-4B) 63.0%Eikos-4B 65.3%Kev-9B no runKev-4B 65.8%GLiNER2.5-Decide 48.8%Kev-0.8B 44.9%Laya-typed 77.4%, trained on this datasetLaya 36.1%Kev-0.6B 44.8%SimpleJev (Qwen3.5-0.8B) 39.0%SemIf (Qwen3-0.6B) 35.4%openJev Verdict 1.4 36.4%
  • Web-agent actions AccuracyJev 70.8%, aheadEikos-27B 68.4%Shisa DE-1 65.2%AutoJev-27B 68.2%SemIf (Qwen3.5-4B) 59.5%Eikos-4B 62.9%Kev-9B no runKev-4B 58.0%GLiNER2.5-Decide 37.0%Kev-0.8B 50.9%Laya-typed 18.3%Laya 17.8%Kev-0.6B 24.8%SimpleJev (Qwen3.5-0.8B) 56.3%SemIf (Qwen3-0.6B) 17.9%openJev Verdict 1.4 18.5%
  • Community sets Accuracy, tiedJev 62.3%, tied with the leaderEikos-27B 59.8%Shisa DE-1 38.2%AutoJev-27B 63.0%, tied with the leaderSemIf (Qwen3.5-4B) 37.3%Eikos-4B 55.8%Kev-9B no runKev-4B 55.8%GLiNER2.5-Decide 41.4%Kev-0.8B 46.4%Laya-typed 47.8%Laya 45.9%Kev-0.6B 45.3%SimpleJev (Qwen3.5-0.8B) 25.0%SemIf (Qwen3-0.6B) 26.2%openJev Verdict 1.4 22.6%
≤30100gray: not ranked (trained on the data, degenerate, not scored or withheld).Dots left to right: Jev, Eikos-27B, Shisa DE-1, AutoJev-27B, SemIf (Qwen3.5-4B), Eikos-4B, Kev-9B, Kev-4B, GLiNER2.5-Decide, Kev-0.8B, Laya-typed, Laya, Kev-0.6B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B), openJev Verdict 1.4.black ring: ahead: the tie test separates the other side’s best from it.green ring: tied with the leader, the tie test cannot separate them.* hatched: trained on this dataset, never ranked. ‡ degenerate score or unreachable threshold, never ranked; the task’s caveats say why. – not scored. ? or = not a score (incomplete run, constant output or under investigation).

How the numbers were made

Frozen test sets, 95% bootstrap intervals, a paired tie test and a contamination audit.

Full method, datasets and footnotes
  1. 1

    Frozen splits

    Each use case has a 60‑item dev split and a test split of up to 1,000 items, frozen with checksums before any model ran.

  2. 2

    Wording chosen on dev

    Every model got each task in two or three wordings; the wording and the decision threshold were picked on the dev split, once, then frozen.

  3. 3

    Intervals and the tie test

    Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The general suites use 10,000 resamples, clustered so the variants of one item move together. The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the metric difference, on the items every model answered.

  4. 4

    Latency

    Open models ran one at a time on a laptop GPU, an L4 or an H100, so latency compares only on the same GPU; every p95 names its GPU. Jev’s latency is measured from a client and includes the network. Cost is not shown on this release.

  5. 5

    Long inputs

    Inputs over each set’s length limit (2,500 to 12,000 characters) were removed for every model alike before sampling. Inputs a model cuts at its own token window stay in the score.

  6. 6

    Contamination

    Each model and task pair was audited against the model’s published training, selection and calibration data. Grayed rows trained on the test items or on the same dataset; footnoted rows trained on a related task. Models are marked on what their authors disclosed; no declaration is not the same as not trained on it.

The contamination rule

Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner. Grayed rows are shown for reference, never ranked.

Jev’s training data is not disclosed, and every open model starts from a pretrained base whose corpus is not itemised. An unmarked row means no known contamination, not proof that there is none.

Shisa DE-1 declares no training data and is grayed on no task; openJev Verdict 1.4 disclosed its data and is grayed on 4 of 14.

Where a grayed model has the best open score, counted:

  • Moderation: AutoJev-27B* (trained on the dataset behind 33% of the test items) scores 90.6%, 0.4 points ahead of Jev.
  • Off-topic: AutoJev-27B* (trained on the dataset behind 33% of the test items) scores 92.7%, 0.7 points behind Jev.
  • Routing, 20 intents: Kev-4B* (trained on the dataset behind 100% of the test items) scores 93.0%, 3.9 points ahead of Jev.
  • Routing, 77 intents: Kev-4B* (trained on the dataset behind 100% of the test items) scores 85.0%, 5.4 points ahead of Jev.

What is left out

  • JevBench public is left out: it is the set open models are tuned against.

What the data‑quality audit found

No trivial baseline beats the best open model on any of the 11 tasks that have one. The notes below say what a better set would change.

  • Prompt injection. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 31.0%; Bag of words (TF-IDF) 72.6%. SPML is separable by length alone (length-only AUROC 0.998): its benign items are short questions and its injections long persona prompts. Without it, the best open model changes. Without spml (750 items): Jev 92.5%, Eikos-27B 92.3%, Shisa DE-1 92.3%.
  • Moderation. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 57.6%; Bag of words (TF-IDF) 76.3%.
  • PII. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 3.2%; Bag of words (TF-IDF) 77.2%; Six regexes 68.7%. The gretel and nemotron negatives carry a redaction stand-in ('on file', 'the customer') that no positive has, so a bag-of-words classifier separates them perfectly (AUROC 1.000). The number without them is on privy alone, the one source without that artefact, where the gap between Jev and the best open model is a fraction of the headline gap. Without gretel and nemotron (333 items): Kev-9B 91.6%, AutoJev-27B 90.4%, Kev-4B 88.6%.
  • RAG faithfulness. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 67.8%; Bag of words (TF-IDF) 54.6%; String-overlap heuristic 68.0%. HaluEval-qa is a length and style detector: 92.8% of its hallucinated answers end in a full stop against 1.2% of the correct ones, and a five-feature string-overlap heuristic (AUROC 0.977) beats Jev (0.934) on it. The number without it is on HaluEval summarization alone. Without halueval-dialogue and halueval-qa (333 items): Eikos-27B 79.0%, Shisa DE-1 78.6%, Kev-9B 76.6%.
  • Off-topic. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 50.8%; Bag of words (TF-IDF) 83.5%. SNIPS is separable by keywords (TF-IDF AUROC 0.999): it holds only two out-of-scope intent types. Without snips (667 items): Jev 91.9%, Kev-9B 87.3%, Shisa DE-1 84.9%.
  • Routing, 20 intents. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 3.1%; Bag of words (TF-IDF) 66.7%.
  • Routing, 77 intents. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 4.6%; Bag of words (TF-IDF) 49.2%.
  • Tool routing. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 48.3%; Bag of words (TF-IDF) 49.5%; Option count only 49.2%; Lexical matcher (unsupervised) 63.2%; Option count only (none-of-these detection, AUROC) 0.793. BFCL's irrelevance half marks a request irrelevant when a call cannot be made (a required argument is missing), while this task asks which tool should handle it: 20.7% of that source is mislabelled by our definition, and 308 of its 500 items offer a single tool, so the option count alone predicts the label. The number without it is on the multiple-tool half. Without bfcl-live-irrelevance (500 items): AutoJev-27B 98.0%, Kev-9B 97.4%, Shisa DE-1 97.4%.
  • Complaint routing. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 13.7%; Bag of words (TF-IDF) 66.8%.
  • Commit type. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 19.8%; Bag of words (TF-IDF) 60.6%; Changed paths only (TF-IDF) 54.4%.
  • Search relevance. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 29.2%; Bag of words (TF-IDF) 30.3%; Product text kind only 27.0%.

Datasets and provenance

Tasks from public datasets, each with its licence. Two tasks are scored without a source licensed for non‑commercial use, and two internal test sets are not shown.

TaskItemsSourcesLicence
Prompt injection1,000deepset/prompt-injections, jackhhao/jailbreak-classification, reshabhs/SPML_Chatbot_Prompt_Injection, djapp18/JailbreaksOverTimeApache-2.0, MIT, CC-BY-4.0
Moderationwithout toxicchat667mmathys/openai-moderation-api-evaluation, nvidia/Aegis-AI-Content-Safety-Dataset-2.0MIT, CC-BY-4.0
PII1,000gretelai/gretel-pii-masking-en-v1, beki/privy, nvidia/Nemotron-PIIApache-2.0, MIT, CC-BY-4.0
RAG faithfulnesswithout halueval-dialogue666pminervini/HaluEvalApache-2.0
Off-topic1,000clinc/clinc_oos, mteb/amazon_massive_intent, benayas/snipsCC-BY-3.0, CC-BY-4.0, CC0-1.0
Routing, 20 intents1,000legacy-datasets/banking77CC-BY-4.0
Routing, 77 intents500legacy-datasets/banking77CC-BY-4.0
Tool routing1,000gorilla-llm/Berkeley-Function-Calling-LeaderboardApache-2.0
Complaint routing1,000BEE-spoke-data/consumer-finance-complaintsCC0-1.0
Commit type1,000github.com/angular/angular, github.com/vitejs/viteMIT
Search relevance1,000tasksource/esciApache-2.0
Typed decisions1,965LocalLLaMA/typed-decisionsApache-2.0
Web-agent actions2,329AndeyTait/JevForge-Mind2WebCC BY 4.0
Community sets1,200Praveenrajus/jev-benchCC BY 4.0
  • JailbreaksOverTime: Piet et al. 2025, CC-BY-4.0.
  • Nemotron-PII: NVIDIA, CC-BY-4.0 (attribution requested when results are published).
  • Aegis AI Content Safety 2.0: NVIDIA, CC-BY-4.0.
  • Banking77: Casanueva et al. 2020, CC-BY-4.0.
  • CLINC150: Larson et al. 2019, CC-BY-3.0. MASSIVE: Amazon, CC-BY-4.0. SNIPS: Sonos, CC0-1.0.
  • Mind2Web: Deng et al. 2023, CC-BY-4.0 (JevForge snapshot).
  • Shopping Queries Dataset (ESCI): Reddy et al. 2022, Apache-2.0.
  • Berkeley Function Calling Leaderboard v3 live: Apache-2.0.
  • CFPB consumer complaint narratives: US government public domain (BEE-spoke copy, CC0-1.0).
  • HaluEval qa and summarization: Li et al. 2023 (HotpotQA CC-BY-SA-4.0, CNN/DailyMail).
  • Banking77: Casanueva et al. 2020, CC-BY-4.0.
  • Community sets, via Praveenrajus/jev-bench: LEDGAR (Tuggener et al. 2020, CC-BY-4.0), MASSIVE (Amazon, CC-BY-4.0), Civil Comments (Jigsaw, CC0-1.0), Measuring Hate Speech (UC Berkeley D-Lab, CC-BY-4.0), GoEmotions (Google, Apache-2.0), HelpSteer2 helpfulness and verbosity (NVIDIA, CC-BY-4.0), FEVER gold evidence (Thorne et al. 2018, CC-BY-SA-3.0).
  • Typed decisions: LocalLLaMA/typed-decisions, Apache-2.0.

Footnotes

  1. *Prompt injection, openJev Verdict 1.4: Verdict 1.4 returns a near-constant probability here: the middle 90% of its answers span 0.12 around 0.5, and on off-topic it returned one identical value for all 1,000 items, so this AUROC ranks numerical noise rather than measuring the model. Reading it upside down would not give a real score either. Scoring polarity was checked against the author's own reference adapter and is correct (harness investigation, 2026-09-23).
  2. *Every task, GLiNER2.5-Decide: GLiNER2.5-Decide ran behind our own shim over the author's gliner2 2.0.0 package, in FP16 on a 4 GB laptop GPU; the checkpoint is FP32, which spills out of GPU memory on long inputs on that card. Against FP32 on CPU, FP16 picked the same answer on 20 of 20 short inputs (max probability difference 0.0002) and on 54 of 54 long ones of 1,962 to 4,053 tokens (max difference 0.0017). Nothing is truncated (gliner2's default), so long inputs run past the 512 positions of the DeBERTa encoder; label names and descriptions count toward that length, and the author's guidance for long documents is chunking, which this run does not do. Requests over 512 encoded tokens: routing-77 100%, routing-20 82%, complaint-routing 81%, x-typed-decisions 75%, tool-routing 58%, x-jevbench-community 29%, rag-faithfulness 22%, prompt-injection 16%, moderation 15%, x-mind2web-actions 14%, search-relevance 8%, pii 7%, commit-type 6%, off-topic 0.2%. Even in FP16, requests above about 2,300 tokens spill out of GPU memory: 2 prompt-injection requests (73% of that task's measured time) and 1 tool-routing request (4%), so those latencies describe our card more than the model. Its latency is laptop-tier and not comparable with the cloud rows. There is no independent reference to reproduce these numbers against: Fastino's own test split is private.
  3. *Prompt injection, GLiNER2.5-Decide: gliner2.5-decide: Related specialty (the card names moderation; the release blog demonstrates prompt injection); no jailbreak dataset is named; no overlap with Fastino's published sample (the training set itself is unpublished).
  4. *Prompt injection, Laya-typed: laya-typed: Trained on unnamed jailbreak and prompt-injection data; deepset/prompt-injections is declared held out.
  5. *Prompt injection, openJev Verdict 1.4: verdict-1.4: fitted its calibrator on deepset/prompt-injections train split (`use: calibration`); fitted its calibrator on TrustAIRLab in-the-wild jailbreaks (inside djapp18/JailbreaksOverTime) (`use: calibration`): that is 2 of this suite's 4 sources, 50% of its items. Shown for reference, not ranked here.
  6. *Prompt injection, Laya: laya: Trained on unnamed jailbreak and prompt-injection data; deepset/prompt-injections is declared held out.
  7. *Moderation, AutoJev-27B: autojev-27b: trained on nvidia/Aegis-AI-Content-Safety-Dataset-2.0 test split (`use: weights`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: data.py TRAIN_QUOTAS.
  8. *Moderation, GLiNER2.5-Decide: gliner2.5-decide: The model card names moderation as a specialty of its unpublished synthetic training data; none of our sources is named; no overlap with Fastino's published sample (the training set itself is unpublished).
  9. *Moderation, Laya-typed: laya-typed: Trained on unnamed human-labelled toxicity data, a closely related task.
  10. *Moderation, Laya: laya: Trained on unnamed human-labelled toxicity data, a closely related task.
  11. *Moderation, openJev Verdict 1.4: verdict-1.4: Its base model's training lineage includes toxicity and hate-speech datasets, a closely related task.
  12. *PII, SimpleJev (Qwen3.5-0.8B): This zero-shot readout saturates at its lowest rating bin (every probability between 0.0102 and 0.0143) and answers 'no PII' to all 1,000 documents. The residual ordering runs backwards because our PII negatives are redacted documents whose stand-in phrases name the very categories the questions ask about ('the address on file', 'the customer'), which is a documented property of the dataset. The score is a dataset artefact, not an inverted signal (harness investigation, 2026-09-23).
  13. *PII, GLiNER2.5-Decide: gliner2.5-decide: Fastino's evaluation suite, made by the same generator as the training data, has a contains-PII question; none of our PII sources is named; no overlap with Fastino's published sample (the training set itself is unpublished).
  14. *RAG faithfulness, Laya, Laya-typed: Laya does carry signal on one of this framing's three sub-questions (contradicts: AUROC 61 / 64), but that question also returns the lowest probabilities, so the pre-registered max-combine rule almost always hands the item's score to one of the two uninformative sub-questions instead. The result is a combined score at or slightly below chance for both Laya and Laya-typed. This is a limitation of the max-combine framing on this model, not a polarity error; the single-question 'broad' framing is also at chance for Laya here (harness investigation, 2026-09-23).
  15. *RAG faithfulness, GLiNER2.5-Decide: GLiNER2.5-Decide's scores are marked degenerate here: its probabilities barely vary (the central 90% spans 0.141 on this subset, just under the 0.15 bar; 0.152 on the full test set, just over it). Its balanced accuracy is 57.8%, and its AUROC of 59.3% [55.1, 63.5] shows the compressed scores order the items only slightly better than chance, so read both numbers as fragile.
  16. *RAG faithfulness, GLiNER2.5-Decide: gliner2.5-decide: The model card names yes/no questions over a passage as a capability of its unpublished synthetic training; no NLI or faithfulness dataset is named; no overlap with Fastino's published sample (the training set itself is unpublished).
  17. *RAG faithfulness, Kev-0.8B: kev-0.8b: Trained on MNLI and BoolQ, closely related entailment and passage-QA tasks (no item overlap found).
  18. *RAG faithfulness, Kev-0.6B: kev-0.6b: Trained on MNLI and BoolQ, closely related entailment and passage-QA tasks (no item overlap found).
  19. *RAG faithfulness, openJev Verdict 1.4: verdict-1.4: Its base model's training lineage includes natural-language-inference datasets, a closely related task.
  20. *RAG faithfulness, Laya: laya: Trained on NLI and fact-verification data (unnamed), closely related to faithfulness checking.
  21. *RAG faithfulness, Laya-typed: laya-typed: Trained on NLI and fact-verification data (unnamed), closely related to faithfulness checking.
  22. *Off-topic, GLiNER2.5-Decide: GLiNER2.5-Decide's balanced accuracy here is 49.3%, chance level, and its scores are marked degenerate: its probabilities barely vary (the central 90% spans 0.13). What ordering they have runs the wrong way (AUROC 35.9%): the model gives higher P(yes) to in-scope messages when asked whether a message is outside the listed scope. With a single question its dev AUROC is 0.29 (broad) and 0.26 (detailed); in the chosen decomposed framing the 'unsupported task' question is inverted the same way (test AUROC 0.29) and supplies the max-combined score on 552 of 1,000 items, while 'other subject' alone is correctly oriented (0.71). Scoring polarity is correct: the same yes/no mapping is correctly oriented on every other binary task. This is how the model reads a negated scope question (harness investigation, 2026-09-25).
  23. *Off-topic, AutoJev-27B: autojev-27b: trained on mteb/amazon_massive_intent test (en-US, de-DE) split (`use: weights`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: data.py TRAIN_QUOTAS.
  24. *Off-topic, Eikos-27B: eikos-27b: selected its checkpoint on clinc/clinc_oos (`use: selection`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: NOTICE: evaluation-only, part of the card's 'general battery'.
  25. *Off-topic, Eikos-4B: eikos-4b: selected its checkpoint on clinc/clinc_oos (`use: selection`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: NOTICE: evaluation-only, part of the card's 'general battery'.
  26. *Off-topic, Laya: laya: Trained on unnamed intent-classification data, the task family of this category's sources.
  27. *Off-topic, Laya-typed: laya-typed: Trained on unnamed intent-classification data, the task family of this category's sources.
  28. *Off-topic, openJev Verdict 1.4: verdict-1.4: trained on clinc/clinc_oos oos_train (100 out-of-scope queries only) split (`use: weights`); fitted its calibrator on clinc/clinc_oos test split (`use: calibration`): that is 1 of this suite's 3 sources, 33% of its items. Shown for reference, not ranked here. Declared evidence: training code.
  29. *Off-topic, GLiNER2.5-Decide: gliner2.5-decide: The model card names intent classification as a specialty of its unpublished synthetic training data; CLINC, MASSIVE and SNIPS are not named as training data (MASSIVE and SNIPS were zero-shot evaluation sets in the GLiNER2 paper); no overlap with Fastino's published sample (the training set itself is unpublished).
  30. *Routing, 20 intents and Routing, 77 intents, Kev-0.8B, Kev-0.6B, openJev Verdict 1.4: Option-definition bug: all 13 routing-20 items whose gold intent is get_physical_card are PIN questions, which contradicts our written definition of that option ('the customer asks about getting a physical card'). Jev, which reads only the definition, gets 1 of 13; Kev-0.8B, trained on Banking77, gets 13 of 13. beneficiary_not_allowed has the same problem. A memorised dataset convention shows up as accuracy. With the affected items and the adjudicated label errors removed, routing-20's first tier is a three-way tie of Kev-0.8B, Jev and Kev-0.6B, and Jev joins routing-77's first tier (data-quality audit, 2026-09-22). The audit itself warns that dropping flagged items favours Jev by construction, because the flagged items are the ones Jev confidently missed, so read that re-scored tier as a bound, not a result.
  31. *Routing, 20 intents and Routing, 77 intents, Kev-0.8B, Kev-0.6B: 27 routing-20 and 18 routing-77 messages also appear verbatim in Kev's own locked test set.
  32. *Routing, 20 intents and Routing, 77 intents, Kev-4B: kev-4b: trained on legacy-datasets/banking77 (`use: weights`); selected its checkpoint on legacy-datasets/banking77 test split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: inherited from kev-0.8b; unverified for this checkpoint.
  33. *Routing, 20 intents and Routing, 77 intents, Kev-9B: kev-9b: trained on legacy-datasets/banking77 (`use: weights`); selected its checkpoint on legacy-datasets/banking77 test split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: inherited from kev-0.8b; unverified for this checkpoint.
  34. *Routing, 20 intents and Routing, 77 intents, Kev-0.8B: kev-0.8b: trained on legacy-datasets/banking77 (`use: weights`); selected its checkpoint on legacy-datasets/banking77 test split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: Kev's own development split; also the served temperature.
  35. *Routing, 20 intents and Routing, 77 intents, Eikos-27B: eikos-27b: selected its checkpoint on legacy-datasets/banking77 (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: NOTICE: evaluation-only, part of the card's 'general battery'.
  36. *Routing, 20 intents and Routing, 77 intents, Kev-0.6B: kev-0.6b: trained on legacy-datasets/banking77 (`use: weights`); selected its checkpoint on legacy-datasets/banking77 test split (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: Kev's own development split.
  37. *Routing, 20 intents and Routing, 77 intents, Eikos-4B: eikos-4b: selected its checkpoint on legacy-datasets/banking77 (`use: selection`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: NOTICE: evaluation-only, part of the card's 'general battery'.
  38. *Routing, 20 intents and Routing, 77 intents, GLiNER2.5-Decide: gliner2.5-decide: The model card names customer and banking intent as a specialty of its unpublished synthetic training data; Banking77 is not named; no overlap with Fastino's published sample (the training set itself is unpublished).
  39. *Routing, 20 intents and Routing, 77 intents, openJev Verdict 1.4: verdict-1.4: trained on legacy-datasets/banking77 train split (`use: weights`); fitted its calibrator on legacy-datasets/banking77 test split (`use: calibration`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: training code.
  40. *Routing, 20 intents and Routing, 77 intents, Laya-typed: laya-typed: Trained on unnamed banking-intent data; its authors list Banking77 itself as held out.
  41. *Routing, 20 intents and Routing, 77 intents, Laya: laya: Trained on unnamed banking-intent data; its authors list Banking77 itself as held out.
  42. *Tool routing, GLiNER2.5-Decide: GLiNER2.5-Decide never picks none_of_these (0 of 1,000 items), so it scores 0% on the bfcl-live-irrelevance half and 64.8% on bfcl-live-multiple; its overall 32.4% is below always-majority (50%) and every trivial baseline.
  43. *Complaint routing, GLiNER2.5-Decide: gliner2.5-decide: The model card names email and ticket routing as a specialty of its unpublished synthetic training data; CFPB complaints are not named; no overlap with Fastino's published sample (the training set itself is unpublished).
  44. *Complaint routing, Laya-typed: laya-typed: Trained on support-ticket queue routing (Tobi-Bueck/customer-support-tickets), a closely related task.
  45. *Complaint routing, Laya: laya: Trained on support-ticket queue routing (Tobi-Bueck/customer-support-tickets), a closely related task.
  46. *Search relevance, Laya: laya: Trained on MS MARCO query-passage relevance, a closely related task.
  47. *Search relevance, Laya-typed: laya-typed: Trained on MS MARCO query-passage relevance, a closely related task.
  48. *Typed decisions, Laya-typed: laya-typed: trained on LocalLLaMA/typed-decisions train split (`use: weights`): that is the whole of this suite. Shown for reference, not ranked here. Declared evidence: model card.
  49. *Typed decisions, GLiNER2.5-Decide: gliner2.5-decide: The model card names typed decisions of the same kind (agent completion, severity, routing) as specialties of its unpublished synthetic training; typed-decisions is not named; no overlap with Fastino's published sample (the training set itself is unpublished).
  50. *Web-agent actions, GLiNER2.5-Decide: GLiNER2.5-Decide answers this suite's yes/no questions 'yes' 70.5% of the time against 22.6% in the gold labels, so its yes/no accuracy (42.3%) is below always answering 'no' (77.4%).
  51. *Community sets, GLiNER2.5-Decide: GLiNER2.5-Decide answers this suite's yes/no questions 'yes' 3% of the time against 50% in the gold labels (yes/no accuracy 49.0%), and 4 of its 8 configs are at or below random: civil_comments 48.7 and fever_evidence 49.3 (random 50), helpsteer2 helpfulness 18.7 and verbosity 16.0 (random 20).
  52. *Community sets, AutoJev-27B: autojev-27b: trained on amazon_massive_intent (inside Praveenrajus/jev-bench) test (en-US, de-DE) split (`use: weights`); trained on google/civil_comments (inside Praveenrajus/jev-bench) (`use: weights`): that is 2 of this suite's 8 sources, 25% of its items. Still ranked: the overlap is below the one-third bar. Declared evidence: data.py TRAIN_QUOTAS.
  53. *Community sets, Laya-typed: laya-typed: Trained on unnamed toxicity, rubric-rating and fact-checking data whose descriptions match four of this suite's eight configs; possibly the same datasets.
  54. *Community sets, Kev-0.8B: kev-0.8b: Trained on MNLI, a task close to this suite's fact-checking config (150 of 1,200 items).
  55. *Community sets, Laya: laya: Trained on unnamed toxicity, rubric-rating and fact-checking data whose descriptions match four of this suite's eight configs; possibly the same datasets.
  56. *Community sets, Kev-0.6B: kev-0.6b: Trained on MNLI, a task close to this suite's fact-checking config (150 of 1,200 items).
  57. *Community sets, GLiNER2.5-Decide: gliner2.5-decide: Partial: the model card names moderation, ordinal scores and intents as specialties, related to some configs of this suite; no overlap with Fastino's published sample (the training set itself is unpublished).
  58. *Community sets, openJev Verdict 1.4: verdict-1.4: Its base model's lineage includes toxicity, hate-speech and NLI data, close to three of this suite's eight configs.

Provenance

Release 2026-09-25.1, benchmark commit c372cd5. All 235 runs behind these numbers were made from a working tree with uncommitted harness changes, so the exact code of those runs is not reproducible from a commit alone.

29 cloud‑GPU result rows were completed from the GPU box’s own reports after a copy error cut the local records short; on the 17,553 items both copies hold, they agree on every one.

Verified against the model cards and the dataset licences on 2026-09-25. Spot something outdated? Tell us.

Questions

Why does LangWatch benchmark the model its own product runs on?

Instant Evals runs on Jev, so we wanted to know how far open models are from it before recommending either. We made real judgment calls along the way (the contamination rule, the framings, the decision thresholds, which caveats to show), and every one of them is written down on this page so you can check it. What we did not touch by hand is the ranking itself: the tie test decides who is ahead.

Which model is the winner?

On each task the tie test decides: one side is ahead, or they are tied when the data cannot separate them. Grayed rows were trained on the task’s data; they are shown for reference and never counted.

What does a grayed row mean?

The model was trained on the test items or on another split of the same dataset, so its score measures memory of that dataset more than the skill. The contamination audit graded every model and task pair; the rule and the evidence are in the method section.

Why is a model with no declared training data never grayed?

Models are marked on what their authors disclosed. A model that declares nothing has nothing to mark, and that is not proof it never saw the data.

Why are some sources left out of the published numbers?

Two sources are licensed for non-commercial use only, so the moderation and RAG numbers here are recomputed without them. Internal test sets are not shown at all. The audit also names sources that a trivial classifier solves; those stay in the headline, with the number without them next to it.

Is the latency comparable between the hosted model and the open ones?

Only on the same GPU. Open models ran on a laptop GPU, an L4 or an H100, and each latency names its GPU. Jev’s includes the network round trip to its API.

Can I run the same judges on my own data?

Partly. Instant Evals runs the same hosted Jev judge over your own production history from the LangWatch CLI, so that half is the same judge. The open models are published weights you can serve yourself, but not with this benchmark’s exact setup: the harness, the framings and the decision thresholds used here are not public yet.

Run the hosted judge on your own data.

Instant Evals asks one question of your whole production history from the CLI.

$npx langwatch instant-eval run