Jev beats every tiny open model

On 15 tasks, out of distribution, Jev leads the best open model under 1B by 12 to 68 points.

LangWatch Instant Evals runs on Jev. Check it outThis benchmark was created using LangWatch. Sign up to create your own

Prompt injection: catch rate at 5% false alarms

1,000 test items
  • Jev, hosted API
  • Best open model
  • Other open models
  • Has a footnote (ring + †)
  • Reference baseline
  1. Jev94.6%, 95% interval 92.2 to 96.7, p95 841 ms, cost not shown, flips 2/300
  2. Kev-0.8B56.2%, 95% interval 46.3 to 62.6, p95 230 ms, $0.010* per 1,000, flips 0/300
  3. Kev-0.6B25.4%, 95% interval 16.9 to 36.1, p95 171 ms, $0.0073* per 1,000, flips 0/300
  4. Laya-typed49.2%, 95% interval 42.6 to 58.2, p95 202 ms, $0.0080* per 1,000, flips 0/300
  5. Laya0.0%, 95% interval 0.0 to 0.0, p95 154 ms, $0.0086* per 1,000, flips 0/300
  6. SimpleJev (Qwen3.5-0.8B)15.2%, 95% interval 10.6 to 19.4, degenerate score: this system's probabilities barely vary (the central 90% spans only 0.0010), so its 69.0% score is a tie-break on numerical noise rather than a measurement, p95 298 ms, $0.019* per 1,000, flips 0/300
  7. openJev Verdict 1.44.8%, 95% interval 2.8 to 7.8, degenerate score: this system's probabilities barely vary (the central 90% spans only 0.1207), so its 36.0% score is a tie-break on numerical noise rather than a measurement, p95 57 ms, $0.0043* per 1,000, flips 0/300
  8. SemIf (Qwen3-0.6B)7.8%, 95% interval 5.1 to 11.5, degenerate score: this system's probabilities barely vary (the central 90% spans only 0.0953), so its 46.8% score is a tie-break on numerical noise rather than a measurement, p95 141 ms, $0.0061* per 1,000, flips 0/300
  1. Length only31.0%
  2. Bag of words (TF-IDF)72.6%
Catch rate at 5% false alarms, %, 1,000 test items. Dot = score, line = 95% bootstrap interval. ‡ Degenerate score: a tie‑break on noise. Hollow squares: trivial baselines, not contenders.

Which model wins each task

Jev’s lead over the best open model not trained on the task’s data, per task.

Prompt injection Catch rate at 5% false alarms

Jev
94.6%
Best open
Kev-0.8B 56.2%
Gap
leads by 38.4 pts
No-model baseline
Bag of words (TF-IDF) 72.6%: beats every open model

Moderation AUROC

Jev
90.3%
Best open
Kev-0.8B 76.2%
Gap
leads by 14.1 pts
No-model baseline
Bag of words (TF-IDF) 76.3%: beats every open model

PII Catch rate at 5% false alarms

Jev
90.8%
Best open
Kev-0.8B 22.8%
Gap
leads by 67.9 pts
No-model baseline
Bag of words (TF-IDF) 77.2%: beats every open model

RAG faithfulness Balanced accuracy

Jev
80.3%
Best open
Kev-0.8B 55.5%
Gap
leads by 24.8 pts
No-model baseline
String-overlap heuristic 68.0%: beats every open model

Off-topic Balanced accuracy

Jev
93.4%
Best open
Kev-0.8B 76.8%
Gap
leads by 16.6 pts
No-model baseline
Bag of words (TF-IDF) 83.5%: beats every open model

Routing, 20 intents Accuracy

Jev
89.1%
Best open
Laya-typed 77.0%
Gap
leads by 12.1 pts
Any open model
Kev-0.8B* 91.3%: Jev trails by 2.2 pts

Routing, 77 intents Accuracy

Jev
79.6%
Best open
Laya 42.4%
Gap
leads by 37.2 pts
Any open model
Kev-0.8B* 83.0%: Jev trails by 3.4 pts
No-model baseline
Bag of words (TF-IDF) 49.2%: beats every open model

Tool routing Accuracy

Jev
78.3%
Best open
Kev-0.8B 57.7%
Gap
leads by 20.6 pts
No-model baseline
Lexical matcher (unsupervised) 63.2%: beats every open model

Complaint routing Accuracy

Jev
78.7%
Best open
Kev-0.6B 59.3%
Gap
leads by 19.4 pts
No-model baseline
Bag of words (TF-IDF) 66.8%: beats every open model

Commit type Accuracy

Jev
68.3%
Best open
Laya-typed 49.7%
Gap
leads by 18.6 pts
No-model baseline
Bag of words (TF-IDF) 60.6%: beats every open model

Search relevance Accuracy

Jev
57.7%
Best open
Laya 31.5%
Gap
leads by 26.2 pts

Typed decisions Accuracy

Jev
73.9%
Best open
Kev-0.8B 44.9%
Gap
leads by 29.0 pts
Any open model
Laya-typed* 77.4%: Jev trails by 3.5 pts

Web-agent actions Accuracy

Jev
70.8%
Best open
SimpleJev (Qwen3.5-0.8B) 56.3%
Gap
leads by 14.5 pts

Community sets Accuracy

Jev
62.3%
Best open
Laya-typed 47.8%
Gap
leads by 14.6 pts

JevBench public Accuracy

Jev
85.7%
Best open
Kev-0.6B 60.6%
Gap
leads by 25.1 pts

* Trained on that task’s dataset: shown for reference, never ranked. A model marked * scores higher than Jev on Routing, 20 intents; Routing, 77 intents; and Typed decisions.

No-model baseline: a trivial classifier with no language model (bag of words, string overlap or a lexical matcher), never ranked. It beats every open model on 9 of 11 tasks that have one.

Pick a task

Overlapping lines mean the data cannot separate two models.

Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.

  • Jev, hosted API
  • Best open model
  • Other open models
  • Has a footnote (ring + †)
  • Reference baseline
  1. Jev94.6%, 95% interval 92.2 to 96.7, p95 841 ms, cost not shown, flips 2/300
  2. Kev-0.8B56.2%, 95% interval 46.3 to 62.6, p95 230 ms, $0.010* per 1,000, flips 0/300
  3. Kev-0.6B25.4%, 95% interval 16.9 to 36.1, p95 171 ms, $0.0073* per 1,000, flips 0/300
  4. Laya-typed49.2%, 95% interval 42.6 to 58.2, p95 202 ms, $0.0080* per 1,000, flips 0/300
  5. Laya0.0%, 95% interval 0.0 to 0.0, p95 154 ms, $0.0086* per 1,000, flips 0/300
  6. SimpleJev (Qwen3.5-0.8B)15.2%, 95% interval 10.6 to 19.4, degenerate score: this system's probabilities barely vary (the central 90% spans only 0.0010), so its 69.0% score is a tie-break on numerical noise rather than a measurement, p95 298 ms, $0.019* per 1,000, flips 0/300
  7. openJev Verdict 1.44.8%, 95% interval 2.8 to 7.8, degenerate score: this system's probabilities barely vary (the central 90% spans only 0.1207), so its 36.0% score is a tie-break on numerical noise rather than a measurement, p95 57 ms, $0.0043* per 1,000, flips 0/300
  8. SemIf (Qwen3-0.6B)7.8%, 95% interval 5.1 to 11.5, degenerate score: this system's probabilities barely vary (the central 90% spans only 0.0953), so its 46.8% score is a tie-break on numerical noise rather than a measurement, p95 141 ms, $0.0061* per 1,000, flips 0/300
  1. Length only31.0%
  2. Bag of words (TF-IDF)72.6%
Catch rate at 5% false alarms, %, 1,000 test items. Dot = score, line = 95% bootstrap interval. ‡ Degenerate score: a tie‑break on noise. Hollow squares: trivial baselines, not contenders.
ModelCoverageCut inputsFlipsp95 ms$ / 1k
Jev100%0%2/300841not shown
Kev-0.8B100%n/a0/300230$0.010*
Kev-0.6B100%n/a0/300171$0.0073*
Laya-typed100%5%0/300202$0.0080*
Laya100%16%0/300154$0.0086*
SimpleJev (Qwen3.5-0.8B)100%n/a0/300298$0.019*
openJev Verdict 1.4100%n/a0/30057$0.0043*
SemIf (Qwen3-0.6B)100%n/a0/300141$0.0061*
Caveats for this task

Flips: answers that changed between two identical runs on the first 300 items.

Without spml (750 items), from the data‑quality audit: SPML is separable by length alone (length-only AUROC 0.998): its benign items are short questions and its injections long persona prompts. Without it, the best open model changes.

Rows without a score

  • SimpleJev (Qwen3.5-0.8B): Degenerate score. this system's probabilities barely vary (the central 90% spans only 0.0010), so its 69.0% score is a tie-break on numerical noise rather than a measurement
  • SemIf (Qwen3-0.6B): Degenerate score. this system's probabilities barely vary (the central 90% spans only 0.0953), so its 46.8% score is a tie-break on numerical noise rather than a measurement
  • openJev Verdict 1.4: Degenerate score. this system's probabilities barely vary (the central 90% spans only 0.1207), so its 36.0% score is a tie-break on numerical noise rather than a measurement
  • Laya: Threshold unreachable. no threshold reaches the false-alarm budget: a block of benign items scores at the maximum, although AUROC 81.6% says the model ranks the classes apart

Caveats

  • Each model's question wording and decision threshold were picked on a separate 60-item dev split, before the 1,000-item test run.
  • Inputs over the set's length limit (2,500 to 12,000 characters, per set) were removed before sampling, for every model alike.
  • Laya and Verdict read at most 512 tokens (Laya-typed 1,024). Longer inputs are cut by the model's own code. Cut items stay in the score; the share of cut inputs per model is shown next to coverage where the server reports token usage (Laya and Laya-typed), and cannot be measured for Verdict, which reports none.
  • Open models ran one at a time on a laptop GPU (NVIDIA RTX 3050 Ti Laptop, 4 GB, WSL2) at concurrency 1: laptop-tier latency. A data-centre GPU is faster.
  • Jev latency is measured at our client on the same laptop and includes the network round trip to the hosted API.
  • Local cost per 1,000 decisions is an assumption: $0.35 per GPU-hour (a cloud-equivalent rate) times the measured time at concurrency 1.
  • Jev cost per 1,000 decisions is the recorded token usage at the vendor's public list price (output tokens free).
  • Flip rate: answers that changed when the same items were sent again.
  • Labels are the datasets' own; no human re-check yet.
  • Reference baselines are trivial classifiers fitted with 5-fold cross-validation on the test items themselves (length only, a bag-of-words TF-IDF logistic regression, a few regexes). They are not contenders: a baseline near a model means the set can be solved without reading it the way a model does.
  • A data-quality audit of the frozen sets (2026-09-22) found sources that a trivial baseline solves; the headline stays on the full frozen set that was run, with the number without that source next to it.

Every task at a glance

Primary score per task, in percent. Darker is better.

  • Prompt injectionCatch rate at 5% false alarmsJev 94.6%, top tierKev-0.8B 56.2%Kev-0.6B 25.4%Laya-typed 49.2%Laya 0.0%SimpleJev (Qwen3.5-0.8B) 15.2%openJev Verdict 1.4 4.8%SemIf (Qwen3-0.6B) 7.8%
  • ModerationAUROCJev 90.3%, top tierKev-0.8B 76.2%Kev-0.6B 65.3%Laya-typed 70.0%Laya 69.8%SimpleJev (Qwen3.5-0.8B) 49.9%openJev Verdict 1.4 51.8%SemIf (Qwen3-0.6B) 63.0%
  • PIICatch rate at 5% false alarmsJev 90.8%, top tierKev-0.8B 22.8%Kev-0.6B 9.6%Laya-typed 21.0%Laya 15.8%SimpleJev (Qwen3.5-0.8B) 3.4%openJev Verdict 1.4 1.0%SemIf (Qwen3-0.6B) 9.6%
  • RAG faithfulnessBalanced accuracyJev 80.3%, top tierKev-0.8B 55.5%Kev-0.6B 54.1%Laya-typed 49.4%Laya 49.8%SimpleJev (Qwen3.5-0.8B) 50.8%openJev Verdict 1.4 50.3%SemIf (Qwen3-0.6B) 52.1%
  • Off-topicBalanced accuracyJev 93.4%, top tierKev-0.8B 76.8%Kev-0.6B 55.8%Laya-typed 52.6%Laya 54.3%SimpleJev (Qwen3.5-0.8B) 58.8%openJev Verdict 1.4 constant outputSemIf (Qwen3-0.6B) 53.5%
  • Routing, 20 intentsAccuracyJev 89.1%, top tierKev-0.8B 91.3%, trained on this datasetKev-0.6B 89.5%, trained on this datasetLaya-typed 77.0%Laya 76.0%SimpleJev (Qwen3.5-0.8B) 60.2%openJev Verdict 1.4 78.4%, trained on this datasetSemIf (Qwen3-0.6B) not scored
  • Routing, 77 intentsAccuracyJev 79.6%, top tierKev-0.8B 83.0%, trained on this datasetKev-0.6B 78.2%, trained on this datasetLaya-typed 40.0%Laya 42.4%SimpleJev (Qwen3.5-0.8B) not scoredopenJev Verdict 1.4 not scoredSemIf (Qwen3-0.6B) not scored
  • Tool routingAccuracyJev 78.3%, top tierKev-0.8B 57.7%Kev-0.6B 44.0%Laya-typed 21.5%Laya 22.9%SimpleJev (Qwen3.5-0.8B) 29.1%openJev Verdict 1.4 16.4%SemIf (Qwen3-0.6B) 42.6%
  • Complaint routingAccuracyJev 78.7%, top tierKev-0.8B 38.5%Kev-0.6B 59.3%Laya-typed 47.1%Laya 44.2%SimpleJev (Qwen3.5-0.8B) 45.2%openJev Verdict 1.4 46.3%SemIf (Qwen3-0.6B) 37.8%
  • Commit typeAccuracyJev 68.3%, top tierKev-0.8B 48.6%Kev-0.6B 41.6%Laya-typed 49.7%Laya 44.8%SimpleJev (Qwen3.5-0.8B) 32.7%openJev Verdict 1.4 26.5%SemIf (Qwen3-0.6B) 23.9%
  • Search relevanceAccuracyJev 57.7%, top tierKev-0.8B 29.3%Kev-0.6B 26.2%Laya-typed 29.8%Laya 31.5%SimpleJev (Qwen3.5-0.8B) 25.0%openJev Verdict 1.4 24.8%SemIf (Qwen3-0.6B) 25.0%
  • Typed decisionsAccuracyJev 73.9%, top tierKev-0.8B 44.9%Kev-0.6B 44.8%Laya-typed 77.4%, trained on this datasetLaya 36.1%SimpleJev (Qwen3.5-0.8B) 39.0%openJev Verdict 1.4 36.4%SemIf (Qwen3-0.6B) 35.4%
  • Web-agent actionsAccuracyJev 70.8%, top tierKev-0.8B 50.9%Kev-0.6B 24.8%Laya-typed 18.3%Laya 17.8%SimpleJev (Qwen3.5-0.8B) 56.3%openJev Verdict 1.4 18.5%SemIf (Qwen3-0.6B) 17.9%
  • Community setsAccuracyJev 62.3%, top tierKev-0.8B 46.4%Kev-0.6B 45.3%Laya-typed 47.8%Laya 45.9%SimpleJev (Qwen3.5-0.8B) 25.0%openJev Verdict 1.4 22.6%SemIf (Qwen3-0.6B) 26.2%
  • JevBench publicAccuracyJev 85.7%, top tierKev-0.8B 60.2%Kev-0.6B 60.6%Laya-typed 53.7%Laya 58.9%SimpleJev (Qwen3.5-0.8B) 55.4%openJev Verdict 1.4 57.6%, trained on this datasetSemIf (Qwen3-0.6B) 48.9%
0100Dots left to right: Jev, Kev-0.8B, Kev-0.6B, Laya-typed, Laya, SimpleJev (Qwen3.5-0.8B), openJev Verdict 1.4, SemIf (Qwen3-0.6B).ringed: top tier (the tie test finds no model above it).* hatched: trained on this dataset, never ranked. – not scored. ? under investigation. = constant output.

How the numbers were made

Frozen test sets, 95% bootstrap intervals, a paired tie test and a contamination audit.

Full method, datasets and footnotes
  1. 1

    Frozen splits

    Each use case has a 60‑item dev split and a 1,000‑item test split (500 for 77‑way routing), frozen with checksums before any model ran.

  2. 2

    Wording chosen on dev

    Every model got each task in two or three wordings; the wording and the decision threshold were picked on the dev split, once, then frozen.

  3. 3

    Intervals and tiers

    Scores carry a 95% percentile bootstrap interval over items (2,000 resamples). Tiers come from a paired McNemar test on accuracy, or a paired bootstrap of the metric difference on the items every model shared.

  4. 4

    Latency and cost

    Open models ran one at a time on a laptop GPU (NVIDIA RTX 3050 Ti, 4 GB), so their latency is laptop tier. Jev’s latency is measured from the same laptop and includes the network. Local cost is $0.35 per GPU‑hour times measured time. Jev has the highest p95 on 10 of 11 real tasks.

  5. 5

    Long inputs

    Inputs over each set’s length limit (2,500 to 12,000 characters) were removed for every model alike before sampling. Inputs a model cuts at its own token window stay in the score; the share of cut inputs is shown next to coverage.

  6. 6

    Contamination

    Each model and task pair was audited against the model’s published training, selection and calibration data. Grayed rows trained on the test items or on the same dataset; footnoted rows trained on a related task.

Applies to the whole page

  • LangWatch's Instant Evals runs on Jev, and LangWatch ran this benchmark.
  • Jev's training data is not disclosed, so its contamination cannot be checked.
  • Runs were made from a working tree with uncommitted changes to the harness; the exact code is not reproducible from a commit alone.

The contamination rule

Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner. Grayed rows are shown for reference, never ranked.

Jev’s training data is not disclosed, and every open model starts from a pretrained base whose corpus is not itemised. An unmarked row means no known contamination, not proof that there is none.

What the data‑quality audit found

A trivial baseline fitted on the test items beats every open model on 9 of 11 tasks that have one. The notes below say what a better set would change.

  • Prompt injection. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 31.0%; Bag of words (TF-IDF) 72.6%. SPML is separable by length alone (length-only AUROC 0.998): its benign items are short questions and its injections long persona prompts. Without it, the best open model changes. Without spml (750 items): Jev 92.5%, Laya-typed 67.7%, Kev-0.8B 53.9%.
  • Moderation. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 57.6%; Bag of words (TF-IDF) 76.3%.
  • PII. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 3.2%; Bag of words (TF-IDF) 77.2%; Six regexes 68.7%. The gretel and nemotron negatives carry a redaction stand-in ('on file', 'the customer') that no positive has, so a bag-of-words classifier separates them perfectly (AUROC 1.000). The number without them is on privy alone, the one source without that artefact, where the gap between Jev and the best open model is a fraction of the headline gap. Without gretel and nemotron (333 items): Jev 80.7%, Kev-0.8B 61.4%, Kev-0.6B 43.4%.
  • RAG faithfulness. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 67.8%; Bag of words (TF-IDF) 54.6%; String-overlap heuristic 68.0%. HaluEval-qa is a length and style detector: 92.8% of its hallucinated answers end in a full stop against 1.2% of the correct ones, and a five-feature string-overlap heuristic (AUROC 0.977) beats Jev (0.934) on it. The number without it is on HaluEval summarization alone. Without halueval-dialogue and halueval-qa (333 items): Jev 75.9%, Kev-0.6B 54.9%, Kev-0.8B 53.3%.
  • Off-topic. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 50.8%; Bag of words (TF-IDF) 83.5%. SNIPS is separable by keywords (TF-IDF AUROC 0.999): it holds only two out-of-scope intent types. Without snips (667 items): Jev 91.9%, Kev-0.8B 72.4%, SimpleJev (Qwen3.5-0.8B) 56.5%.
  • Routing, 20 intents. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 3.1%; Bag of words (TF-IDF) 66.7%.
  • Routing, 77 intents. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 4.6%; Bag of words (TF-IDF) 49.2%.
  • Tool routing. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 48.3%; Bag of words (TF-IDF) 49.5%; Option count only 49.2%; Lexical matcher (unsupervised) 63.2%; Option count only (none-of-these detection, AUROC) 0.793. BFCL's irrelevance half marks a request irrelevant when a call cannot be made (a required argument is missing), while this task asks which tool should handle it: 20.7% of that source is mislabelled by our definition, and 308 of its 500 items offer a single tool, so the option count alone predicts the label. The number without it is on the multiple-tool half. Without bfcl-live-irrelevance (500 items): Jev 96.8%, Kev-0.8B 83.8%, Kev-0.6B 77.0%.
  • Complaint routing. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 13.7%; Bag of words (TF-IDF) 66.8%.
  • Commit type. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 19.8%; Bag of words (TF-IDF) 60.6%; Changed paths only (TF-IDF) 54.4%.
  • Search relevance. Reference baselines, fitted with 5‑fold cross‑validation on the test items themselves: Length only 29.2%; Bag of words (TF-IDF) 30.3%; Product text kind only 27.0%.

Datasets and provenance

15 tasks from public datasets, each with its licence. Two sources licensed for non‑commercial use are left out of the published numbers, and two internal test sets are not shown.

TaskItemsSourcesLicence
Prompt injection1,000deepset/prompt-injections, jackhhao/jailbreak-classification, reshabhs/SPML_Chatbot_Prompt_Injection, djapp18/JailbreaksOverTimeApache-2.0, MIT, CC-BY-4.0
Moderationwithout toxicchat667mmathys/openai-moderation-api-evaluation, nvidia/Aegis-AI-Content-Safety-Dataset-2.0MIT, CC-BY-4.0
PII1,000gretelai/gretel-pii-masking-en-v1, beki/privy, nvidia/Nemotron-PIIApache-2.0, MIT, CC-BY-4.0
RAG faithfulnesswithout halueval-dialogue666pminervini/HaluEvalApache-2.0
Off-topic1,000clinc/clinc_oos, mteb/amazon_massive_intent, benayas/snipsCC-BY-3.0, CC-BY-4.0, CC0-1.0
Routing, 20 intents1,000legacy-datasets/banking77CC-BY-4.0
Routing, 77 intents500legacy-datasets/banking77CC-BY-4.0
Tool routing1,000gorilla-llm/Berkeley-Function-Calling-LeaderboardApache-2.0
Complaint routing1,000BEE-spoke-data/consumer-finance-complaintsCC0-1.0
Commit type1,000github.com/angular/angular, github.com/vitejs/viteMIT
Search relevance1,000tasksource/esciApache-2.0
Typed decisions1,965LocalLLaMA/typed-decisionsApache-2.0
Web-agent actions2,329AndeyTait/JevForge-Mind2WebCC BY 4.0
Community sets1,200Praveenrajus/jev-benchCC BY 4.0
JevBench public231fstandhartinger/jevbenchMIT
  • JailbreaksOverTime: Piet et al. 2025, CC-BY-4.0.
  • Nemotron-PII: NVIDIA, CC-BY-4.0 (attribution requested when results are published).
  • Aegis AI Content Safety 2.0: NVIDIA, CC-BY-4.0.
  • Banking77: Casanueva et al. 2020, CC-BY-4.0.
  • CLINC150: Larson et al. 2019, CC-BY-3.0. MASSIVE: Amazon, CC-BY-4.0. SNIPS: Sonos, CC0-1.0.
  • Mind2Web: Deng et al. 2023, CC-BY-4.0 (JevForge snapshot).
  • Shopping Queries Dataset (ESCI): Reddy et al. 2022, Apache-2.0.
  • Berkeley Function Calling Leaderboard v3 live: Apache-2.0.
  • CFPB consumer complaint narratives: US government public domain (BEE-spoke copy, CC0-1.0).
  • HaluEval qa and summarization: Li et al. 2023 (HotpotQA CC-BY-SA-4.0, CNN/DailyMail).
  • Banking77: Casanueva et al. 2020, CC-BY-4.0.
  • Community sets, via Praveenrajus/jev-bench: LEDGAR (Tuggener et al. 2020, CC-BY-4.0), MASSIVE (Amazon, CC-BY-4.0), Civil Comments (Jigsaw, CC0-1.0), Measuring Hate Speech (UC Berkeley D-Lab, CC-BY-4.0), GoEmotions (Google, Apache-2.0), HelpSteer2 helpfulness and verbosity (NVIDIA, CC-BY-4.0), FEVER gold evidence (Thorne et al. 2018, CC-BY-SA-3.0).
  • Typed decisions: LocalLLaMA/typed-decisions, Apache-2.0.
  • JevBench public items: fstandhartinger/jevbench, MIT.

Footnotes

  1. *Prompt injection, openJev Verdict 1.4: Verdict 1.4 returns a near-constant probability here: the middle 90% of its answers span 0.12 around 0.5, and on off-topic it returned one identical value for all 1,000 items, so this AUROC ranks numerical noise rather than measuring the model. Reading it upside down would not give a real score either. Scoring polarity was checked against the author's own reference adapter and is correct (harness investigation, 2026-09-23).
  2. *Prompt injection, Laya-typed, Laya: Trained on unnamed jailbreak and prompt-injection data; deepset/prompt-injections is declared held out.
  3. *Prompt injection, openJev Verdict 1.4: Its confidence calibrator (not the model weights) was fitted on prompts from two of this category's sources.
  4. *Moderation, Laya-typed, Laya: Trained on unnamed human-labelled toxicity data, a closely related task.
  5. *Moderation, openJev Verdict 1.4: Its base model's training lineage includes toxicity and hate-speech datasets, a closely related task.
  6. *PII, SimpleJev (Qwen3.5-0.8B): This zero-shot readout saturates at its lowest rating bin (every probability between 0.0102 and 0.0143) and answers 'no PII' to all 1,000 documents. The residual ordering runs backwards because our PII negatives are redacted documents whose stand-in phrases name the very categories the questions ask about ('the address on file', 'the customer'), which is a documented property of the dataset. The score is a dataset artefact, not an inverted signal (harness investigation, 2026-09-23).
  7. *RAG faithfulness, Laya, Laya-typed: Laya does carry signal on one of this framing's three sub-questions (contradicts: AUROC 61 / 64), but that question also returns the lowest probabilities, so the pre-registered max-combine rule almost always hands the item's score to one of the two uninformative sub-questions instead. The result is a combined score at or slightly below chance for both Laya and Laya-typed. This is a limitation of the max-combine framing on this model, not a polarity error; the single-question 'broad' framing is also at chance for Laya here (harness investigation, 2026-09-23).
  8. *RAG faithfulness, Kev-0.8B, Kev-0.6B: Trained on MNLI and BoolQ, closely related entailment and passage-QA tasks (no item overlap found).
  9. *RAG faithfulness, openJev Verdict 1.4: Its base model's training lineage includes natural-language-inference datasets, a closely related task.
  10. *RAG faithfulness, Laya, Laya-typed: Trained on NLI and fact-verification data (unnamed), closely related to faithfulness checking.
  11. *Off-topic, Laya, Laya-typed: Trained on unnamed intent-classification data, the task family of this category's sources.
  12. *Off-topic, openJev Verdict 1.4: Trained on CLINC150's out-of-scope queries, the source of a third of this category, and calibrated on CLINC test queries; shown for reference, not ranked.
  13. *Routing, 20 intents, Kev-0.8B, Kev-0.6B, openJev Verdict 1.4: Option-definition bug: all 13 routing-20 items whose gold intent is get_physical_card are PIN questions, which contradicts our written definition of that option ('the customer asks about getting a physical card'). Jev, which reads only the definition, gets 1 of 13; Kev-0.8B, trained on Banking77, gets 13 of 13. beneficiary_not_allowed has the same problem. A memorised dataset convention shows up as accuracy. With the affected items and the adjudicated label errors removed, routing-20's first tier is a three-way tie of Kev-0.8B, Jev and Kev-0.6B, and Jev joins routing-77's first tier (data-quality audit, 2026-09-22). The audit itself warns that dropping flagged items favours Jev by construction, because the flagged items are the ones Jev confidently missed, so read that re-scored tier as a bound, not a result.
  14. *Routing, 20 intents, Kev-0.8B, Kev-0.6B: 27 routing-20 and 18 routing-77 messages also appear verbatim in Kev's own locked test set.
  15. *Routing, 20 intents, Kev-0.8B, Kev-0.6B: Trained on Banking77, the dataset behind this category (train split), and its checkpoint was picked on a set holding 27 of these 1,000 test messages; shown for reference, not ranked.
  16. *Routing, 20 intents, openJev Verdict 1.4: Trained on Banking77, the dataset behind this category (train split), and calibrated on Banking77 test messages; shown for reference, not ranked.
  17. *Routing, 20 intents, Laya-typed, Laya: Trained on unnamed banking-intent data; its authors list Banking77 itself as held out.
  18. *Routing, 77 intents, Kev-0.8B, Kev-0.6B: Trained on Banking77, the dataset behind this category (train split), and its checkpoint was picked on a set holding 16 of these 500 test messages; shown for reference, not ranked.
  19. *Complaint routing, Laya-typed, Laya: Trained on support-ticket queue routing (Tobi-Bueck/customer-support-tickets), a closely related task.
  20. *Search relevance, Laya, Laya-typed: Trained on MS MARCO query-passage relevance, a closely related task.
  21. *Typed decisions, Laya-typed: Fine-tuned on this dataset's training split (same generator, workflows and questions); shown for reference, not ranked.
  22. *Community sets, Laya-typed, Laya: Trained on unnamed toxicity, rubric-rating and fact-checking data whose descriptions match four of this suite's eight configs; possibly the same datasets.
  23. *Community sets, Kev-0.8B, Kev-0.6B: Trained on MNLI, a task close to this suite's fact-checking config (150 of 1,200 items).
  24. *Community sets, openJev Verdict 1.4: Its base model's lineage includes toxicity, hate-speech and NLI data, close to three of this suite's eight configs.
  25. *JevBench public, Kev-0.6B, Kev-0.8B: Trained on template-generated policy-rule cases like this set's policy items (no item overlap found).
  26. *JevBench public, openJev Verdict 1.4: openJev Verdict 1.4's inference settings were chosen by measuring on these same 231 public items; shown for reference, not ranked.

Provenance

Release 2026-09-23.1, benchmark commit 58b54bf. 120 of the runs behind these numbers were made from a working tree with uncommitted harness changes, so the exact code of those runs is not reproducible from a commit alone.

Jev’s cost per 1,000 decisions is not shown on this release.

Verified against the model cards, the dataset licences and the public list price on 2026-09-22. Spot something outdated? Tell us.

Questions

Why does LangWatch benchmark the model its own product runs on?

Instant Evals runs on Jev, so we wanted to know how far the open Jev-class models are from it before recommending either. We made real judgment calls along the way (the contamination rule, the framings, the decision thresholds, which caveats to show), and every one of them is written down on this page so you can check it. What we did not touch by hand is the ranking itself: the tie test decides who leads on the numbers.

Which model is the winner?

Each task shows its own leader, and two models in the same tier cannot be separated by the data. Grayed rows were trained on the task’s dataset; they are shown for reference and are not counted as the best open model.

What does a grayed row mean?

The model was trained on the test items or on another split of the same dataset, so its score measures memory of that dataset more than the skill. The contamination audit graded every model and task pair; the rule and the evidence are in the method section.

Why are some sources left out of the published numbers?

Two sources are licensed for non-commercial use only, so the moderation and RAG numbers here are recomputed without them. Two internal test sets are not shown at all. The audit also names sources that a trivial classifier solves; those stay in the headline, with the number without them next to it.

Is the latency comparable between the hosted model and the open ones?

Only within its kind. Open models ran one at a time on a laptop GPU, so their latency is laptop tier and a data-centre GPU is faster. Jev’s latency is measured from the same laptop and includes the network round trip to its API.

How is the cost per 1,000 decisions computed?

For the open models it is an assumption: a cloud-equivalent GPU rate times the measured time at concurrency 1, marked as assumed wherever it appears. Jev’s cost is not shown on this release: it is confidential under the vendor’s customer agreement until legal clears it for a public page.

Can I run the same judges on my own data?

Partly. Instant Evals runs the same hosted Jev judge over your own production history from the LangWatch CLI, so that half is the same judge. The open models are published weights you can serve yourself, but not with this benchmark’s exact setup: the harness, the framings and the decision thresholds used here are not public yet.

Run the hosted judge on your own data.

Instant Evals asks one question of your whole production history from the CLI.

$npx langwatch instant-eval run