Jev against 7 open models under 1B parameters on 15 tasks, with 95% intervals and tie tiers. Rows trained on a task's data are grayed and never lead. Pick a task, filter, or click a model.
Laya, convaiinnovations: 0.0%, interval 0.0 to 0.0, not ranked
Ahead aloneTied for the leadThe restTrained on the test data*95% interval
One task at a time
Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.
Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.
Is a message harmful content a moderation filter should flag (sexual, hate, violence, harassment, self-harm, sexual content involving minors)? Real user prompts sent to chatbots plus a curated set with per-category human labels.
Does a business document (invoice, ticket, record, transaction) contain personal data? Negatives are real documents with the personal data replaced by neutral wording.
Does the answer (or summary, or dialogue reply) add or contradict facts the provided context does not support? The positive class is a hallucinated answer, so a higher score means the model catches more hallucinations at the same rate of false alarms, not that its answers are more faithful.
Is a user request outside what the assistant supports? 150 supported intents across 10 domains versus real out-of-scope requests.
Pick the right banking intent for a customer message from 20 candidates (gold plus 19 distractors).
The same messages with all 77 banking intents as options.
Real user requests from BFCL "live" with the tools an application exposed (2-13 tools each, with their descriptions and parameters); pick the tool that should be called, or "none of these". Half the items have a matching tool, half are requests no listed tool can serve.
Real consumer complaints filed with the US CFPB, routed to one of 11 current product queues (credit reporting, cards, mortgages, loans, money transfer, ...). Balanced across products and years 2015-2024.
Commits from Angular and Vite: from the subject (type prefix removed), body and changed files, pick the conventional-commit type (build, ci, docs, feat, fix, perf, refactor, test). Balanced across types.
Amazon ESCI: for a real shopping query and a product listing, decide Exact, Substitute, Complement or Irrelevant, using the dataset's own guideline definitions. Balanced across the four labels.
400 synthetic workflow cases (agent-trace observability, customer service, invoice processing, security incidents), each one state with 5 typed questions (choice, noul and score) asked in one request.
Web-agent steps from Mind2Web converted by JevForge: given a task, the operation and 12 candidate page elements (accessibility-tree text), pick the element to act on next (choice over e1..e12) and judge whether one named candidate is a correct target (noul). 800 records from 8 in-domain websites and 386 from 4 held-out websites.
Up to 150 test items from each of 8 licence-clean configs of the community jev-bench (Jevify v0.1.1): LEDGAR contract clauses (100-way choice), MASSIVE assistant intents (60-way), GoEmotions (28-way), Civil Comments toxicity and FEVER claim-vs-evidence (noul), Measuring Hate Speech (3-level score) and HelpSteer2 helpfulness and verbosity (5-level scores).
231 public items from the community JevBench leaderboard (v1.2.14): 48 easy (intent, explicit facts, enum extraction, tool choice), 72 standard (short policy, intent, extraction, ordinal, adequacy and routing questions, 36 paraphrase pairs) and 111 hard (2-4k-token policy documents, multi-hop reasoning, date and number arithmetic, probability, traps, adversarial wording).
Prompt injection: Catch rate at 5% false alarms, higher is better
SimpleJev (Qwen3.5-0.8B)SimpleJev (Qwen3.5-0.8B): 15.2% [10.6, 19.4]. Not ranked.SimpleJev (Qwen3.5-0.8B): 15.2% [10.6, 19.4]. Not ranked.15.2%, interval 15.2% [10.6, 19.4], not ranked
SemIf (Qwen3-0.6B)SemIf (Qwen3-0.6B): 7.8% [5.1, 11.5]. Not ranked.SemIf (Qwen3-0.6B): 7.8% [5.1, 11.5]. Not ranked.7.8%, interval 7.8% [5.1, 11.5], not ranked
openJev Verdict 1.4openJev Verdict 1.4: 4.8% [2.8, 7.8]. Not ranked.openJev Verdict 1.4: 4.8% [2.8, 7.8]. Not ranked.4.8%, interval 4.8% [2.8, 7.8], not ranked
LayaLaya: 0.0% [0.0, 0.0]. Not ranked.Laya: 0.0% [0.0, 0.0]. Not ranked.0.0%, interval 0.0% [0.0, 0.0], not ranked
020406080100
Dot: catch rate at 5% false alarms. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead.
x
Dot: catch rate at 5% false alarms. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: auroc. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: catch rate at 5% false alarms. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
Plot
Hardware
Prompt injection: catch rate at 5% false alarms (higher is better) vs cost
SimpleJev (Qwen3.5-0.8B): 15.2% [10.6, 19.4], ~$0.019, not ranked.
SemIf (Qwen3-0.6B): 7.8% [5.1, 11.5], ~$0.006, not ranked.
$0.001$0.002$0.005$0.01$0.02$0.05$0.1← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev, SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev, SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev, openJev Verdict 1.4, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev, openJev Verdict 1.4, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~). Not plotted (no cost measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Laptop GPU. Not plotted (no latency on Laptop GPU measured): Jev.
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): Laya, Laya-typed, openJev Verdict 1.4, Kev-0.6B, Kev-0.8B, SimpleJev (Qwen3.5-0.8B), SemIf (Qwen3-0.6B).
Every model, every task
Every visible model on every task, sorted by Prompt injection. Value with half its interval; bands are tie tiers.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Training data not disclosed, so contamination can't be ruled out.
How we measured
Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the metric difference, on the items every model answered.
Tie tiers
Models in one tier can't be separated by the tie test (paired bootstrap of the TPR@5%FPR difference on shared items, 95% CI includes 0; paired bootstrap of the AUROC difference on shared items, 95% CI includes 0; paired bootstrap of the Balanced acc. difference on shared items, 95% CI includes 0; paired McNemar on per-item correctness at the dev threshold, p >= 0.05; paired McNemar on per-item correctness, p >= 0.05). Tiers are computed in the export over every model; filters on this page never change them.
Contamination
Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner.
Moderation (667 test items): mmathys/openai-moderation-api-evaluation (MIT), nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (CC-BY-4.0)
PII (1,000 test items): gretelai/gretel-pii-masking-en-v1 (Apache-2.0), beki/privy (MIT), nvidia/Nemotron-PII (CC-BY-4.0)
RAG faithfulness (666 test items): pminervini/HaluEval (Apache-2.0)
Off-topic (1,000 test items): clinc/clinc_oos (CC-BY-3.0), mteb/amazon_massive_intent (CC-BY-4.0), benayas/snips (CC0-1.0)
Routing, 20 intents (1,000 test items): legacy-datasets/banking77 (CC-BY-4.0)
Routing, 77 intents (500 test items): legacy-datasets/banking77 (CC-BY-4.0)
Tool routing (1,000 test items): gorilla-llm/Berkeley-Function-Calling-Leaderboard (Apache-2.0)
Complaint routing (1,000 test items): BEE-spoke-data/consumer-finance-complaints (CC0-1.0)
Commit type (1,000 test items): github.com/angular/angular (MIT), github.com/vitejs/vite (MIT)
Search relevance (1,000 test items): tasksource/esci (Apache-2.0)
Typed decisions (1,965 test items): LocalLLaMA/typed-decisions (Apache-2.0)
Web-agent actions (2,329 test items): AndeyTait/JevForge-Mind2Web (CC BY 4.0)
Community sets (1,200 test items): Praveenrajus/jev-bench (CC BY 4.0)
JevBench public (231 test items): fstandhartinger/jevbench (MIT)
Trained on unnamed jailbreak and prompt-injection data; deepset/prompt-injections is declared held out.
Its confidence calibrator (not the model weights) was fitted on prompts from two of this category's sources.
Verdict 1.4 returns a near-constant probability here: the middle 90% of its answers span 0.12 around 0.5, and on off-topic it returned one identical value for all 1,000 items, so this AUROC ranks numerical noise rather than measuring the model. Reading it upside down would not give a real score either. Scoring polarity was checked against the author's own reference adapter and is correct (harness investigation, 2026-09-23).
Trained on unnamed human-labelled toxicity data, a closely related task.
Its base model's training lineage includes toxicity and hate-speech datasets, a closely related task.
This zero-shot readout saturates at its lowest rating bin (every probability between 0.0102 and 0.0143) and answers 'no PII' to all 1,000 documents. The residual ordering runs backwards because our PII negatives are redacted documents whose stand-in phrases name the very categories the questions ask about ('the address on file', 'the customer'), which is a documented property of the dataset. The score is a dataset artefact, not an inverted signal (harness investigation, 2026-09-23).
Trained on MNLI and BoolQ, closely related entailment and passage-QA tasks (no item overlap found).
Its base model's training lineage includes natural-language-inference datasets, a closely related task.
Trained on NLI and fact-verification data (unnamed), closely related to faithfulness checking.
Laya does carry signal on one of this framing's three sub-questions (contradicts: AUROC 61 / 64), but that question also returns the lowest probabilities, so the pre-registered max-combine rule almost always hands the item's score to one of the two uninformative sub-questions instead. The result is a combined score at or slightly below chance for both Laya and Laya-typed. This is a limitation of the max-combine framing on this model, not a polarity error; the single-question 'broad' framing is also at chance for Laya here (harness investigation, 2026-09-23).
Trained on unnamed intent-classification data, the task family of this category's sources.
Trained on CLINC150's out-of-scope queries, the source of a third of this category, and calibrated on CLINC test queries; shown for reference, not ranked.
Trained on Banking77, the dataset behind this category (train split), and its checkpoint was picked on a set holding 27 of these 1,000 test messages; shown for reference, not ranked.
Trained on Banking77, the dataset behind this category (train split), and calibrated on Banking77 test messages; shown for reference, not ranked.
Trained on unnamed banking-intent data; its authors list Banking77 itself as held out.
Option-definition bug: all 13 routing-20 items whose gold intent is get_physical_card are PIN questions, which contradicts our written definition of that option ('the customer asks about getting a physical card'). Jev, which reads only the definition, gets 1 of 13; Kev-0.8B, trained on Banking77, gets 13 of 13. beneficiary_not_allowed has the same problem. A memorised dataset convention shows up as accuracy. With the affected items and the adjudicated label errors removed, routing-20's first tier is a three-way tie of Kev-0.8B, Jev and Kev-0.6B, and Jev joins routing-77's first tier (data-quality audit, 2026-09-22). The audit itself warns that dropping flagged items favours Jev by construction, because the flagged items are the ones Jev confidently missed, so read that re-scored tier as a bound, not a result.
27 routing-20 and 18 routing-77 messages also appear verbatim in Kev's own locked test set.
Trained on Banking77, the dataset behind this category (train split), and its checkpoint was picked on a set holding 16 of these 500 test messages; shown for reference, not ranked.
Trained on support-ticket queue routing (Tobi-Bueck/customer-support-tickets), a closely related task.
Trained on MS MARCO query-passage relevance, a closely related task.
Fine-tuned on this dataset's training split (same generator, workflows and questions); shown for reference, not ranked.
Trained on unnamed toxicity, rubric-rating and fact-checking data whose descriptions match four of this suite's eight configs; possibly the same datasets.
Trained on MNLI, a task close to this suite's fact-checking config (150 of 1,200 items).
Its base model's lineage includes toxicity, hate-speech and NLI data, close to three of this suite's eight configs.
Trained on template-generated policy-rule cases like this set's policy items (no item overlap found).
openJev Verdict 1.4's inference settings were chosen by measuring on these same 231 public items; shown for reference, not ranked.