GPT-6 Luna, OpenAI: 88.0%, interval 86.1 to 89.8, tier 4
Ahead aloneTied for the leadThe rest95% interval
One task at a time
Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.
Is a user message an attempt to override, bypass or extract the assistant's instructions? Items mix short injections (some in German), long "DAN-style" jailbreaks, and injections aimed at a specific system prompt, against ordinary requests including role-play that is legitimate.
Is a message harmful content a moderation filter should flag (sexual, hate, violence, harassment, self-harm, sexual content involving minors)? Real user prompts sent to chatbots plus a curated set with per-category human labels.
Does a business document (invoice, ticket, record, transaction) contain personal data? Negatives are real documents with the personal data replaced by neutral wording.
Does the answer (or summary, or dialogue reply) add or contradict facts the provided context does not support? The positive class is a hallucinated answer, so a higher score means the model catches more hallucinations at the same rate of false alarms, not that its answers are more faithful.
Is a user request outside what the assistant supports? 150 supported intents across 10 domains versus real out-of-scope requests.
Pick the right banking intent for a customer message from 20 candidates (gold plus 19 distractors).
The same messages with all 77 banking intents as options.
Real user requests from BFCL "live" with the tools an application exposed (2-13 tools each, with their descriptions and parameters); pick the tool that should be called, or "none of these". Half the items have a matching tool, half are requests no listed tool can serve.
Real consumer complaints filed with the US CFPB, routed to one of 11 current product queues (credit reporting, cards, mortgages, loans, money transfer, ...). Balanced across products and years 2015-2024.
Commits from Angular and Vite: from the subject (type prefix removed), body and changed files, pick the conventional-commit type (build, ci, docs, feat, fix, perf, refactor, test). Balanced across types.
Amazon ESCI: for a real shopping query and a product listing, decide Exact, Substitute, Complement or Irrelevant, using the dataset's own guideline definitions. Balanced across the four labels.
Prompt injection: Balanced accuracy, higher is better
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 86%, not 0.
x
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: balanced accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
x
Dot: accuracy. Line: its 95% interval. Higher is better. Orange: ahead by the tie test; green: tied for the lead. Axis starts at 00.0%, not 0.
Plot
Prompt injection: balanced accuracy (higher is better) vs cost
Claude Haiku 4.5: 92.7% [91.1, 94.3], $0.61, tier 2.
GPT-5.6 Terra: 88.2% [86.3, 90.0], $1.04, tier 4.
Claude Opus 5.5: 90.9% [89.2, 92.5], $3.85, tier 2.
Claude Sonnet 5.5: 88.3% [86.4, 90.2], $1.69, tier 4.
GPT-6.1 Sol: 91.3% [89.6, 92.9], ~$1.02, tier 2.
$0.01$0.1$1$10← Cost per 1,000 decisions, USD (log scale). Lower is betterHigher is better ↑
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Dashed line: the Pareto frontier (no other model is both better and cheaper). Hollow mark: an estimate (~).
Dashed line: the Pareto frontier (no other model is both better and faster). Hollow mark: an estimate (~). Latency is only compared within one hardware tier: Hosted API. Not plotted (no latency on Hosted API measured): GPT-6.1 Sol.
Every model, every task
Every visible model on every task, sorted by Prompt injection. Value with half its interval; bands are tie tiers.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Scores in percent, higher is better, ± half the 95% interval. T1 is the leading tie tier: models in one tier can't be told apart by the tie test. Orange: the one model ahead; green: tied for the lead. No overall column: tasks use different metrics, so we don't average them.
Training data not disclosed, so contamination can't be ruled out.
~ marks an approximate latency or an estimated cost.
How we measured
Scores carry a 95% percentile bootstrap interval over items (2,000 resamples on the use cases). The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the metric difference, on the items every model answered.
Tie tiers
Models in one tier can't be separated by the tie test (paired bootstrap of the Balanced accuracy difference on shared items, 95% CI includes 0; paired bootstrap of the Balanced acc. difference on shared items, 95% CI includes 0; paired McNemar on per-item correctness at the dev threshold, p >= 0.05). Tiers are computed in the export over every model; filters on this page never change them.
Contamination
Grayed: trained on the test items or on the same dataset for at least a third of the category. Footnoted: trained on a related task. Unknown: training data not disclosed. Grayed entries are shown for reference and never ranked as the winner.
Latency and hardware
Hosted API: the laptop (Denmark) through the LangWatch AI Gateway, concurrency 6; latency is the time of each call during the scoring run and includes the network and the gateway hop. the laptop (Denmark), concurrency 1; approximate latency: p50 is the client time minus the network round trip measured in the same window (warm keep-alive requests to the API host); p95 is the API's server-reported time.
Latencies are only compared within one hardware tier.
Moderation (667 test items): mmathys/openai-moderation-api-evaluation (MIT), nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (CC-BY-4.0)
PII (1,000 test items): gretelai/gretel-pii-masking-en-v1 (Apache-2.0), beki/privy (MIT), nvidia/Nemotron-PII (CC-BY-4.0)
RAG faithfulness (666 test items): pminervini/HaluEval (Apache-2.0)
Off-topic (1,000 test items): clinc/clinc_oos (CC-BY-3.0), mteb/amazon_massive_intent (CC-BY-4.0), benayas/snips (CC0-1.0)
Routing, 20 intents (1,000 test items): legacy-datasets/banking77 (CC-BY-4.0)
Routing, 77 intents (500 test items): legacy-datasets/banking77 (CC-BY-4.0)
Tool routing (1,000 test items): gorilla-llm/Berkeley-Function-Calling-Leaderboard (Apache-2.0)
Complaint routing (1,000 test items): BEE-spoke-data/consumer-finance-complaints (CC0-1.0)
Commit type (1,000 test items): github.com/angular/angular (MIT), github.com/vitejs/vite (MIT)
Search relevance (1,000 test items): tasksource/esci (Apache-2.0)
Jev's latency on RAG faithfulness comes from its latency run, whose 300 items are all from HaluEval dialogue, the non-commercial source this task is scored without: only the timing is used, no item or score.