OpenAIProprietary model

OpenAI Decisions API benchmarks

All models
  • General rank
    #6of 28Leads 3 of 6 (50%)non-coding tasks
  • Speed#1 / 4
    226 msMedian latency, p50Hosted API only
  • Cost#2 / 7
    $1.23Per 1,000 decisionsHosted price
  • LLM-as-a-judge#1 / 21 tied
    77.0%Best task · Accuracyvs Jev-class models

Comparison summary

Strongest
Judging and evals
Leads 3 of 6 (50%)

A typed decision model: it answers a typed question, such as yes or no or one of a few options, with a probability for each answer and no text. Strong at judging.

Specifications

Lab
OpenAI
Weights
Proprietary
Parameters
Not disclosed
Licence
Proprietary

Overall standing

Share of non-coding tasks led · #6 of 28

Judging and evals

Set

Share of tasks led · #9 of 26

  1. Claude Opus 5.5Leads 3 of 3
  2. ClefLeads 3 of 3
  3. Decision 2.0 Vega 27BLeads 3 of 3
  4. GLiDELeads 3 of 3
  5. Jev-OmniLeads 3 of 3
  6. Kev-27BLeads 3 of 3
  7. WinnowLeads 3 of 3
  8. OpenAI Decisions API
    Leads 3 of 6
See the 3 tasks

LLM-as-a-judge

Tier 3 of 4#6 / 7

Accuracy, % · Higher is better

  1. GPT-6.1 Sol92.6, interval 89.1 to 95.8, tied for the lead
  2. Claude Opus 5.589.1, interval 83.9 to 93.9, tied for the lead
  3. Claude Sonnet 5.583.5, interval 78.4 to 88.5, tier 2
  4. GPT-6 Luna82.6, interval 77.5 to 87.8, tier 2
  5. Gemini 3.8 Flash80.9, interval 76.2 to 85.5, tier 2
  6. OpenAI Decisions API
    71.7, interval 66.1 to 77.2, tier 3
  7. Jev61.3, interval 54.8 to 67.9, tier 4
Tested against Jev and 5 frontier LLMsAlso tested against Jev-class models: #1 of 21 tied

Scenario judge

Tied for the lead#1 / 7

Accuracy, % · Higher is better

  1. Claude Opus 5.586.8, interval 80.2 to 92.3, tied for the lead
  2. Gemini 3.8 Flash83.3, interval 76.8 to 89.1, tied for the lead
  3. OpenAI Decisions API
    83.3, interval 76.1 to 89.9, tied for the lead
  4. Claude Sonnet 5.581.3, interval 74.1 to 87.4, tier 2
  5. GPT-6.1 Sol81.3, interval 74.1 to 87.3, tier 2
  6. GPT-6 Luna79.9, interval 72.8 to 86.2, tier 2
  7. Jev66.7, interval 57.3 to 75.3, tier 3
Tested against Jev and 5 frontier LLMsAlso tested against Jev-class models: #9 of 21 tied

Search

Tier 3 of 4#5 / 7

Accuracy, % · Higher is better

  1. Claude Opus 5.591.5, interval 87.6 to 94.8, tied for the lead
  2. GPT-6.1 Sol89.6, interval 85.6 to 93.1, tied for the lead
  3. Claude Sonnet 5.589.1, interval 85.8 to 92.1, tied for the lead
  4. Gemini 3.8 Flash85.0, interval 81.3 to 88.4, tier 2
  5. GPT-6 Luna75.2, interval 70.7 to 79.6, tier 3
  6. OpenAI Decisions API
    70.4, interval 66.0 to 74.8, tier 3
  7. Jev63.3, interval 58.2 to 68.1, tier 4
Tested against Jev and 5 frontier LLMsAlso tested against Jev-class models: #1 of 21 tied

Speed

Median latency, p50 · Hosted API only · lower is better

  1. OpenAI Decisions API
    226 ms
  2. Jev408 ms
  3. GPT-6 Luna3,085 ms
  4. Gemini 3.8 Flash5,493 ms

Cost

Per 1,000 decisions, hosted price · lower is better

  1. Jev$0.43
  2. OpenAI Decisions API
    $1.23
  3. GPT-6 Luna$1.42
  4. Gemini 3.8 Flash$15
  5. GPT-6.1 Sol~$28
  6. Claude Sonnet 5.5~$30
  7. Claude Opus 5.5~$69
How we count

A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.

OpenAI Decisions APIOther modelsWhisker: 95% interval