CloudflareOpen weights model27.4B parameters

Clef benchmarks

All models
  • General rank
    Too few tasks3 non-coding tasks ranked5 needed for a rank
  • Speed#4 / 4
    887 msMedian latency, p50Hosted API only
  • Cost#4 / 4
    $0.54Per 1,000 decisionsHosted price
  • LLM-as-a-judge#1 / 20 tied
    73.0%Best task · Accuracyvs Jev-class models

Comparison summary

Strongest
Judging and evals
Leads 3 of 3 (100%)

Strong at judging. It is ranked on 3 non-coding tasks, too few for a general rank (5 needed).

Specifications

Lab
Cloudflare
Weights
Open weights
Parameters
27.4B
Licence
Apache-2.0

Judging and evals

Share of tasks led · #1 of 20

  1. Clef
    Leads 3 of 3
  2. Decision 2.0 Vega 27BLeads 3 of 3
  3. GLiDELeads 3 of 3
  4. Jev-OmniLeads 3 of 3
  5. Kev-27BLeads 3 of 3
  6. WinnowLeads 3 of 3
  7. JevLeads 3 of 6
  8. DiffusionGemma-JevLeads 1 of 3
See the 3 tasks

LLM-as-a-judge

Tied for the lead#1 / 20

Accuracy, % · Higher is better

  1. Decision 2.0 Vega 27B75.5, interval 69.8 to 81.3, tied for the lead
  2. Clef
    73.0, interval 66.2 to 79.7, tied for the lead
  3. Jev72.5, interval 66.7 to 78.3, tied for the lead
  4. GLiDE72.0, interval 65.7 to 78.1, tied for the lead
  5. Jev-Omni72.0, interval 65.5 to 78.5, tied for the lead
  6. Kev-27B72.0, interval 65.6 to 78.4, tied for the lead
  7. DiffusionGemma-Jev70.5, interval 63.8 to 77.4, tied for the lead
  8. Winnow70.5, interval 64.0 to 76.3, tied for the lead

Scenario judge

Tied for the lead#1 / 20

Accuracy, % · Higher is better

  1. Clef
    83.0, interval 77.4 to 88.4, tied for the lead
  2. Decision 2.0 Vega 27B81.0, interval 75.5 to 86.1, tied for the lead
  3. GLiDE79.5, interval 74.4 to 84.7, tied for the lead
  4. Winnow79.0, interval 73.1 to 85.0, tied for the lead
  5. Kev-27B78.0, interval 71.6 to 84.0, tied for the lead
  6. Jev76.5, interval 70.4 to 82.1, tied for the lead
  7. Hopper 12B75.5, interval 69.5 to 81.5, tied for the lead
  8. Jev-Omni75.5, interval 69.3 to 81.7, tied for the lead

Search

Tied for the lead#1 / 20

Accuracy, % · Higher is better

  1. Jev-Omni72.5, interval 65.7 to 78.7, tied for the lead
  2. GLiDE71.5, interval 64.4 to 77.8, tied for the lead
  3. Decision 2.0 Vega 27B71.0, interval 64.5 to 77.0, tied for the lead
  4. Winnow69.5, interval 63.2 to 75.6, tied for the lead
  5. Jev69.0, interval 62.3 to 75.5, tied for the lead
  6. Clef
    67.5, interval 61.3 to 73.1, tied for the lead
  7. Kev-27B66.0, interval 59.7 to 71.8, tied for the lead
  8. DiffusionGemma-Jev63.0, interval 56.3 to 69.5, tier 2

Speed

Median latency, p50 · Hosted API only · lower is better

  1. Jev258 ms
  2. Clef Flash548 ms
  3. GLiDE763 ms
  4. Clef
    887 ms

Cost

Per 1,000 decisions, hosted price · lower is better

  1. Jev$0.11
  2. Clef Flash$0.20
  3. GLiDE$0.34
  4. Clef
    $0.54
How we count

A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.

ClefOther modelsWhisker: 95% interval