vLLM Semantic RouterOpen weights model2B parameters

Decision 2.0 Sol 2B benchmarks

All models
  • General rank
    Too few tasks3 non-coding tasks ranked5 needed for a rank
  • Speed#3 / 8
    269 msMedian latency, p50L4 only
  • Cost
    Self-hostedRan on L4No hosted price
  • LLM-as-a-judge#15 / 20 tied
    51.0%Best task · Accuracyvs Jev-class models

Comparison summary

Strongest
Judging and evals
Leads 0 of 3 (0%)

Strongest at judging. It is ranked on 3 non-coding tasks, too few for a general rank (5 needed).

Specifications

Lab
vLLM Semantic Router
Weights
Open weights
Parameters
2B
Licence
Apache-2.0

Judging and evals

Share of tasks led · #20 of 20

  1. ClefLeads 3 of 3
  2. Decision 2.0 Vega 27BLeads 3 of 3
  3. GLiDELeads 3 of 3
  4. Jev-OmniLeads 3 of 3
  5. Kev-27BLeads 3 of 3
  6. WinnowLeads 3 of 3
  7. JevLeads 3 of 6
  8. Decision 2.0 Sol 2B
    Leads 0 of 3
See the 3 tasks

LLM-as-a-judge

Tier 3 of 3#15 / 20

Accuracy, % · Higher is better

  1. Decision 2.0 Vega 27B75.5, interval 69.8 to 81.3, tied for the lead
  2. Clef73.0, interval 66.2 to 79.7, tied for the lead
  3. Jev72.5, interval 66.7 to 78.3, tied for the lead
  4. GLiDE72.0, interval 65.7 to 78.1, tied for the lead
  5. Jev-Omni72.0, interval 65.5 to 78.5, tied for the lead
  6. Kev-27B72.0, interval 65.6 to 78.4, tied for the lead
  7. DiffusionGemma-Jev70.5, interval 63.8 to 77.4, tied for the lead
  8. Decision 2.0 Sol 2B
    51.0, interval 44.5 to 57.5, tier 3

Scenario judge

Tier 5 of 5#18 / 20

Accuracy, % · Higher is better

  1. Clef83.0, interval 77.4 to 88.4, tied for the lead
  2. Decision 2.0 Vega 27B81.0, interval 75.5 to 86.1, tied for the lead
  3. GLiDE79.5, interval 74.4 to 84.7, tied for the lead
  4. Winnow79.0, interval 73.1 to 85.0, tied for the lead
  5. Kev-27B78.0, interval 71.6 to 84.0, tied for the lead
  6. Jev76.5, interval 70.4 to 82.1, tied for the lead
  7. Hopper 12B75.5, interval 69.5 to 81.5, tied for the lead
  8. Decision 2.0 Sol 2B
    48.0, interval 41.7 to 54.5, tier 5

Search

Tier 4 of 4#19 / 20

Accuracy, % · Higher is better

  1. Jev-Omni72.5, interval 65.7 to 78.7, tied for the lead
  2. GLiDE71.5, interval 64.4 to 77.8, tied for the lead
  3. Decision 2.0 Vega 27B71.0, interval 64.5 to 77.0, tied for the lead
  4. Winnow69.5, interval 63.2 to 75.6, tied for the lead
  5. Jev69.0, interval 62.3 to 75.5, tied for the lead
  6. Clef67.5, interval 61.3 to 73.1, tied for the lead
  7. Kev-27B66.0, interval 59.7 to 71.8, tied for the lead
  8. Decision 2.0 Sol 2B
    46.5, interval 39.8 to 53.6, tier 4

Speed

Median latency, p50 · L4 only · lower is better

  1. Decision 2.0 Kai 0.6B114 ms
  2. Decision 2.0 Eos 0.8B250 ms
  3. Decision 2.0 Sol 2B
    269 ms
  4. decider-4b v2409 ms
  5. Lev 4B620 ms
  6. JevK5665 ms
  7. Decision 2.0 Nox 4B691 ms
  8. Winnow1,356 ms
How we count

A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.

Decision 2.0 Sol 2BOther modelsWhisker: 95% interval