vLLM Semantic RouterOpen weights model0.6B parameters
Decision 2.0 Kai 0.6B benchmarks
- General rankToo few tasks3 non-coding tasks ranked5 needed for a rank
- Speed#1 / 8114 msMedian latency, p50L4 only
- CostSelf-hostedRan on L4No hosted price
- LLM-as-a-judge#15 / 20 tied50.0%Best task · Accuracyvs Jev-class models
Comparison summary
- Strongest
- Judging and evals
- Leads 0 of 3 (0%)
Strongest at judging. It is ranked on 3 non-coding tasks, too few for a general rank (5 needed).
Specifications
- Lab
- vLLM Semantic Router
- Weights
- Open weights
- Parameters
- 0.6B
- Licence
- Apache-2.0
Judging and evals
Share of tasks led · #19 of 20
See the 3 tasks
LLM-as-a-judge
Tier 3 of 3#15 / 20Accuracy, % · Higher is better
- Decision 2.0 Vega 27B75.5, interval 69.8 to 81.3, tied for the lead
- Clef73.0, interval 66.2 to 79.7, tied for the lead
- Jev72.5, interval 66.7 to 78.3, tied for the lead
- GLiDE72.0, interval 65.7 to 78.1, tied for the lead
- Jev-Omni72.0, interval 65.5 to 78.5, tied for the lead
- Kev-27B72.0, interval 65.6 to 78.4, tied for the lead
- DiffusionGemma-Jev70.5, interval 63.8 to 77.4, tied for the lead
- Decision 2.0 Kai 0.6B50.0, interval 43.1 to 57.1, tier 3
Scenario judge
Tier 4 of 5#15 / 20Accuracy, % · Higher is better
- Clef83.0, interval 77.4 to 88.4, tied for the lead
- Decision 2.0 Vega 27B81.0, interval 75.5 to 86.1, tied for the lead
- GLiDE79.5, interval 74.4 to 84.7, tied for the lead
- Winnow79.0, interval 73.1 to 85.0, tied for the lead
- Kev-27B78.0, interval 71.6 to 84.0, tied for the lead
- Jev76.5, interval 70.4 to 82.1, tied for the lead
- Hopper 12B75.5, interval 69.5 to 81.5, tied for the lead
- Decision 2.0 Kai 0.6B50.0, interval 42.9 to 57.2, tier 4
Search
Tier 4 of 4#19 / 20Accuracy, % · Higher is better
- Jev-Omni72.5, interval 65.7 to 78.7, tied for the lead
- GLiDE71.5, interval 64.4 to 77.8, tied for the lead
- Decision 2.0 Vega 27B71.0, interval 64.5 to 77.0, tied for the lead
- Winnow69.5, interval 63.2 to 75.6, tied for the lead
- Jev69.0, interval 62.3 to 75.5, tied for the lead
- Clef67.5, interval 61.3 to 73.1, tied for the lead
- Kev-27B66.0, interval 59.7 to 71.8, tied for the lead
- Decision 2.0 Kai 0.6B45.0, interval 38.5 to 51.6, tier 4
Speed
Median latency, p50 · L4 only · lower is better
- Decision 2.0 Kai 0.6B114 ms
- Decision 2.0 Eos 0.8B250 ms
- Decision 2.0 Sol 2B269 ms
- decider-4b v2409 ms
- Lev 4B620 ms
- JevK5665 ms
- Decision 2.0 Nox 4B691 ms
- Winnow1,356 ms
How we count
A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.
Decision 2.0 Kai 0.6BOther modelsWhisker: 95% interval
