akhilaaa3Open weights model12B parameters
Jev-Omni benchmarks
- General rankToo few tasks3 non-coding tasks ranked5 needed for a rank
- Speed#4 / 7248 msMedian latency, p50H100 only
- CostSelf-hostedRan on H100No hosted price
- LLM-as-a-judge#1 / 20 tied72.0%Best task · Accuracyvs Jev-class models
Comparison summary
- Strongest
- Judging and evals
- Leads 3 of 3 (100%)
Strong at judging. It is ranked on 3 non-coding tasks, too few for a general rank (5 needed).
Specifications
- Lab
- akhilaaa3
- Weights
- Open weights
- Parameters
- 12B
- Licence
- Apache-2.0
Judging and evals
Share of tasks led · #4 of 20
- ClefLeads 3 of 3
- Decision 2.0 Vega 27BLeads 3 of 3
- GLiDELeads 3 of 3
- Jev-OmniLeads 3 of 3
- Kev-27BLeads 3 of 3
- WinnowLeads 3 of 3
- JevLeads 3 of 6
- DiffusionGemma-JevLeads 1 of 3
See the 3 tasks
LLM-as-a-judge
Tied for the lead#1 / 20Accuracy, % · Higher is better
- Decision 2.0 Vega 27B75.5, interval 69.8 to 81.3, tied for the lead
- Clef73.0, interval 66.2 to 79.7, tied for the lead
- Jev72.5, interval 66.7 to 78.3, tied for the lead
- GLiDE72.0, interval 65.7 to 78.1, tied for the lead
- Jev-Omni72.0, interval 65.5 to 78.5, tied for the lead
- Kev-27B72.0, interval 65.6 to 78.4, tied for the lead
- DiffusionGemma-Jev70.5, interval 63.8 to 77.4, tied for the lead
- Winnow70.5, interval 64.0 to 76.3, tied for the lead
Scenario judge
Tied for the lead#1 / 20Accuracy, % · Higher is better
- Clef83.0, interval 77.4 to 88.4, tied for the lead
- Decision 2.0 Vega 27B81.0, interval 75.5 to 86.1, tied for the lead
- GLiDE79.5, interval 74.4 to 84.7, tied for the lead
- Winnow79.0, interval 73.1 to 85.0, tied for the lead
- Kev-27B78.0, interval 71.6 to 84.0, tied for the lead
- Jev76.5, interval 70.4 to 82.1, tied for the lead
- Hopper 12B75.5, interval 69.5 to 81.5, tied for the lead
- Jev-Omni75.5, interval 69.3 to 81.7, tied for the lead
Search
Tied for the lead#1 / 20Accuracy, % · Higher is better
- Jev-Omni72.5, interval 65.7 to 78.7, tied for the lead
- GLiDE71.5, interval 64.4 to 77.8, tied for the lead
- Decision 2.0 Vega 27B71.0, interval 64.5 to 77.0, tied for the lead
- Winnow69.5, interval 63.2 to 75.6, tied for the lead
- Jev69.0, interval 62.3 to 75.5, tied for the lead
- Clef67.5, interval 61.3 to 73.1, tied for the lead
- Kev-27B66.0, interval 59.7 to 71.8, tied for the lead
- DiffusionGemma-Jev63.0, interval 56.3 to 69.5, tier 2
Speed
Median latency, p50 · H100 only · lower is better
- Cygnet162 ms
- DiffusionGemma-Jev188 ms
- Decision 2.0 Lux 9B213 ms
- Jev-Omni248 ms
- Kev-27B318 ms
- Hopper 12B397 ms
- Decision 2.0 Vega 27B724 ms
How we count
A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.
Jev-OmniOther modelsWhisker: 95% interval
