Supersonic LabsOpen weights model0.144B parameters
Julia-1 benchmarks
- General rankToo few tasks3 non-coding tasks ranked5 needed for a rank
- Speed#1 / 11,623 msMedian latency, p50Server CPU only
- CostSelf-hostedRan on Server CPUNo hosted price
- Search#12 / 20 tied49.5%Best task · Accuracyvs Jev-class models
Comparison summary
- Strongest
- Judging and evals
- Leads 0 of 3 (0%)
Strongest at judging. It is ranked on 3 non-coding tasks, too few for a general rank (5 needed).
Specifications
- Lab
- Supersonic Labs
- Weights
- Open weights
- Parameters
- 0.144B
- Licence
- Apache-2.0
Judging and evals
Share of tasks led · #18 of 20
See the 3 tasks
LLM-as-a-judge
Tier 3 of 3#15 / 20Accuracy, % · Higher is better
- Decision 2.0 Vega 27B75.5, interval 69.8 to 81.3, tied for the lead
- Clef73.0, interval 66.2 to 79.7, tied for the lead
- Jev72.5, interval 66.7 to 78.3, tied for the lead
- GLiDE72.0, interval 65.7 to 78.1, tied for the lead
- Jev-Omni72.0, interval 65.5 to 78.5, tied for the lead
- Kev-27B72.0, interval 65.6 to 78.4, tied for the lead
- DiffusionGemma-Jev70.5, interval 63.8 to 77.4, tied for the lead
- Julia-153.0, interval 47.7 to 58.8, tier 3
Scenario judge
Tier 5 of 5#18 / 20Accuracy, % · Higher is better
- Clef83.0, interval 77.4 to 88.4, tied for the lead
- Decision 2.0 Vega 27B81.0, interval 75.5 to 86.1, tied for the lead
- GLiDE79.5, interval 74.4 to 84.7, tied for the lead
- Winnow79.0, interval 73.1 to 85.0, tied for the lead
- Kev-27B78.0, interval 71.6 to 84.0, tied for the lead
- Jev76.5, interval 70.4 to 82.1, tied for the lead
- Hopper 12B75.5, interval 69.5 to 81.5, tied for the lead
- Julia-147.5, interval 40.5 to 54.3, tier 5
Search
Tier 3 of 4#12 / 20Accuracy, % · Higher is better
- Jev-Omni72.5, interval 65.7 to 78.7, tied for the lead
- GLiDE71.5, interval 64.4 to 77.8, tied for the lead
- Decision 2.0 Vega 27B71.0, interval 64.5 to 77.0, tied for the lead
- Winnow69.5, interval 63.2 to 75.6, tied for the lead
- Jev69.0, interval 62.3 to 75.5, tied for the lead
- Clef67.5, interval 61.3 to 73.1, tied for the lead
- Kev-27B66.0, interval 59.7 to 71.8, tied for the lead
- Julia-149.5, interval 43.4 to 56.1, tier 3
How we count
A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.
Julia-1Other modelsWhisker: 95% interval
