Models

General ranking

Models by the share of their non-coding tasks where they are in the leading tie tier, with the benchmarks counted, median latency and cost per 1,000 decisions.
#ModelTasks led
1Claude Opus 5.52,194 ms · $3.44Frontier LLMs 8/10LLM judges 3/311 of 13 (85%)
2Jev~65 ms · $0.029Frontier LLMs 2/10Open models to 28B 11/14Tiny open models 14/14LLM judges 0/3Jev-class models 3/330 of 44 (68%)
3AutoJev-27B99 ms · Self-hostedOpen models to 28B 8/128 of 12 (67%)
4Eikos-27B82 ms · Self-hostedOpen models to 28B 6/106 of 10 (60%)
5GPT-6.1 SolNot measured · ~$1.02Frontier LLMs 5/10LLM judges 2/37 of 13 (54%)
6Claude Sonnet 5.51,646 ms · $1.52Frontier LLMs 5/10LLM judges 1/36 of 13 (46%)
7Shisa DE-146 ms · Self-hostedOpen models to 28B 5/125 of 12 (42%)
8Gemini 3.8 Flash1,936 ms · $0.57Frontier LLMs 4/10LLM judges 1/35 of 13 (38%)
9GPT-5.6 Terra1,565 ms · $1.03Frontier LLMs 3/103 of 10 (30%)
10Qwen3.8 27BNot measured · Self-hostedPII detection 1/51 of 5 (20%)
11GPT-6 Luna1,124 ms · $0.049Frontier LLMs 2/102 of 10 (20%)
12Claude Haiku 4.51,306 ms · $0.69Frontier LLMs 1/101 of 10 (10%)
13SemIf (Qwen3.5-4B)108 ms · Self-hostedOpen models to 28B 1/111 of 11 (9%)
14Kev-9B43 ms · Self-hostedOpen models to 28B 0/80 of 8 (0%)
15Eikos-4B95 ms · Self-hostedOpen models to 28B 0/90 of 9 (0%)
16Ministral 3 14BNot measured · Self-hostedPII detection 0/50 of 5 (0%)
17Kev-4B85 ms · Self-hostedOpen models to 28B 0/120 of 12 (0%)
18Kev-0.8B85 ms · Self-hostedOpen models to 28B 0/12Tiny open models 0/120 of 24 (0%)
19Laya-typed53 ms · Self-hostedOpen models to 28B 0/13Tiny open models 0/130 of 26 (0%)
20Kev-0.6B61 ms · Self-hostedOpen models to 28B 0/12Tiny open models 0/120 of 24 (0%)
21Laya50 ms · Self-hostedOpen models to 28B 0/13Tiny open models 0/130 of 26 (0%)
22SimpleJev (Qwen3.5-0.8B)196 ms · Self-hostedOpen models to 28B 0/7Tiny open models 0/70 of 14 (0%)
23Qwen3.5 9BNot measured · Self-hostedPII detection 0/50 of 5 (0%)
24GLiNER2.5-Decide61 ms · Self-hostedOpen models to 28B 0/120 of 12 (0%)
25SemIf (Qwen3-0.6B)60 ms · Self-hostedOpen models to 28B 0/7Tiny open models 0/70 of 14 (0%)
26openJev Verdict 1.445 ms · Self-hostedOpen models to 28B 0/5Tiny open models 0/50 of 10 (0%)
27Gemma 3 4B ITNot measured · Self-hostedPII detection 0/50 of 5 (0%)
Too few tasks to rank (under 5)
Gemma 4 12BNot measured · Self-hostedPII detection 0/40 of 4 (0%)
Gemma 4 26B A4BNot measured · Self-hostedPII detection 0/40 of 4 (0%)
Gemma 4 31BNot measured · Self-hostedPII detection 3/43 of 4 (75%)
Granite 4.2 8BNot measured · Self-hostedPII detection 0/40 of 4 (0%)
Qwen3.6 35B A3BNot measured · Self-hostedPII detection 0/40 of 4 (0%)
Phi-4 Mini InstructNot measured · Self-hostedPII detection 0/30 of 3 (0%)
Llama 3.2 3B InstructNot measured · Self-hostedNo ranked task

Share of tasks in the leading tie tier, not an averaged score. Each task counts inside its own benchmark, against that benchmark's models. Ties go to the lower average rank; under 5 ranked tasks, no rank. Coding tasks (commit types, code and logs) are left out. Latency is compared only within one hardware tier; ~ marks an estimate.

All models