QwenOpen weights model2B parameters

Qwen3.5 2B benchmarks

All models
  • Cost
    Self-hostedSmall GPU, ~1.9 GBNo hosted price

Comparison summary

Tested on 1 task, too few to compare it in general.

Specifications

Lab
Qwen
Weights
Open weights
Parameters
2B
Licence
Apache-2.0
Type
small LLM

Safety and guardrails

Leads 0 of 1, too few tasks to compare here.

See the task

LangWatch policy

Tier 13 of 13#39 / 42

PII left in, % · Lower is better

  1. Gemma 4 31B6.9, interval 2.8 to 11.3, ahead alone
  2. Ministral 3 14B10.2, interval 5.8 to 15.4, tier 2
  3. Gemma 4 12B16.5, interval 10.3 to 23.2, tier 3
  4. Qwen3.8 27B16.7, interval 9.8 to 25.0, tier 3
  5. Gemma 4 26B A4B20.5, interval 13.3 to 28.8, tier 4
  6. Qwen3.5 9B21.5, interval 13.4 to 30.6, tier 4
  7. Gemma 3 4B IT22.4, interval 14.3 to 32.1, tier 4
  8. Qwen3.5 2B
    99.8, interval 99.5 to 100.0, tier 13

Cost

Self-hosted · GPU memory needed, among small LLMs · lower is better

How we count

A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.

Qwen3.5 2BOther modelsWhisker: 95% interval