Not recordedOpen weights model0.49B parameters

GDPR Anonymization 0.5B benchmarks

All models
  • Cost
    Self-hostedSmall GPU, ~1.5 GBNo hosted price

Comparison summary

Tested on 1 task, too few to compare it in general.

Specifications

Lab
Not recorded
Weights
Open weights
Parameters
0.49B
Licence
Apache-2.0
Type
small LLM

Safety and guardrails

Tested on 1 task here, none ranked.

See the task

LangWatch policy

Not ranked

PII left in, % · Lower is better

  1. Gemma 4 31B6.9, interval 2.8 to 11.3, ahead alone
  2. Ministral 3 14B10.2, interval 5.8 to 15.4, tier 2
  3. Gemma 4 12B16.5, interval 10.3 to 23.2, tier 3
  4. Qwen3.8 27B16.7, interval 9.8 to 25.0, tier 3
  5. Gemma 4 26B A4B20.5, interval 13.3 to 28.8, tier 4
  6. Qwen3.5 9B21.5, interval 13.4 to 30.6, tier 4
  7. Gemma 3 4B IT22.4, interval 14.3 to 32.1, tier 4
  8. GDPR Anonymization 0.5B
    64.2, interval 55.3 to 72.1, not ranked

Cost

Self-hosted · GPU memory needed, among small LLMs · lower is better

  1. Qwen3 0.6B~1 GB
  2. Qwen3.5 0.8B~1.1 GB
  3. LFM2 1.2B~1.3 GB
  4. LFM2.5 1.2B Instruct~1.3 GB
  5. Gemma 3 1B IT~1.4 GB
  6. Llama 3.2 1B Instruct~1.4 GB
  7. GDPR Anonymization 0.5B
    ~1.5 GB
  8. NuExtract 2.0 2B~1.6 GB
How we count

A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.

GDPR Anonymization 0.5BOther modelsWhisker: 95% interval