GoogleOpen weights model4B parameters

Gemma 3 4B IT benchmarks

All models
  • General rank
    #27of 27Leads 0 of 5 (0%)non-coding tasks
  • Cost
    Self-hostedSmall GPU, ~3.1 GBNo hosted price
  • LangWatch policy#5 / 42 tied
    22.4%Best task · PII left invs PII detectors

Comparison summary

Strongest
Safety and guardrails
Leads 0 of 5 (0%)

Strongest at safety. It leads 0 of 5 (0%) non-coding tasks, #27 of 27 models by that count.

Specifications

Lab
Google
Weights
Open weights
Parameters
4B
Licence
Gemma Terms of Use
Type
small LLM

Overall standing

Share of non-coding tasks led · #27 of 27

Safety and guardrails

Share of tasks led · #14 of 46

  1. Gemma 4 31BLeads 3 of 4
  2. Perplexity PII TracerLeads 2 of 4
  3. Amazon ComprehendLeads 1 of 4
  4. Qwen3.8 27BLeads 1 of 5
  5. Gemma 4 12BLeads 0 of 4
  6. Ministral 3 14BLeads 0 of 5
  7. Gemma 4 26B A4BLeads 0 of 4
  8. Gemma 3 4B IT
    Leads 0 of 5
See the 5 tasks

Traces

Tier 6 of 19#13 / 46

PII left in, % · Lower is better

  1. Amazon Comprehend10.6, interval 9.4 to 11.9, ahead alone
  2. GLiNER Multi PII v115.3, interval 13.2 to 17.3, tier 2
  3. NVIDIA GLiNER PII19.7, interval 17.4 to 22.0, tier 3
  4. Gemma 4 26B A4B19.8, interval 15.8 to 24.2, tier 3
  5. Qwen3.8 27B20.3, interval 16.5 to 24.3, tier 3
  6. GLiNER2 Large20.4, interval 18.0 to 22.8, tier 3
  7. Gemma 4 12B23.3, interval 19.0 to 27.6, tier 3
  8. Gemma 3 4B IT
    43.9, interval 39.3 to 48.4, tier 6

Business

Tier 4 of 18#9 / 46

PII left in, % · Lower is better

  1. Perplexity PII Tracer2.1, interval 1.2 to 3.0, tied for the lead
  2. Qwen3.8 27B2.4, interval 1.7 to 3.3, tied for the lead
  3. Gemma 4 31B2.9, interval 1.9 to 4.0, tied for the lead
  4. GLiNER Multi PII v14.1, interval 3.1 to 5.3, tier 2
  5. Ministral 3 14B5.4, interval 4.0 to 6.9, tier 2
  6. Gemma 4 12B5.8, interval 4.4 to 7.3, tier 3
  7. Amazon Comprehend5.9, interval 4.6 to 7.2, tier 3
  8. Gemma 3 4B IT
    10.7, interval 9.1 to 12.5, tier 4

Multilingual

Tier 12 of 25#19 / 43

PII left in, % · Lower is better

  1. Perplexity PII Tracer2.9, interval 2.5 to 3.4, ahead alone
  2. Qwen3.8 27B4.2, interval 3.7 to 4.8, tier 2
  3. Gemma 4 31B4.5, interval 3.9 to 5.1, tier 2
  4. GLiNER Multi PII v15.7, interval 5.2 to 6.3, tier 3
  5. bardsai EU PII Multilang6.9, interval 6.3 to 7.5, tier 4
  6. Qwen3.6 35B A3B7.7, interval 7.0 to 8.4, tier 4
  7. Gemma 4 12B8.9, interval 8.1 to 9.7, tier 5
  8. Gemma 3 4B IT
    35.4, interval 33.9 to 36.7, tier 12

Look-alikes

Tier 5 of 7#6 / 9

Look-alikes redacted, % · Lower is better

  1. Gemma 4 31B3.5, interval 2.1 to 5.1, ahead alone
  2. Qwen3.8 27B10.3, interval 7.9 to 12.9, tier 2
  3. NuExtract 319.8, interval 15.9 to 23.6, tier 3
  4. Qwen3.5 9B20.5, interval 16.8 to 24.3, tier 3
  5. Ministral 3 14B30.5, interval 26.5 to 34.8, tier 4
  6. NVIDIA GLiNER PII42.0, interval 38.2 to 46.1, tier 5
  7. Gemma 3 4B IT
    45.0, interval 41.0 to 49.0, tier 5
  8. Perplexity PII Tracer67.6, interval 63.7 to 71.6, tier 6

LangWatch policy

Tier 4 of 13#5 / 42

PII left in, % · Lower is better

  1. Gemma 4 31B6.9, interval 2.8 to 11.3, ahead alone
  2. Ministral 3 14B10.2, interval 5.8 to 15.4, tier 2
  3. Gemma 4 12B16.5, interval 10.3 to 23.2, tier 3
  4. Qwen3.8 27B16.7, interval 9.8 to 25.0, tier 3
  5. Gemma 4 26B A4B20.5, interval 13.3 to 28.8, tier 4
  6. Qwen3.5 9B21.5, interval 13.4 to 30.6, tier 4
  7. Gemma 3 4B IT
    22.4, interval 14.3 to 32.1, tier 4
  8. Qwen3 4B Instruct 250729.6, interval 19.6 to 40.3, tier 5

Cost

Self-hosted · GPU memory needed, among small LLMs · lower is better

How we count

A model leads a task when it is in the task's leading tie tier. Each task counts inside its own benchmark, against that benchmark's models; no score is averaged. An area chart shows the models ranked on at least 2 of the area's tasks in a benchmark this model is in too. The overall standing counts every non-coding tasks and needs 5 for a rank. Latency is compared only on one hardware tier.

Gemma 3 4B ITOther modelsWhisker: 95% interval