Gemma 4 31B leaves the least PII in on the LangWatch policy test
58 PII detectors on 6 tests, scored by how much PII each one leaves in. Lower is better. The best 6 per test, against the best trivial baseline.
This benchmark was created using LangWatch. Sign up to create your own · Read the method
Showing 58 of 58 models. Filters hide rows; tiers and intervals stay as measured on every model.
LangWatch policy
PII left in, % · Lower is better
- Gemma 4 31B6.9, interval 2.8 to 11.3, ahead alone
- Ministral 3 14B10.2, interval 5.8 to 15.4, tier 2
- Gemma 4 12B16.5, interval 10.3 to 23.2, tier 3
- Qwen3.8 27B16.7, interval 9.8 to 25.0, tier 3
- Gemma 4 26B A4B20.5, interval 13.3 to 28.8, tier 4
- Qwen3.5 9B21.5, interval 13.4 to 30.6, tier 4
- Baseline: regex47.4, a trivial baseline
Ahead aloneTied for the leadThe rest95% intervalBest trivial baseline
What can I run
Best on LangWatch policy for each hardware class: pii left in, lower is better
Small GPU (laptop, up to ~8 GB)
32 models: Qwen3 4B Instruct 2507, GLiNER2 Large, Ministral 3 3B Instruct, GLiNER Multi PII v1, NuExtract 2.0 2B, OpenMed Privacy Filter Multilingual, ...
Mid GPU (~16 to 24 GB)
5 models: Qwen3.5 9B, NuExtract 2.0 8B, Granite 4.2 8B, NuExtract 3
Large GPU (40 GB and up)
5 models: Gemma 4 12B, Qwen3.8 27B, Gemma 4 26B A4B, Qwen3.6 35B A3B
Need not stated
16 models: DataFog, Presidio (LangWatch config), Presidio (all recognizers), Presidio (default), spaCy en_core_web_lg, CredSweeper, ...
GPU needs are estimates from weight size plus a serving margin, at the precision each model was run in.
How we measured
Scores carry a 95% percentile bootstrap interval over items. The tie test is a paired McNemar test on accuracy, or a paired bootstrap of the metric difference, on the items every model answered.
Tie tiers
Models in one tier can't be separated by the tie test (tie tiers as the benchmark report computed them (reports/SCORES.md)). Tiers are computed in the export over every model; filters on this page never change them.
Contamination
A model trained on a test set is grayed and never ranked. Overlap with a training set or a home advantage (the model comes from the dataset authors) is footnoted; a model that does not disclose its training data is marked unknown.
Datasets
- Traces (950 test items)
- Business (990 test items)
- Code and logs (1,000 test items)
- Multilingual (1,769 test items)
- Look-alikes (500 test items)
- LangWatch policy (787 test items)
- * Overlaps with a dataset the model was trained on.
- The model comes from the authors of one of the test sets.
- Amazon Comprehend: Scored with AWS Comprehend DetectPiiEntities in eu-central-1, LanguageCode es for Spanish records and en for all others. The confidence threshold was chosen on the dev split per category with the harness's select-threshold and frozen before the test run; test calls ran at concurrency 1, so latency is that of the hosted API tier, and the price is USD 0.0001 per 100-character unit (3-unit minimum). The frozen datasets and the harness config (benchmarks/pii/benchmark.yaml) define the rest of the replication. On hard-negatives no dev threshold passed the entity-leak gate, so the frozen threshold redacts nothing and that cell is gate-failed, not a score.
- Not scored: nemotron-3-nano-30b-a3b (all span categories): vLLM reasoning-plugin import error on box A; the server exited twice, no run exists.
- Not scored: mistral-small-4-119b (all span categories): TP2 flashinfer JIT build failed on box A, no run exists.
- Not scored: shisa-de-1 (doc-level only): Gemma-4 base that vLLM could not serve; no run.
- Not scored: kev-9b (doc-level only): CUDA out of memory on the 24 GB L4 at model load (twice); no run.
- Not scored: kev-4b (doc-level only): needs 8-9 GB of weights, does not fit the 4 GB laptop GPU, not tried; would also be contaminated on business (trained on complaint narratives).
- Not scored: laya, laya-typed, verdict-1.4, kev-0.6b, kev-0.8b, simplejev-0.8b, semif-0.6b (doc-level only): laptop Jev-like runs exited rc=99 (loader/server failure, work/phase6/jev/SKIPPED.md); no scores.
- Not scored: gemma-4-31b (traces; doc-natural): threshold selection failed on traces dev and the job gave up; doc-natural has no run. Scored on the other five span categories.
- Not scored: lfm2.5-1.2b-instruct, llama-3.2-1b-instruct, qwen3.5-2b (business; code-logs; hard-negatives; multilingual; traces (see table)): invalid answers on >= 90% of dev items (60/60 or 120/120 for most), so no test run in those categories (scored, if at all, only where listed in the main table).
- Not scored: gemma-3-1b-it, lfm2-1.2b, llama-3.2-3b-instruct, ministral-3-3b-instruct-2512, nuextract-2.0-2b, qwen3-0.6b, qwen3.5-0.8b (300-item subsample (‡) in 5 span categories): plan amendment A8: not among the six best LLM rows on dev, so only a 300-item subsample of each test set was run; never ranked against full-set rows.
- Not scored: aws-comprehend (doc-twins; doc-natural): span-only service (DetectPiiEntities), not run on the document-level sets.
- Not scored: gpt-oss-120b, gpt-oss-20b (all span categories): the vLLM server could not load the tokenizer vocab file on box A: every dev and test request returned HTTP 500 or a connection reset (0/787 answered on langwatch-policy test). The 100% leak in results.json is a setup failure, not a model result; no score.
- Not scored: gdpr-anonymization-0.5b (five span categories other than langwatch-policy): dev completed everywhere (0 invalid), but only langwatch-policy was run on test; no ledger or event entry gives a reason for the other five, so treat them as not run (flagged for follow-up).
- Not scored: azure-language-pii, google-dlp, nightfall, lakera-guard, pangea-redact, gpt-6-luna, gpt-6-sol, claude-haiku-4-5, claude-sonnet-5, claude-opus-5-5, gemini-2.5-flash-lite, gemini-3.5-flash-lite, gemini-3.1-pro-preview, mistral-small-4, mistral-medium-3.5, mistral-large-3 (all): paid or signup-gated vendor APIs: registered for the record, never approved and never called ($0 spent on them).
Release pii-local-203f836324, benchmark commit 203f836, 3 input files, integrity none (parsed from markdown).
