iiiorg / token-classification updated 10 months ago

piiranha-v1-detect-personal-information

Piiranha (cc-by-nc-nd-4.0 license) is trained to detect 17 types of Personally Identifiable Information (PII) across six languages. It successfully catches 98.27% of PII tokens, with an overall classification accuracy of 99.44%. Piiranha is especially accurate at detecting passwords, emails (100%), phone numbers, and usernames.

Params
280 M
Context
512
Downloads 30d
236 K
Likes
249
Commercial use: not allowed · cc-by-nc-nd-4.0 Not gated SAFETENSORS 6 languages View on Hugging Face ↗

Download history

daily snapshots · 24 days
309 K236 K
Aug 28Sep 5Sep 13Sep 20

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors f32 1.1 GB 1.8 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
DebertaV2ForTokenClassification
Parameters
280 M
Tensor type
F32
Context length
512
Vocabulary
251,000
Layers / heads
12 / 12
Licence
cc-by-nc-nd-4.0
First seen on the Hub
2024-09-12
Training datasets
ai4privacy/pii-masking-400k
Added to our catalog
2026-08-28
Compare with any token-classification model