distilbert_pii_ner_yalen — PII & NER Detector (9 classes · ONNX INT8)
Fine-tuned and quantized version of distilbert-base-multilingual-cased for PII and Named Entity Recognition detection.
Developed by Yalen AI as part of the Yalen Sentinel Pulse privacy protection platform.
Model format: ONNX INT8 quantized — 74% smaller than the original PyTorch model (139 MB vs 543 MB), with minimal accuracy loss.
Model Description
This model detects 9 classes of sensitive personal and financial information using BIO tagging. It supports multilingual input with strong performance on French, English, and Tunisian Arabic-Latin mixed text.
Detected Entities
| Entity | Description | Examples |
|---|---|---|
| PER | Person name | Ahmed Ben Salah, Marie Dupont |
| ORG | Organization | TechCorp Tunisie, Banque de France |
| LOC | Location | Tunis, Paris, avenue Habib Bourguiba |
| MISC | Miscellaneous named entity | Visa, MasterCard |
| IBAN | International Bank Account Number | TN59 1000 6035 1835 9848 3270, FR76 3000... |
| CARD | Credit / Debit card number | 4532 1234 5678 9012 |
| PHONE | Phone number (international formats) | +216 71 234 567, +33 1 23 45 67 89 |
| Email address | ahmed.bensalah@techcorp.tn | |
| DATE | Date (any format) | 15/03/2025, 14 mars 2025, 12/2028 |
Performance
Global Scores (epoch 4 / 14,944 steps)
| Precision | Recall | F1-Score | Accuracy |
|---|---|---|---|
| 89.39% | 91.47% | 90.42% | 98.87% |
Per-Class F1 (evaluation set)
| Entity | F1 |
|---|---|
| PER | 100% |
| ORG | 88.9% |
| LOC | 80.0% |
| IBAN | 100% |
| CARD | 100% |
| PHONE | 100% |
| 100% | |
| DATE | 66.7% |
Training Data
| Source | Volume | Classes |
|---|---|---|
| Jean-Baptiste/wikiner_fr | 120,682 sentences | PER, ORG, LOC, MISC |
| ai4privacy/pii-masking-200k | 82,545 examples | IBAN, CARD, PHONE, EMAIL, DATE |
| Faker (synthetic) | 40,000 examples (8,000 × 5 classes) | IBAN, CARD, PHONE, EMAIL, DATE |
| TOTAL | 239,099 train · 17,538 val | 9 classes |
Quick Start
With Optimum (recommended for ONNX)
from optimum.onnxruntime import ORTModelForTokenClassification
from transformers import AutoTokenizer, pipeline
model = ORTModelForTokenClassification.from_pretrained(
"yalen-ai/distilbert_pii_ner_yalen",
file_name="model_quantized.onnx"
)
tokenizer = AutoTokenizer.from_pretrained("yalen-ai/distilbert_pii_ner_yalen")
ner = pipeline("token-classification", model=model, tokenizer=tokenizer, aggregation_strategy="simple")
text = "Contact: ahmed.bensalah@techcorp.tn | Tel: +216 71 234 567 | IBAN: TN59 1000 6035 1835 9848 3270"
results = ner(text)
for entity in results:
print(f"[{entity['entity_group']}] '{entity['word']}' (score: {entity['score']:.3f})")
Output:
[EMAIL] 'ahmed.bensalah@techcorp.tn' (score: 0.998)
[PHONE] '+216 71 234 567' (score: 0.984)
[IBAN] 'TN59 1000 6035 1835 9848 3270' (score: 0.757)
With ONNX Runtime directly
import onnxruntime as ort
from transformers import AutoTokenizer
import numpy as np
import json
# Load tokenizer and label map
tokenizer = AutoTokenizer.from_pretrained("yalen-ai/distilbert_pii_ner_yalen")
with open("config.json") as f:
cfg = json.load(f)
id2label = cfg["id2label"]
# Load ONNX session
session = ort.InferenceSession("model_quantized.onnx", providers=["CPUExecutionProvider"])
def predict(text):
inputs = tokenizer(text, return_tensors="np", truncation=True, max_length=512)
outputs = session.run(None, dict(inputs))
logits = outputs[0][0]
token_ids = inputs["input_ids"][0]
tokens = tokenizer.convert_ids_to_tokens(token_ids)
labels = [id2label[str(np.argmax(l))] for l in logits]
return [(tok, lbl) for tok, lbl in zip(tokens, labels) if lbl != "O" and not tok.startswith("[")]
results = predict("Ahmed Ben Salah travaille chez TechCorp, IBAN: TN59 1000 6035 1835 9848 3270")
for token, label in results:
print(f" {token:<30} {label}")
Installation
# For Optimum (ONNX Runtime)
pip install optimum[onnxruntime] transformers
# For direct ONNX Runtime usage
pip install onnxruntime transformers
Model Architecture
| Parameter | Value |
|---|---|
| Base model | distilbert-base-multilingual-cased |
| Architecture | DistilBertForTokenClassification |
| Hidden size | 768 |
| Attention heads | 12 |
| Hidden layers | 6 |
| Max tokens | 512 |
| Vocab size | 119,547 |
| Labels | 19 (O + 9×BIO) |
| Format | ONNX INT8 (avx2 quantization) |
Model Size Comparison
| Stage | Format | Size |
|---|---|---|
| Fine-tuning | PyTorch FP32 | 543 MB |
| Export | ONNX FP32 | 539 MB |
| Quantization | ONNX INT8 avx2 | 139 MB |
Label Map
{
"0": "O",
"1": "B-PER", "2": "I-PER",
"3": "B-ORG", "4": "I-ORG",
"5": "B-LOC", "6": "I-LOC",
"7": "B-MISC", "8": "I-MISC",
"9": "B-IBAN", "10": "I-IBAN",
"11": "B-CARD", "12": "I-CARD",
"13": "B-PHONE", "14": "I-PHONE",
"15": "B-EMAIL", "16": "I-EMAIL",
"17": "B-DATE", "18": "I-DATE"
}
Intended Use
- Privacy compliance (GDPR, Tunisian Data Protection Law)
- Document redaction — anonymize sensitive documents before sharing
- Data loss prevention (DLP) — detect accidental PII leaks in logs or messages
- Financial document processing — extract IBAN/card numbers for validation
- Healthcare & insurance — detect names, dates and contact information
- Edge deployment — ONNX INT8 runs efficiently on CPU without GPU
Limitations
- Maximum input length: 512 tokens (long documents should be split by sentence or paragraph)
- DATE detection is the weakest class (F1 ~67%) — dates in full French text ("14 mars 2025") are harder to detect than numeric formats
- Card number detection works best with standard spacing (
XXXX XXXX XXXX XXXX) - MISC class is inherited from WikiNER and may catch general named entities beyond PII
About Yalen AI
Yalen Sentinel Pulse is an AI-powered platform for PII detection and data privacy protection, developed by the Yalen AI team. It combines regex patterns, ML models, and NER to provide comprehensive sensitive data identification.
- Platform: Yalen Sentinel Pulse
- Contact: hello@yalen.ai
License
MIT — free for commercial and research use.
- Downloads last month
- 22
Evaluation results
- F1 (test set)self-reported0.904