piiguard / README.md
bogdanraduta's picture
Add OpenNER model, card, NOTICE (Apache-2.0, FlowX.AI)
018e7f0 verified
|
Raw
History Blame Contribute Delete
1.8 kB
---
license: apache-2.0
library_name: transformers
pipeline_tag: token-classification
base_model: FacebookAI/xlm-roberta-base
language:
- en
- ro
- bg
- hu
- sl
- hr
- de
- it
- fr
tags:
- ner
- on-device
- privacy
- flowx
- openner
- cross
- de-identification
- token-classification
metrics:
- f1
---
# PiiGuard
**PiiGuard** is a small, on-device cross NER model from the FlowX **OpenNER** family. Developed by **FlowX.AI**. Runs 100% on-premise / air-gapped, so no data leaves your boundary.
## What it does
- **Task:** token-classification
- **Base model:** `FacebookAI/xlm-roberta-base`
- **Entity types (7):** CARD, DATE, EMAIL, IBAN, NATIONAL_ID, PERSON, PHONE
- **Held-out F1:** 1.0000
- **Runtime:** CPU, Apple Silicon, one GPU, or browser/edge via ONNX (INT8). ~100-160 ms/doc on CPU.
## Why a small model
Fine-tuned encoders match or beat frontier LLMs on structured, convention-bound extraction, at a fraction of the latency and cost, with **zero data egress**. Identifiers are validated by checksum (IBAN mod-97, card Luhn, ISIN/LEI, container ISO-6346, VIN, national IDs), a correctness guarantee general LLMs lack. See the FlowX OpenNER benchmark for measured results.
## Usage
```python
from transformers import AutoTokenizer, AutoModelForTokenClassification
tok = AutoTokenizer.from_pretrained("flowxai/piiguard")
model = AutoModelForTokenClassification.from_pretrained("flowxai/piiguard")
```
## License & attribution
Licensed under the **Apache License 2.0**. Copyright 2026 **FlowX.AI** (https://flowx.ai). See the `NOTICE` file. Trained on synthetic, checksum-validated data.
_Part of the FlowX OpenNER model family. Synthetic-data F1 reflects an in-distribution synthetic distribution; validate on real documents before production use._