Instructions to use GOSHUNCLE/pii-masking-zh-tw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use GOSHUNCLE/pii-masking-zh-tw with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("C:\Users\chris.ma_cycraft\.cache\huggingface\hub\models--Qwen--Qwen2.5-1.5B-Instruct\snapshots\989aa7980e4cf806f80c7fef2b1adb7bc71aa306") model = PeftModel.from_pretrained(base_model, "GOSHUNCLE/pii-masking-zh-tw") - Notebooks
- Google Colab
- Kaggle
PII Masking ZH-TW
A LoRA adapter on top of Qwen2.5-1.5B-Instruct that detects and masks 8 types of Taiwan-specific personal information (PII), returning structured JSON output. Designed for compliance with Taiwan's Personal Data Protection Act (PDPA) and similar privacy requirements.
Why this model
Most regex-based PII tools handle structured types (phone numbers, ID numbers, emails) well, but struggle with NAME and ADDRESS — the two types that require semantic understanding. This model is tuned to handle exactly those cases, including:
- Names embedded naturally in Chinese sentences
- Addresses with non-standard administrative divisions and floor (鎮 / 鄉, no 區, B1, B2, 5F)
- Mixed character systems (halfwidth, fullwidth, Chinese numerals)
- Distinguishing real addresses from place-name mentions (e.g., "we'll meet at a café in 大安區" → no PII)
Supported PII types
| Type | Description | Example |
|---|---|---|
NAME |
Personal names | 王小明 → 王OO |
ROC_ID |
National ID number | A123456789 → A1******89 |
PHONE |
Mobile number | 0912345678 → 0912****** |
EMAIL |
Email address | alex@gmail.com → al***@gmail.com |
ADDRESS |
Full address | 台北市信義區忠孝東路100號 → 台北市信義區********* |
NHI_ID |
NHI card number | 000812345678 → 0008******** |
BANK_ACCOUNT |
Bank account number | 0123456789 → 0123****** |
CREDIT_CARD |
Credit card number | 4311220012345678 → 4311************ |
Performance
Evaluated on 70 held-out Taiwan-style test cases:
| Metric | Score |
|---|---|
| JSON format accuracy | 100% |
| PII judgment accuracy | 98.6% |
| Entity recall | 100% |
| Mask accuracy (with included safeguards) | ~100% |
Inference safeguards
The included inference.py applies two filters and one mask corrector:
- Structural validation — drops entities whose value shape doesn't match the claimed type (e.g.,
BANK_ACCOUNTvalue containing non-digits, orADDRESSvalue lacking "號" or "樓") - Provenance check — drops entities whose value doesn't verbatim appear in the input (rejects hallucinations and mistranslations)
- Mask correction — recomputes the
maskedfield deterministically fromvalue, guaranteeing correct star count
These guarantee output integrity even when the model occasionally misclassifies or fabricates.
Quick start
pip install -r requirements.txt
from inference import PIIDetector
detector = PIIDetector()
result = detector.detect_and_replace(
"客戶王小明來電諮詢,身分證A123456789,手機0912345678"
)
print(result["masked_text"])
# 客戶王OO來電諮詢,身分證A1******89,手機0912******
print(result["entities"])
# [
# {"type": "NAME", "value": "王小明", "masked": "王OO"},
# {"type": "ROC_ID", "value": "A123456789", "masked": "A1******89"},
# {"type": "PHONE", "value": "0912345678", "masked": "0912******"},
# ]
Try the demo
Output format
The model returns strict JSON:
{
"pii_found": true,
"entities": [
{"type": "NAME", "value": "原始值", "masked": "遮罩值"}
]
}
When no PII is found:
{"pii_found": false, "entities": []}
Requirements
- Python 3.9+
- ~6 GB RAM (CPU inference works without GPU)
- See
requirements.txtfor package versions
License
Apache License 2.0
Disclaimer
All training and demonstration data used in this project is synthetic. The model was not trained on any real personal information. Output values in examples and test cases are programmatically generated and do not correspond to any real individuals.
繁體中文
針對 Qwen2.5-1.5B-Instruct 的 LoRA adapter,偵測並遮罩 8 種台灣常見個人 資訊(PII),輸出結構化 JSON。適用於符合個資法(PDPA)的資料去識別需求。
為什麼要用這個模型
大多數規則型 PII 工具能處理結構化類型(電話、身分證、Email),但對需要 語意理解的 NAME(姓名)與 ADDRESS(地址)束手無策。本模型專門 強化以下案例:
- 自然語句中嵌入的姓名(不再依賴姓氏規則表)
- 非標準行政區地址與樓層表示(鎮、鄉、無區、B1、B3、5F...等)
- 混合字元系統(半形、全形、中文數字)
- 區分完整地址與地名提及(例如「約在大安區的咖啡店」→ 不算地址)
支援的 PII 類型
| 類型 | 說明 | 範例 |
|---|---|---|
NAME |
姓名 | 王小明 → 王OO |
ROC_ID |
身分證字號 | A123456789 → A1******89 |
PHONE |
手機號碼 | 0912345678 → 0912****** |
EMAIL |
電子郵件 | alex@gmail.com → al***@gmail.com |
ADDRESS |
完整地址 | 台北市信義區忠孝東路100號 → 台北市信義區********* |
NHI_ID |
健保卡號 | 000812345678 → 0008******** |
BANK_ACCOUNT |
銀行帳號 | 0123456789 → 0123****** |
CREDIT_CARD |
信用卡號 | 4311220012345678 → 4311************ |
效能
在 70 筆台灣情境測試集評估:
| 指標 | 分數 |
|---|---|
| JSON 格式正確率 | 100% |
| PII 判斷正確率 | 98.6% |
| 實體召回率 | 100% |
| 遮罩正確率(含內建安全機制) | ~100% |
推論安全機制
內附的 inference.py 套用兩道過濾與一道遮罩修正:
- 結構驗證 — 丟棄 value 結構不符其宣稱類型的 entity(例如
BANK_ACCOUNT含非數字字元、或ADDRESS缺「號」或「樓」) - 來源驗證 — 丟棄 value 未逐字出現在原文中的 entity(擋幻覺與翻譯錯誤)
- 遮罩修正 — 從 value 確定性地重算 masked 欄位,保證星號數量精準
這幾層確保即使模型偶有誤判或捏造,輸出仍然可信。
快速使用
pip install -r requirements.txt
from inference import PIIDetector
detector = PIIDetector()
result = detector.detect_and_replace(
"客戶王小明來電諮詢,身分證A123456789,手機0912345678"
)
print(result["masked_text"])
# 客戶王OO來電諮詢,身分證A1******89,手機0912******
線上 Demo
輸出格式
模型輸出嚴格 JSON:
{
"pii_found": true,
"entities": [
{"type": "NAME", "value": "原始值", "masked": "遮罩值"}
]
}
無 PII 時:
{"pii_found": false, "entities": []}
系統需求
- Python 3.9+
- 約 6 GB 記憶體(CPU 即可推論,不需 GPU)
- 套件版本見
requirements.txt
授權
Apache License 2.0
聲明
本專案所有訓練資料與展示資料皆為合成資料,未使用任何真實個人資訊 進行訓練。範例與測試集中的 PII 值皆為程式產生,與任何真實人物無關。
- Downloads last month
- 8