Token Classification
Transformers
Safetensors
Thai
camembert
ner
thai
phayathaibert
job-posting
apple-silicon
Eval Results (legacy)
Instructions to use chayuto/thai-job-ner-phayathaibert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use chayuto/thai-job-ner-phayathaibert with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="chayuto/thai-job-ner-phayathaibert")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("chayuto/thai-job-ner-phayathaibert") model = AutoModelForTokenClassification.from_pretrained("chayuto/thai-job-ner-phayathaibert", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 6,400 Bytes
1700972 b2e8ebd 1700972 b2e8ebd 92e14d4 1700972 b2e8ebd 92e14d4 1700972 b2e8ebd 1700972 92e14d4 1700972 92e14d4 1700972 92e14d4 1700972 b2e8ebd 92e14d4 1700972 92e14d4 b2e8ebd 92e14d4 b2e8ebd 92e14d4 b2e8ebd 92e14d4 b2e8ebd 92e14d4 b2e8ebd 92e14d4 1700972 92e14d4 1700972 92e14d4 b2e8ebd 92e14d4 b2e8ebd 92e14d4 1700972 b2e8ebd 1700972 92e14d4 1700972 92e14d4 1700972 92e14d4 b2e8ebd 1700972 92e14d4 1700972 92e14d4 1700972 92e14d4 b2e8ebd 92e14d4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 | ---
language:
- th
license: mit
library_name: transformers
tags:
- token-classification
- ner
- thai
- phayathaibert
- job-posting
- apple-silicon
datasets:
- chayuto/thai-job-ner-dataset
metrics:
- f1
- precision
- recall
pipeline_tag: token-classification
model-index:
- name: thai-job-ner-phayathaibert
results:
- task:
type: token-classification
name: Named Entity Recognition
metrics:
- name: F1 (real-world)
type: f1
value: 0.975
- name: F1 (mixed test)
type: f1
value: 0.929
- name: Precision
type: precision
value: 0.960
- name: Recall
type: recall
value: 0.990
---
# Thai Job NER — Fine-tuned PhayaThaiBERT (v2: Mixed Training)
Named Entity Recognition model for extracting structured HR data from informal Thai job postings (e.g., Facebook groups, Line chats). Fine-tuned from [PhayaThaiBERT](https://huggingface.co/clicknext/phayathaibert) (~122M params).
**v2 update:** Trained on style-matched synthetic + real-world data with upsampling. Real-world F1 improved from 0.143 → **0.975**.
## Model Description
This model extracts 7 entity types from Thai job-related text:
| Entity | Description | Example |
|--------|-------------|---------|
| `HARD_SKILL` | Skills or procedures | ดูแลผู้สูงอายุ, CPR, Python |
| `PERSON` | Names | คุณสมชาย, พี่แจน |
| `LOCATION` | Places | สีลม, ลาดพร้าว, บางนา |
| `COMPENSATION` | Pay amounts | 18,000 บาท/เดือน |
| `EMPLOYMENT_TERMS` | Job structure | part-time, กะกลางวัน |
| `CONTACT` | Phone, Line, email | 081-234-5678, @care123 |
| `DEMOGRAPHIC` | Age, gender | อายุ 25-40, หญิง |
## Usage
```python
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
model_name = "chayuto/thai-job-ner-phayathaibert"
ner = pipeline("ner", model=model_name, aggregation_strategy="simple")
text = "รับสมัครคนดูแลผู้สูงอายุ ย่านสีลม เงินเดือน 18,000 บาท โทร 081-234-5678"
results = ner(text)
for entity in results:
print(f"{entity['entity_group']}: {entity['word']} ({entity['score']:.2%})")
```
## Training
- **Base model:** `clicknext/phayathaibert` (CamemBERT architecture, ~122M params, XLM-R-derived vocabulary)
- **Training data:** 619 posts (510 style-matched synthetic + 54 real-world posts upsampled 5x), silver labels from GPT-4o, fuzzy-aligned to IOB2
- **Real data proportion:** 34.6% (key factor for real-world performance)
- **Hardware:** Apple Silicon MPS backend, FP32
- **Hyperparameters:** LR=3e-5, warmup=0.1, batch=2, grad_accum=4, 10 epochs, gradient checkpointing, frozen embeddings
- **Training time:** ~35 min
### Data Pipeline
Raw Thai text + GPT-4o entity extractions → fuzzy alignment with rapidfuzz + pythainlp TCC boundary snapping → subword token mapping via offset_mapping → IOB2-formatted HuggingFace Dataset.
### Training Strategy
Style-matched synthetic data (v3) was generated to mimic real-world post characteristics: varied formatting, code-switching, informal language. Real-world posts were upsampled 5x to achieve a 20-35% real data proportion — the sweet spot identified through systematic experiments (see below).
## Evaluation
### Real-World Test Set (8 held-out real posts, 98 entities)
| Metric | Score |
|--------|-------|
| **F1** | **0.975** |
| Precision | 0.960 |
| Recall | 0.990 |
### Real-World Per-Entity F1
| Entity | F1 | Precision | Recall | Support |
|--------|-----|-----------|--------|---------|
| COMPENSATION | 1.000 | 1.000 | 1.000 | 15 |
| CONTACT | 1.000 | 1.000 | 1.000 | 10 |
| DEMOGRAPHIC | 1.000 | 1.000 | 1.000 | 14 |
| LOCATION | 1.000 | 1.000 | 1.000 | 8 |
| PERSON | 1.000 | 1.000 | 1.000 | 3 |
| HARD_SKILL | 0.973 | 0.947 | 1.000 | 36 |
| EMPLOYMENT_TERMS | 0.880 | 0.846 | 0.917 | 12 |
5 of 7 entity types achieve **perfect F1** on real-world data.
### Mixed Test Set (78 examples)
| Metric | Score |
|--------|-------|
| **F1** | **0.929** |
| Precision | 0.915 |
| Recall | 0.944 |
### Experiment History: Closing the Real-World Gap
| Version | Training Data | Real-World F1 | Key Insight |
|---------|--------------|---------------|-------------|
| v1 (synthetic-only) | 1,253 synthetic | 0.143 | Synthetic alone fails on real data |
| v1 (real-only) | 37 real posts | 0.558 | Small but 4x more effective per sample |
| v1 (mixed, 4.1% real) | 1,057 mixed | 0.935 | Even a few real posts help dramatically |
| **v2 (v3+real5x, 34.6% real)** | **619 mixed** | **0.975** | **Style-matched synthetic + upsampling** |
Key finding: real data proportion is the #1 factor. Style-matched synthetic data is 3x more efficient than generic synthetic data.
## Links
- **Model:** [chayuto/thai-job-ner-phayathaibert](https://huggingface.co/chayuto/thai-job-ner-phayathaibert)
- **WangchanBERTa variant:** [chayuto/thai-job-ner-wangchanberta](https://huggingface.co/chayuto/thai-job-ner-wangchanberta)
- **Dataset:** [chayuto/thai-job-ner-dataset](https://huggingface.co/datasets/chayuto/thai-job-ner-dataset)
- **Source Code:** [github.com/chayuto/thai-job-nlp-ner](https://github.com/chayuto/thai-job-nlp-ner)
## Limitations
- Real-world evaluation is based on 8 held-out posts (98 entities) — confidence intervals are wide
- EMPLOYMENT_TERMS remains the weakest entity (F1=0.880) — boundary ambiguity in schedule/contract terms
- Embeddings were frozen during training (MPS memory constraint) — unfreezing on a larger GPU may yield further gains
- 256 token max sequence length (covers >95% of real posts)
- Larger model file size due to 248K vocabulary (vs WangchanBERTa's 25K)
## Technical Notes
- **FP16 is broken on MPS** — always use FP32 for Apple Silicon training
- PhayaThaiBERT's 248K vocab (XLM-R-derived) requires frozen embeddings + gradient checkpointing to fit on 18GB MPS
- Uses `offset_mapping` for tokenizer-agnostic subword-to-character alignment
- Thai Character Cluster (TCC) boundary snapping prevents Unicode grapheme splitting during alignment
- Real data upsampling (5-10x) is a simple, effective technique for low-resource scenarios
## License
MIT
|