Token Classification
Transformers
Safetensors
English
bert
ner
named-entity-recognition
job-postings
distilled
Instructions to use AchrafSoltani/jobbert-ner-haiku-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AchrafSoltani/jobbert-ner-haiku-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="AchrafSoltani/jobbert-ner-haiku-v1")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("AchrafSoltani/jobbert-ner-haiku-v1") model = AutoModelForTokenClassification.from_pretrained("AchrafSoltani/jobbert-ner-haiku-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| license: cc-by-nc-4.0 | |
| language: | |
| - en | |
| library_name: transformers | |
| pipeline_tag: token-classification | |
| tags: | |
| - ner | |
| - named-entity-recognition | |
| - job-postings | |
| - distilled | |
| - bert | |
| metrics: | |
| - f1 | |
| - precision | |
| - recall | |
| # jobbert-ner-haiku-v1 | |
| Distilled Named Entity Recognition model for English-language job postings. One of six students produced for the paper *Distributed Agentic NER on Spark: A Teacher-Student Pipeline for Large-Scale Entity Extraction from Job Postings* (Soltani 2026). | |
| - **Teacher:** Claude Haiku 4.5 (labels acquired via AWS Bedrock) | |
| - **Architecture:** jjzha/jobbert-base-cased fine-tuned for 8-class token classification | |
| - **Student identifier:** `s4_jobbert_haiku` | |
| - **Artefact size:** ~820 MB | |
| Latency is materially lower than the spaCy students (~15 ms vs ~44 ms per document on the gold set). F1 is materially lower because the 512-token window cannot reach the trailing part of most postings. | |
| ## Intended use | |
| Entity extraction from English-language job-posting descriptions into an eight-type schema: | |
| `SKILL`, `JOB_TITLE`, `COMPANY`, `LOCATION`, `EXPERIENCE_LEVEL`, `EDUCATION`, `CERT`, `COMPENSATION`. | |
| Appropriate downstream applications include posting indexing for search and analytics, skill-demand aggregation for labour-market research, cost-quality-speed benchmarking of distilled NER, and teaching use in NLP / distillation courses. | |
| ## Out-of-scope use | |
| Not suitable for: | |
| - CVs or résumés (different register; a CV-trained model should be used instead). | |
| - Non-English postings. | |
| - Fully-automated candidate screening or hiring decisions; downstream ranking or filtering should be built only after an application-side schema and bias review (see *Ethical considerations*). | |
| - Medical, legal, financial or other high-stakes decision support. | |
| - Posting text from languages or locales for which the underlying teacher labels were not representative. | |
| ## Training | |
| - **Teacher labels:** 5,000 stratified postings labelled by Claude Haiku 4.5 in a single run at temperature 0. `max_tokens` was raised from 4,096 to 8,192 mid-run after two truncation failures on entity-dense postings; final labels from the fixed-ceiling run were used. | |
| - **Curator:** 80/10/10 train/dev/test split by `md5(job_link) mod 10`, so Sonnet- and Haiku-trained students see the same posting partitions. | |
| - **Hardware:** one NVIDIA A10G 24 GB GPU (AWS g5.xlarge). | |
| - **Training seed:** 42. | |
| - **Principal hyperparameters and full training spec:** `pipeline/training/experiments/specs/s4_jobbert_haiku.yaml` in the accompanying project repository. | |
| ## Evaluation | |
| Sonnet-trained students evaluate on all 516 gold postings; Haiku-trained students evaluate on 515 because one posting was dropped by the curator for zero-entity teacher output during the Haiku run. Metric: micro-F1 over exact `(text, type)` tuples; character-offset matching is relaxed. Entities are deduplicated within a posting before comparison. | |
| | Overall | Value | | |
| |---|---| | |
| | Micro-F1 | **0.2844** | | |
| | Precision | 0.3838 | | |
| | Recall | 0.2259 | | |
| | 95% CI | [0.276, 0.293] (entity-level delta method) | | |
| | Latency mean (eval hardware) | 15.28 ms / document | | |
| | Latency p99 (eval hardware) | 17.6 ms / document | | |
| | Text coverage | first 512 BERT tokens (~first third of an average posting at 3,996 characters mean length) | | |
| | Postings evaluated | 515 (of the 516-posting gold set) | | |
| ### Per-entity-type | |
| | Entity type | P | R | F1 | | |
| |---|---|---|---| | |
| | COMPANY | 0.636 | 0.470 | 0.540 | | |
| | JOB_TITLE | 0.615 | 0.473 | 0.535 | | |
| | LOCATION | 0.552 | 0.458 | 0.501 | | |
| | COMPENSATION | 0.235 | 0.067 | 0.104 | | |
| | EDUCATION | 0.192 | 0.030 | 0.051 | | |
| | CERT | 0.202 | 0.062 | 0.095 | | |
| | EXPERIENCE_LEVEL | 0.066 | 0.014 | 0.023 | | |
| | SKILL | 0.106 | 0.096 | 0.101 | | |
| ### Teacher comparison | |
| The teacher (Claude Haiku 4.5) reaches micro-F1 = 0.5411 against the same gold set (95% bootstrap CI [0.524, 0.558]). The student trails the teacher by 0.257 points absolute (47.4% relative). See paper §4.3 for the full comparison and the error-mode analysis of this student's residuals. | |
| ## Usage | |
| ```python | |
| from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline | |
| tokenizer = AutoTokenizer.from_pretrained("AchrafSoltani/jobbert-ner-haiku-v1") | |
| model = AutoModelForTokenClassification.from_pretrained("AchrafSoltani/jobbert-ner-haiku-v1") | |
| ner = pipeline("token-classification", model=model, tokenizer=tokenizer, | |
| aggregation_strategy="simple") | |
| text = 'Senior Machine Learning Engineer at Acme Corp in Berlin. Requires 5+ years of experience with PyTorch, AWS, and Kubernetes. MSc in Computer Science preferred. Salary $140,000 – $180,000.' | |
| for ent in ner(text): | |
| print(ent["word"], "->", ent["entity_group"]) | |
| # Note: the BERT base tokeniser has a 512-token window; very long postings | |
| # are truncated. For full-text coverage, use the spaCy variant. | |
| ``` | |
| ## Ethical considerations | |
| This model extracts entities from job postings, a document class whose downstream consumers are typically hiring, ranking, or matching systems. Three cautions are transplanted from paper §6: | |
| - **Schema-induced bias.** SKILL over-extraction is inherited from the LLM teacher; soft-skill phrases ("communication skills", "interpersonal skills") and generic tools ("Excel", "CRM") are over-represented relative to a tighter gold standard. A downstream ranker that treats such phrases as filters is encoding the teacher's lexical habits as a hiring criterion and is not recommended without a schema review at the application layer. | |
| - **Contested ground truth.** A vendor benchmark in the paper against LinkedIn's own `job_skills.csv` on 938,028 jointly-present postings yielded 9.56% agreement and 56.44% discovery: the two extraction schemas produce largely non-overlapping views of the same corpus. Neither constitutes a ground truth; the numbers measure schema divergence, not model quality. | |
| - **Consent and licensing.** The training corpus is a publicly-released Kaggle redistribution of scraped LinkedIn postings. Individuals named in postings (recruiters, hiring managers) did not consent to having their role descriptions re-processed for research. The model is licensed CC BY-NC 4.0 for research and non-commercial evaluation only; any commercial deployment requires a separate legal and ethical review against the data-provenance chain. | |
| ## Limitations | |
| - Trained and evaluated on English-language LinkedIn postings from a publicly-released 2024 Kaggle redistribution; generalisation to other platforms (Indeed, Stack Overflow, regional job boards) or other languages is unevaluated. | |
| - Gold set is single-annotator (516 postings). Intra-annotator stability was scheduled to be measured one week after the main annotation pass; users should treat the reported F1 as having an un-quantified annotator-noise floor until that number lands. | |
| - Output schema is locked to the eight types above. Finer-grained or taxonomy-aligned schemas require re-training against new labels. | |
| - The underlying BERT tokeniser has a 512-token window; the average gold-set posting is 3,996 characters, so approximately the first third of a typical posting is in context. Entities that appear only in the tail (often qualifications, certifications, benefits) are systematically missed. For full-text coverage, prefer the spaCy variant. | |
| ## Citation | |
| ```bibtex | |
| @unpublished{soltani2026distilledner, | |
| author = {Achraf Soltani}, | |
| title = {Distributed Agentic NER on Spark: A Teacher-Student Pipeline for Large-Scale Entity Extraction from Job Postings}, | |
| year = {2026}, | |
| note = {Advisor: Prof.\ Hanine Mohamed}, | |
| url = {https://github.com/AchrafSoltani/distributed-agentic-ner-on-spark}, | |
| } | |
| ``` | |
| ## Licence | |
| - Model weights: **CC BY-NC 4.0** — research and non-commercial evaluation only. | |
| - Source code in the accompanying repository: **Apache 2.0**. | |