Instructions to use AchrafSoltani/spacy-lg-jobposting-ner-sonnet-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- spaCy
How to use AchrafSoltani/spacy-lg-jobposting-ner-sonnet-v1 with spaCy:
!pip install https://huggingface.co/AchrafSoltani/spacy-lg-jobposting-ner-sonnet-v1/resolve/main/spacy-lg-jobposting-ner-sonnet-v1-any-py3-none-any.whl # Using spacy.load(). import spacy nlp = spacy.load("spacy-lg-jobposting-ner-sonnet-v1") # Importing as module. import spacy-lg-jobposting-ner-sonnet-v1 nlp = spacy-lg-jobposting-ner-sonnet-v1.load() - Notebooks
- Google Colab
- Kaggle
spacy-lg-jobposting-ner-sonnet-v1
Distilled Named Entity Recognition model for English-language job postings. One of six students produced for the paper Distributed NER on Spark: A Teacher-Student Pipeline for Large-Scale Entity Extraction from Job Postings (Soltani and Hanine 2026).
- Teacher: Claude Sonnet 4.6 (labels acquired via AWS Bedrock)
- Architecture: spaCy en_core_web_lg with NER head fine-tuned on teacher labels
- Student identifier:
s1_spacy_sonnet - Artefact size: ~835 MB
Intended use
Entity extraction from English-language job-posting descriptions into an eight-type schema:
SKILL, JOB_TITLE, COMPANY, LOCATION, EXPERIENCE_LEVEL, EDUCATION, CERT, COMPENSATION.
Appropriate downstream applications include posting indexing for search and analytics, skill-demand aggregation for labour-market research, cost-quality-speed benchmarking of distilled NER, and teaching use in NLP / distillation courses.
Out-of-scope use
Not suitable for:
- CVs or résumés (different register; a CV-trained model should be used instead).
- Non-English postings.
- Fully-automated candidate screening or hiring decisions; downstream ranking or filtering should be built only after an application-side schema and bias review (see Ethical considerations).
- Medical, legal, financial or other high-stakes decision support.
- Posting text from languages or locales for which the underlying teacher labels were not representative.
Training
- Teacher labels: 5,000 stratified postings labelled by Claude Sonnet 4.6 in a single run at temperature 0.
max_tokenswas raised from 4,096 to 8,192 mid-run after two truncation failures on entity-dense postings; final labels from the fixed-ceiling run were used. - Curator: 80/10/10 train/dev/test split by
md5(job_link) mod 10, so Sonnet- and Haiku-trained students see the same posting partitions. - Hardware: one NVIDIA A10G 24 GB GPU (AWS g5.xlarge).
- Training seed: 42.
- Principal hyperparameters and full training spec:
pipeline/training/experiments/specs/s1_spacy_sonnet.yamlin the accompanying project repository.
Evaluation
Sonnet-trained students evaluate on all 516 gold postings; Haiku-trained students evaluate on 515 because one posting was dropped by the curator for zero-entity teacher output during the Haiku run. Metric: micro-F1 over exact (text, type) tuples; character-offset matching is relaxed. Entities are deduplicated within a posting before comparison.
| Overall | Value |
|---|---|
| Micro-F1 | 0.5082 |
| Precision | 0.4201 |
| Recall | 0.6430 |
| 95% CI | [0.494, 0.523] (posting-level bootstrap, 10,000 resamples) |
| Latency mean (eval hardware) | 44.05 ms / document |
| Latency p99 (eval hardware) | 97.08 ms / document |
| Text coverage | full text (no token-window truncation) |
| Postings evaluated | 516 (of the 516-posting gold set) |
Per-entity-type
| Entity type | P | R | F1 |
|---|---|---|---|
| COMPANY | 0.701 | 0.601 | 0.647 |
| JOB_TITLE | 0.669 | 0.621 | 0.644 |
| LOCATION | 0.608 | 0.794 | 0.689 |
| COMPENSATION | 0.537 | 0.723 | 0.616 |
| EDUCATION | 0.460 | 0.407 | 0.432 |
| CERT | 0.522 | 0.492 | 0.507 |
| EXPERIENCE_LEVEL | 0.278 | 0.296 | 0.287 |
| SKILL | 0.209 | 0.686 | 0.320 |
Teacher comparison
The teacher (Claude Sonnet 4.6) reaches micro-F1 = 0.5171 against the same gold set (95% bootstrap CI [0.503, 0.530]). The student trails the teacher by 0.009 points absolute (1.7% relative). See paper §4.3 for the full comparison and the error-mode analysis of this student's residuals.
Usage
import spacy
from huggingface_hub import snapshot_download
local = snapshot_download(repo_id="AchrafSoltani/spacy-lg-jobposting-ner-sonnet-v1")
nlp = spacy.load(local)
text = 'Senior Machine Learning Engineer at Acme Corp in Berlin. Requires 5+ years of experience with PyTorch, AWS, and Kubernetes. MSc in Computer Science preferred. Salary $140,000 – $180,000.'
doc = nlp(text)
for ent in doc.ents:
print(ent.text, "->", ent.label_)
# Produces (verified on this release):
# Senior Machine Learning Engineer -> JOB_TITLE
# Acme Corp -> COMPANY
# Berlin -> LOCATION
# 5+ years of experience -> EXPERIENCE_LEVEL
# PyTorch -> SKILL
# AWS -> SKILL
# Kubernetes -> SKILL
# MSc in Computer Science -> EDUCATION
# $140,000 – $180,000 -> COMPENSATION
Ethical considerations
This model extracts entities from job postings, a document class whose downstream consumers are typically hiring, ranking, or matching systems. Three cautions are transplanted from paper §6:
- Schema-induced bias. SKILL over-extraction is inherited from the LLM teacher; soft-skill phrases ("communication skills", "interpersonal skills") and generic tools ("Excel", "CRM") are over-represented relative to a tighter gold standard. A downstream ranker that treats such phrases as filters is encoding the teacher's lexical habits as a hiring criterion and is not recommended without a schema review at the application layer.
- Contested ground truth. A vendor benchmark in the paper against LinkedIn's own
job_skills.csvon 938,028 jointly-present postings yielded 9.56% agreement and 56.44% discovery: the two extraction schemas produce largely non-overlapping views of the same corpus. Neither constitutes a ground truth; the numbers measure schema divergence, not model quality. - Consent and licensing. The training corpus is a publicly-released Kaggle redistribution of scraped LinkedIn postings. Individuals named in postings (recruiters, hiring managers) did not consent to having their role descriptions re-processed for research. The model is licensed CC BY-NC 4.0 for research and non-commercial evaluation only; any commercial deployment requires a separate legal and ethical review against the data-provenance chain.
Limitations
- Trained and evaluated on English-language LinkedIn postings from a publicly-released 2024 Kaggle redistribution; generalisation to other platforms (Indeed, Stack Overflow, regional job boards) or other languages is unevaluated.
- Gold set is single-annotator (516 postings). Intra-annotator stability was scheduled to be measured one week after the main annotation pass; users should treat the reported F1 as having an un-quantified annotator-noise floor until that number lands.
- Output schema is locked to the eight types above. Finer-grained or taxonomy-aligned schemas require re-training against new labels.
Citation
@unpublished{soltani2026distilledner,
author = {Achraf Soltani and Mohamed Hanine},
title = {Distributed NER on Spark: A Teacher-Student Pipeline for Large-Scale Entity Extraction from Job Postings},
year = {2026},
note = {Advisor: Prof.\ Hanine Mohamed},
url = {https://github.com/achrafsoltani/distributed-ner-on-spark},
}
Licence
- Model weights: CC BY-NC 4.0 — research and non-commercial evaluation only.
- Source code in the accompanying repository: Apache 2.0.
- Downloads last month
- -