jobbert-ner-haiku-v1-onnx

Distilled Named Entity Recognition model for English-language job postings. One of six students produced for the paper Distributed NER on Spark: A Teacher-Student Pipeline for Large-Scale Entity Extraction from Job Postings (Soltani and Hanine 2026).

  • Teacher: Claude Haiku 4.5 (labels acquired via AWS Bedrock)
  • Architecture: S4 (JobBERT fine-tune) exported and quantised to int8 ONNX via optimum + onnxruntime
  • Student identifier: s6_jobbert_onnx_haiku
  • Artefact size: ~105 MB

Weights compressed to ~4× smaller than the dense JobBERT. In our evaluation, ONNX Runtime fell back to a CPU int8 kernel that ran slower than the dense FP16 GPU path; on hardware with a proper VNNI or GPU int8 kernel the latency relationship is expected to invert. Size is the clear win; latency is hardware-dependent.

Intended use

Entity extraction from English-language job-posting descriptions into an eight-type schema:

SKILL, JOB_TITLE, COMPANY, LOCATION, EXPERIENCE_LEVEL, EDUCATION, CERT, COMPENSATION.

Appropriate downstream applications include posting indexing for search and analytics, skill-demand aggregation for labour-market research, cost-quality-speed benchmarking of distilled NER, and teaching use in NLP / distillation courses.

Out-of-scope use

Not suitable for:

  • CVs or résumés (different register; a CV-trained model should be used instead).
  • Non-English postings.
  • Fully-automated candidate screening or hiring decisions; downstream ranking or filtering should be built only after an application-side schema and bias review (see Ethical considerations).
  • Medical, legal, financial or other high-stakes decision support.
  • Posting text from languages or locales for which the underlying teacher labels were not representative.

Training

  • Teacher labels: 5,000 stratified postings labelled by Claude Haiku 4.5 in a single run at temperature 0. max_tokens was raised from 4,096 to 8,192 mid-run after two truncation failures on entity-dense postings; final labels from the fixed-ceiling run were used.
  • Curator: 80/10/10 train/dev/test split by md5(job_link) mod 10, so Sonnet- and Haiku-trained students see the same posting partitions.
  • Hardware: one NVIDIA A10G 24 GB GPU (AWS g5.xlarge).
  • Training seed: 42.
  • Principal hyperparameters and full training spec: pipeline/training/experiments/specs/s6_jobbert_onnx_haiku.yaml in the accompanying project repository.

Evaluation

Sonnet-trained students evaluate on all 516 gold postings; Haiku-trained students evaluate on 515 because one posting was dropped by the curator for zero-entity teacher output during the Haiku run. Metric: micro-F1 over exact (text, type) tuples; character-offset matching is relaxed. Entities are deduplicated within a posting before comparison.

Overall Value
Micro-F1 0.2609
Precision 0.3431
Recall 0.2105
95% CI [0.252, 0.269] (entity-level delta method)
Latency mean (eval hardware) 264.14 ms / document
Latency p99 (eval hardware) 346.99 ms / document
Text coverage first 512 BERT tokens
Postings evaluated 515 (of the 516-posting gold set)

Per-entity-type

Per-entity numbers below reflect the coverage constraint as much as the model's per-type quality. Entities that appear only in the trailing part of a long posting (typically CERT, EDUCATION, EXPERIENCE_LEVEL, and COMPENSATION in many English templates) are systematically outside the model's input window and therefore missed at the recall metric even when the model would classify them correctly on shorter text. For full-text coverage, use the spaCy variant.

Entity type P R F1
COMPANY 0.575 0.406 0.476
JOB_TITLE 0.553 0.455 0.499
LOCATION 0.524 0.448 0.483
COMPENSATION 0.200 0.055 0.087
EDUCATION 0.154 0.024 0.041
CERT 0.184 0.059 0.089
EXPERIENCE_LEVEL 0.051 0.012 0.019
SKILL 0.094 0.094 0.094

Teacher comparison

The teacher (Claude Haiku 4.5) reaches micro-F1 = 0.5411 against the same gold set (95% bootstrap CI [0.524, 0.558]). The student trails the teacher by 0.280 points absolute (51.8% relative). See paper §4.3 for the full comparison and the error-mode analysis of this student's residuals.

Usage

from optimum.onnxruntime import ORTModelForTokenClassification
from transformers          import AutoTokenizer, pipeline

tokenizer = AutoTokenizer.from_pretrained("AchrafSoltani/jobbert-ner-haiku-v1-onnx")
model     = ORTModelForTokenClassification.from_pretrained(
                "AchrafSoltani/jobbert-ner-haiku-v1-onnx", file_name="model_quantized.onnx")
ner       = pipeline("token-classification", model=model, tokenizer=tokenizer,
                     aggregation_strategy="simple")

text = 'Senior Machine Learning Engineer at Acme Corp in Berlin. Requires 5+ years of experience with PyTorch, AWS, and Kubernetes. MSc in Computer Science preferred. Salary $140,000 – $180,000.'
for ent in ner(text):
    print(ent["word"], "->", ent["entity_group"])

# Produces (verified on this release; note the BERT wordpiece tokenisation
# artefacts in numeric spans):
# Senior Machine Learning Engineer     -> JOB_TITLE         (0.68)
# Acme Corp                            -> COMPANY           (0.74)
# Berlin                               -> LOCATION          (0.62)
# 5 + years of experience              -> EXPERIENCE_LEVEL  (0.76)
# PyTorch                              -> SKILL             (0.60)
# AWS                                  -> SKILL             (0.51)
# Kubernetes                           -> SKILL             (0.57)
# MSc in Computer Science              -> EDUCATION         (0.57)
# $ 140, 000 – $ 180, 000              -> COMPENSATION      (0.91)

# Note: latency depends on the ONNX Runtime execution provider available on
# the host. On CPUs without VNNI support, the int8 kernel is slower than the
# dense JobBERT on FP16 GPU; on VNNI CPUs or GPU ORT, the ordering is
# expected to invert. Size is the consistent win: ~4× smaller artefact.

Ethical considerations

This model extracts entities from job postings, a document class whose downstream consumers are typically hiring, ranking, or matching systems. Three cautions are transplanted from paper §6:

  • Schema-induced bias. SKILL over-extraction is inherited from the LLM teacher; soft-skill phrases ("communication skills", "interpersonal skills") and generic tools ("Excel", "CRM") are over-represented relative to a tighter gold standard. A downstream ranker that treats such phrases as filters is encoding the teacher's lexical habits as a hiring criterion and is not recommended without a schema review at the application layer.
  • Contested ground truth. A vendor benchmark in the paper against LinkedIn's own job_skills.csv on 938,028 jointly-present postings yielded 9.56% agreement and 56.44% discovery: the two extraction schemas produce largely non-overlapping views of the same corpus. Neither constitutes a ground truth; the numbers measure schema divergence, not model quality.
  • Consent and licensing. The training corpus is a publicly-released Kaggle redistribution of scraped LinkedIn postings. Individuals named in postings (recruiters, hiring managers) did not consent to having their role descriptions re-processed for research. The model is licensed CC BY-NC 4.0 for research and non-commercial evaluation only; any commercial deployment requires a separate legal and ethical review against the data-provenance chain.

Limitations

  • Trained and evaluated on English-language LinkedIn postings from a publicly-released 2024 Kaggle redistribution; generalisation to other platforms (Indeed, Stack Overflow, regional job boards) or other languages is unevaluated.
  • Gold set is single-annotator (516 postings). Intra-annotator stability was scheduled to be measured one week after the main annotation pass; users should treat the reported F1 as having an un-quantified annotator-noise floor until that number lands.
  • Output schema is locked to the eight types above. Finer-grained or taxonomy-aligned schemas require re-training against new labels.
  • The underlying BERT tokeniser has a 512-token window; the average gold-set posting is 3,996 characters, so approximately the first third of a typical posting is in context. Entities that appear only in the tail (often qualifications, certifications, benefits) are systematically missed. For full-text coverage, prefer the spaCy variant.
  • Latency is hardware-dependent. On CPUs without VNNI support, the ONNX Runtime int8 kernel falls back to a slower generic path; in our evaluation this was slower than the dense JobBERT on FP16 GPU. The latency number in the evaluation table reflects that CPU fallback on our evaluation hardware.

Citation

@unpublished{soltani2026distilledner,
  author = {Achraf Soltani and Mohamed Hanine},
  title  = {Distributed NER on Spark: A Teacher-Student Pipeline for Large-Scale Entity Extraction from Job Postings},
  year   = {2026},
  note   = {Advisor: Prof.\ Hanine Mohamed},
  url    = {https://github.com/achrafsoltani/distributed-ner-on-spark},
}

Licence

  • Model weights: CC BY-NC 4.0 — research and non-commercial evaluation only.
  • Source code in the accompanying repository: Apache 2.0.
Downloads last month
7
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support