Instructions to use rafmacalaba/gliner2-datause-large-v1-deval-synth-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use rafmacalaba/gliner2-datause-large-v1-deval-synth-v2 with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("rafmacalaba/gliner2-datause-large-v1-deval-synth-v2") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
File size: 2,447 Bytes
d3fe46d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 | ---
tags:
- gliner2
- ner
- data-mention-extraction
- lora
base_model: fastino/gliner2-large-v1
library_name: gliner2
---
# GLiNER2 Data Mention Extractor (v1-deval-synth-v2)
Fine-tuned GLiNER2 LoRA adapter for extracting structured data mentions from
development economics and humanitarian research documents.
**v4 key change:** Trained with a patched GLiNER2 library
(`rafmacalaba/GLiNER2@feat/main-mirror`) that feeds mean-pooled passage token embeddings
into `count_pred` instead of the schema `[P]` token. This is the first adapter
version where multi-mention recall is trained via a meaningful gradient.
## Task
Given a document passage, extracts structured information about each dataset mentioned:
- **Extractive field**: `mention_name` (verbatim from text)
- **Classification fields** (fixed choices):
- `specificity_tag`: named / descriptive / vague
- `typology_tag`: survey / census / database / administrative / indicator / geospatial / microdata / report / other
- `is_used`: True / False
- `usage_context`: primary / supporting / background
## Training
- **Base model**: `fastino/gliner2-large-v1`
- **Method**: LoRA (r=16, alpha=32.0)
- **Target modules**: ['encoder', 'span_rep', 'classifier', 'count_embed', 'count_pred']
- **Training examples**: 8791
- **Val examples**: 651
- **Best val loss**: 439.4476
- **GLiNER2 branch**: `rafmacalaba/GLiNER2@feat/main-mirror`
## Usage
```python
from gliner2 import GLiNER2
# Install the patched library first
# pip install git+https://github.com/rafmacalaba/GLiNER2.git@feat/main-mirror
extractor = GLiNER2.from_pretrained("fastino/gliner2-large-v1")
extractor.load_adapter("rafmacalaba/gliner2-datause-large-v1-deval-synth-v2")
schema = (
extractor.create_schema()
.structure("data_mention")
.field("mention_name", dtype="str")
.field("specificity_tag", dtype="str", choices=["named", "descriptive", "vague", "na"])
.field("typology_tag", dtype="str", choices=["survey", "census", "administrative",
"database", "indicator", "geospatial",
"microdata", "report", "other", "na"])
.field("is_used", dtype="str", choices=["True", "False", "na"])
.field("usage_context", dtype="str", choices=["primary", "supporting", "background", "na"])
)
result = extractor.extract(text, schema)
```
|