Sentence Similarity
sentence-transformers
Safetensors
Turkish
xlm-roberta
feature-extraction
embeddings
turkish
türkçe
e5
retrieval
semantic-search
mteb
tr-mteb
Eval Results (legacy)
text-embeddings-inference
Instructions to use thealper2/intfloat-multilingual-e5-base-tr-nli with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use thealper2/intfloat-multilingual-e5-base-tr-nli with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("thealper2/intfloat-multilingual-e5-base-tr-nli") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
| language: | |
| - tr | |
| license: cc-by-sa-4.0 | |
| library_name: sentence-transformers | |
| pipeline_tag: sentence-similarity | |
| tags: | |
| - sentence-transformers | |
| - sentence-similarity | |
| - feature-extraction | |
| - embeddings | |
| - turkish | |
| - türkçe | |
| - e5 | |
| - retrieval | |
| - semantic-search | |
| - mteb | |
| - tr-mteb | |
| base_model: intfloat/multilingual-e5-base | |
| datasets: | |
| - mertcobanov/all-nli-triplets-turkish | |
| - emrecan/stsb-mt-turkish | |
| metrics: | |
| - spearmanr | |
| - ndcg | |
| - accuracy | |
| - v_measure | |
| model-index: | |
| - name: e5-tr-nli | |
| results: | |
| - task: | |
| type: semantic-textual-similarity | |
| name: Semantic Textual Similarity | |
| dataset: | |
| name: STSb-TR (test) | |
| type: trmteb/stsb-tr | |
| metrics: | |
| - type: spearman_cosine | |
| value: 0.7907 | |
| - task: | |
| type: retrieval | |
| name: Retrieval | |
| dataset: | |
| name: TQuad | |
| type: trmteb/tquad | |
| metrics: | |
| - type: ndcg_at_10 | |
| value: 0.8222 | |
| - task: | |
| type: retrieval | |
| name: Retrieval | |
| dataset: | |
| name: Quora-TR | |
| type: trmteb/quora-tr | |
| metrics: | |
| - type: ndcg_at_10 | |
| value: 0.7323 | |
| - task: | |
| type: retrieval | |
| name: Retrieval | |
| dataset: | |
| name: MS MARCO-TR | |
| type: trmteb/msmarco-tr | |
| metrics: | |
| - type: ndcg_at_10 | |
| value: 0.2734 | |
| - task: | |
| type: classification | |
| name: Classification | |
| dataset: | |
| name: News-Cat | |
| type: trmteb/news-cat | |
| metrics: | |
| - type: accuracy | |
| value: 0.9600 | |
| # e5-tr-nli — Turkish Sentence Embedding Model (Bi-Encoder) | |
| A **Turkish sentence-embedding (bi-encoder) model** for **retrieval and semantic | |
| search**, fine-tuned from [`intfloat/multilingual-e5-base`](https://huggingface.co/intfloat/multilingual-e5-base) | |
| with contrastive learning (`MultipleNegativesRankingLoss`) on Turkish NLI triplets. | |
| - **Embedding dimension:** 768 | |
| - **Max sequence length:** 256 tokens (trained); backbone supports 512 | |
| - **Pooling:** mean pooling | |
| - **Similarity function:** cosine | |
| - **Parameters:** ~278M (XLM-RoBERTa-base backbone) | |
| - **Prefix convention:** e5-style `query:` / `passage:` (see [Usage](#usage) — **required**) | |
| --- | |
| ## Table of Contents | |
| 1. [Intended Use](#intended-use) | |
| 2. [Usage](#usage) | |
| 3. [Training Data](#training-data) | |
| 4. [Training Procedure](#training-procedure) | |
| 5. [Evaluation](#evaluation) | |
| 6. [TR-MTEB Results](#tr-mteb-results-full) | |
| 7. [Limitations & Biases](#limitations--biases) | |
| 8. [Compute & Environmental Impact](#compute--environmental-impact) | |
| 9. [Licensing](#licensing) | |
| 10. [Citation](#citation) | |
| 11. [Reproduction](#reproduction) | |
| --- | |
| ## Intended Use | |
| **In scope** | |
| - Turkish semantic search / passage retrieval (dense retrieval, RAG) | |
| - Semantic textual similarity, paraphrase & duplicate detection | |
| - Clustering and topic grouping of Turkish text | |
| - Feature extraction for downstream Turkish NLP classifiers | |
| **Out of scope** | |
| - Re-ranking with query–document cross-attention (use a cross-encoder instead) | |
| - Long-document embedding beyond 256 tokens without chunking | |
| - High-stakes decisions (legal, medical, hiring) without human oversight | |
| - Non-Turkish text (backbone is multilingual, but this model is tuned for Turkish) | |
| --- | |
| ## Usage | |
| ### ⚠️ e5 prefixes are mandatory | |
| The model was trained with the e5 asymmetric prefixes and **must** be used with them, | |
| consistently at training / evaluation / inference: | |
| | Role | Prefix | | |
| |------|--------| | |
| | Query / anchor / any symmetric-task sentence (STS, classification, clustering) | `query: ` | | |
| | Passage / document / positive & negative candidates | `passage: ` | | |
| Omitting the prefixes degrades quality noticeably. | |
| ### With `sentence-transformers` | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| from sentence_transformers.util import cos_sim | |
| model = SentenceTransformer("thealper2/intfloat-multilingual-e5-base-tr-nli") | |
| # Asymmetric retrieval: query vs. candidate passages | |
| query = "query: Türkiye'nin başkenti neresidir?" | |
| passages = [ | |
| "passage: Ankara, Türkiye'nin başkentidir.", | |
| "passage: İstanbul Türkiye'nin en kalabalık şehridir.", | |
| "passage: Muz tropikal bir meyvedir.", | |
| ] | |
| q = model.encode(query, normalize_embeddings=True) | |
| p = model.encode(passages, normalize_embeddings=True) | |
| print(cos_sim(q, p)) # highest score -> the Ankara passage | |
| ``` | |
| ### Semantic similarity (symmetric — use `query:` on both sides) | |
| ```python | |
| a = model.encode("query: Bugün hava çok güzel.", normalize_embeddings=True) | |
| b = model.encode("query: Hava bugün oldukça güzel.", normalize_embeddings=True) | |
| print(float(cos_sim(a, b))) # ~0.9 | |
| ``` | |
| ### With 🤗 Transformers (mean pooling) | |
| ```python | |
| import torch, torch.nn.functional as F | |
| from transformers import AutoTokenizer, AutoModel | |
| tok = AutoTokenizer.from_pretrained("thealper2/intfloat-multilingual-e5-base-tr-nli") | |
| mdl = AutoModel.from_pretrained("thealper2/intfloat-multilingual-e5-base-tr-nli") | |
| def encode(texts): | |
| batch = tok(texts, padding=True, truncation=True, max_length=256, return_tensors="pt") | |
| with torch.no_grad(): | |
| out = mdl(**batch) | |
| mask = batch["attention_mask"].unsqueeze(-1).float() | |
| emb = (out.last_hidden_state * mask).sum(1) / mask.sum(1) # mean pooling | |
| return F.normalize(emb, p=2, dim=1) | |
| emb = encode(["query: örnek cümle", "passage: örnek pasaj"]) | |
| ``` | |
| --- | |
| ## Training Data | |
| - **Source:** [`mertcobanov/all-nli-triplets-turkish`](https://huggingface.co/datasets/mertcobanov/all-nli-triplets-turkish) | |
| — a machine-translated Turkish version of the AllNLI (SNLI + MultiNLI) triplet set. | |
| - **Format:** `(anchor, positive, negative)` triplets, where the `negative` acts as a | |
| **hard negative** for the contrastive objective. | |
| - **Column handling:** **Only the Turkish columns** (`anchor_translated`, | |
| `positive_translated`, `negative_translated`) were used and renamed to | |
| `anchor` / `positive` / `negative`. **All English columns were discarded.** | |
| - **Cleaning:** rows with `None` / empty / whitespace-only fields were filtered out. | |
| - **Resulting sizes (after filtering):** | |
| | Split | Triplets | | |
| |-------|---------:| | |
| | train | **277,167** | | |
| | dev | 6,584 | | |
| | test | 6,609 | | |
| --- | |
| ## Training Procedure | |
| ### Objective | |
| - **Loss:** `CachedMultipleNegativesRankingLoss` (MNRL with in-batch + hard negatives; | |
| the cached variant is used to simulate a large effective batch on limited VRAM). | |
| - **Batch sampler:** `NO_DUPLICATES` (required for MNRL to avoid trivial in-batch collisions). | |
| - **In-batch negatives:** every other positive/negative in the batch serves as a negative | |
| for a given anchor, so larger batches yield a stronger contrastive signal. | |
| ### Hyperparameter search | |
| An **Optuna** search (TPE sampler) was run **before** the final training: | |
| | Setting | Value | | |
| |--------|-------| | |
| | Trials | 10 | | |
| | Subset | 40,000 random training triplets | | |
| | Steps / trial | 600 | | |
| | Search space | `lr ∈ {1e-5, 2e-5, 3e-5}`, `batch ∈ {32, 64, 128}`, `warmup_ratio ∈ {0.0, 0.1}` | | |
| | Selection metric | `stsb-tr` **dev** `spearman_cosine` | | |
| | **Best config** | **lr = 2e-5, batch = 32, warmup_ratio = 0.1** (dev spearman = **0.8285**) | | |
| ### Final training | |
| | Setting | Value | | |
| |--------|-------| | |
| | Base model | `intfloat/multilingual-e5-base` | | |
| | Epochs | 1 | | |
| | Learning rate | 2e-5 | | |
| | Batch size | 32 (cached MNRL) | | |
| | Warmup ratio | 0.1 | | |
| | Max sequence length | 256 | | |
| | Precision | bf16 | | |
| | Seed | 42 | | |
| | Eval strategy | steps, monitoring `eval_stsb-tr-dev_spearman_cosine` | | |
| | Checkpoint selection | `load_best_model_at_end=True` (best dev spearman) | | |
| ### Frameworks & hardware | |
| - `sentence-transformers` 5.2.2, `transformers` 5.0.0, `torch` 2.11.0+cu128, | |
| `datasets` 4.4.1, Python 3.12 | |
| - 1× NVIDIA GeForce RTX 5050 Laptop GPU (8 GB), CUDA 12.8 | |
| --- | |
| ## Evaluation | |
| ### Baseline (base model) vs. Final (this model) | |
| Measured on the dataset's own test triplets (`TripletEvaluator`, cosine accuracy) and | |
| on the external Turkish STS set [`emrecan/stsb-mt-turkish`](https://huggingface.co/datasets/emrecan/stsb-mt-turkish) | |
| (`EmbeddingSimilarityEvaluator`, Spearman; scores normalised 0–5 → 0–1): | |
| | Metric | Baseline (e5-base) | Final | Δ | | |
| |--------|:---:|:---:|:---:| | |
| | stsb-tr **dev** spearman_cosine | 0.8092 | **0.8312** | +0.0220 | | |
| | stsb-tr **test** spearman_cosine | 0.7761 | **0.7814** | +0.0053 | | |
| | nli-tr **test** cosine_accuracy | 0.8968 | **0.9260** | +0.0292 | | |
| ### Methodology notes for TR-MTEB | |
| - Datasets pulled **directly from HuggingFace** (`trmteb/*`); metrics computed with a | |
| standalone script (no `mteb` package dependency), following standard MTEB protocols: | |
| - **STS** → Spearman of cosine vs. gold; **Retrieval** → nDCG@10 / Recall@10 / MAP@10 | |
| via exact cosine search; **Classification** → logistic regression on frozen | |
| embeddings (accuracy + macro-F1); **PairClassification** → average precision of cosine; | |
| **Clustering** → V-measure (MiniBatchKMeans, k = #labels); **BitextMining** → top-1 | |
| nearest-neighbour accuracy. | |
| - e5 prefixes applied throughout (`query:` for queries/symmetric sides, `passage:` for corpus). | |
| - Retrieval search runs on-GPU in fp16. | |
| - `75haber`, `thy_sa`, `irony-tr` have **no train split** in the `trmteb` org, so a seeded | |
| **70/30 stratified self-split** was used for those classification tasks (marked `note` in | |
| `trmteb_results.json`). All other classification sets use their official train/test. | |
| ### Category summary (macro averages) | |
| | Task type | # datasets | Avg. main metric | | |
| |-----------|:---:|:---:| | |
| | STS (spearman_cosine) | 1 | **0.791** | | |
| | BitextMining (accuracy) | 1 | **0.975** | | |
| | Classification (accuracy) | 6 | **0.817** | | |
| | PairClassification (AP) | 3 | **0.573** | | |
| | Clustering (v-measure) | 2 | **0.487** | | |
| | Retrieval (nDCG@10) | 10 | **0.407** | | |
| | **Overall (mean of category means)** | 23 | **0.675** | | |
| --- | |
| ## TR-MTEB Results (full) | |
| Evaluated on the [TR-MTEB](https://huggingface.co/trmteb) datasets | |
| (Baysan & Güngör, *TR-MTEB*, Findings of EMNLP 2025). | |
| | Task type | Dataset | Main metric | Value | | |
| |---|---|---|---:| | |
| | STS | stsb-tr | spearman_cosine | 0.7907 | | |
| | BitextMining | wmt16_en_tr | accuracy | 0.9753 | | |
| | Classification | news-cat | accuracy | 0.9600 | | |
| | Classification | 75haber ᵃ | accuracy | 0.8696 | | |
| | Classification | ts_timeline_news_category | accuracy | 0.7250 | | |
| | Classification | thy_sa ᵃ | accuracy | 0.8386 | | |
| | Classification | offenseval | accuracy | 0.8427 | | |
| | Classification | irony-tr ᵃ | accuracy | 0.6667 | | |
| | PairClassification | snli_tr | ap_cosine | 0.5280 | | |
| | PairClassification | multinli_tr | ap_cosine | 0.6324 | | |
| | PairClassification | xnli_tr | ap_cosine | 0.5591 | | |
| | Clustering | ts_abstract_corpus | v_measure | 0.5762 | | |
| | Clustering | 630koseyazisi | v_measure | 0.3978 | | |
| | Retrieval | tquad | ndcg@10 | 0.8222 | | |
| | Retrieval | quora-tr | ndcg@10 | 0.7323 | | |
| | Retrieval | squad-tr | ndcg@10 | 0.6584 | | |
| | Retrieval | scifact-tr | ndcg@10 | 0.5309 | | |
| | Retrieval | cqadupstack-gaming-tr | ndcg@10 | 0.3584 | | |
| | Retrieval | msmarco-tr | ndcg@10 | 0.2734 | | |
| | Retrieval | arguana-tr | ndcg@10 | 0.2583 | | |
| | Retrieval | nfcorpus-tr | ndcg@10 | 0.2272 | | |
| | Retrieval | fiqa-tr | ndcg@10 | 0.1237 | | |
| | Retrieval | scidocs-tr | ndcg@10 | 0.0887 | | |
| --- | |
| ## Limitations & Biases | |
| - **Translated training data.** The NLI triplets are machine-translated; translation | |
| artefacts and noise are inherited, which can cap fine-grained semantic precision. | |
| - **Short-text bias.** Trained on NLI-style short sentences at 256 tokens; long-document | |
| retrieval requires chunking and may underperform (see the lower `fiqa`/`scidocs` scores). | |
| - **Prefix sensitivity.** Using the model without `query:`/`passage:` prefixes degrades results. | |
| - **Domain gaps.** Scientific/financial retrieval (scidocs, fiqa) is weak; the model is | |
| strongest on general-domain QA-style retrieval (tquad, squad-tr, quora-tr). | |
| - **Societal bias.** The backbone and NLI data may encode social biases; validate before | |
| deployment in sensitive settings. | |
| - **Single epoch.** Trained for 1 epoch; longer schedules or larger effective batches may | |
| improve retrieval further. | |
| ## Compute & Environmental Impact | |
| - Single consumer laptop GPU (RTX 5050, 8 GB). Full pipeline (Optuna sweep + 1-epoch | |
| training on 277k triplets + evaluation) ran in a few GPU-hours. No multi-GPU/cluster | |
| training was used. | |
| ## Licensing | |
| - **Backbone** `intfloat/multilingual-e5-base`: MIT. | |
| - **Training data** derives from **AllNLI** (SNLI is **CC BY-SA 4.0**; MultiNLI has mixed | |
| source licenses), translated to Turkish. Because the most restrictive component is | |
| CC BY-SA 4.0, this model card is released under **CC BY-SA 4.0**; verify that this suits | |
| your use case and comply with the ShareAlike terms. | |
| - **Evaluation datasets** belong to their respective authors (see the `trmteb` org and | |
| `emrecan/stsb-mt-turkish`). | |
| ## Citation | |
| This model: | |
| ```bibtex | |
| @misc{e5-tr-nli, | |
| title = {e5-tr-nli: A Turkish Sentence Embedding Model}, | |
| note = {Fine-tuned from intfloat/multilingual-e5-base on Turkish NLI triplets}, | |
| year = {2026} | |
| } | |
| ``` | |
| TR-MTEB benchmark: | |
| ```bibtex | |
| @inproceedings{baysan-gungor-2025-trmteb, | |
| title = {{TR-MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations}, | |
| author = {Baysan, Mehmet Selman and G{\"u}ng{\"o}r, Tunga}, | |
| booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025}, | |
| year = {2025} | |
| } | |
| ``` | |
| Base model (E5): | |
| ```bibtex | |
| @article{wang2024multilingual, | |
| title = {Multilingual E5 Text Embeddings: A Technical Report}, | |
| author = {Wang, Liang and Yang, Nan and Huang, Xiaolong and Yang, Linjun and Majumder, Rangan and Wei, Furu}, | |
| journal = {arXiv preprint arXiv:2402.05672}, | |
| year = {2024} | |
| } | |
| ``` | |
| ## Reproduction | |
| The full pipeline (data prep, Optuna sweep, training, and TR-MTEB evaluation) is scripted: | |
| - `train_embedding_tr.py` — `--mode sweep | train | eval | mteb` | |
| - `eval_trmteb_hf.py` — TR-MTEB evaluation straight from HuggingFace `trmteb/*` datasets | |
| - `config.py`, `Makefile`, `requirements.txt` | |
| To publish (disabled by default — no secrets are used): | |
| ```python | |
| from sentence_transformers import SentenceTransformer | |
| SentenceTransformer("models/e5-tr-nli-final").push_to_hub("thealper2/intfloat-multilingual-e5-base-tr-nli") | |
| ``` | |