Instructions to use andreasmartin/apertus-v1.1-swiss-embed-0.4b-bidir-langmoe with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use andreasmartin/apertus-v1.1-swiss-embed-0.4b-bidir-langmoe with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("andreasmartin/apertus-v1.1-swiss-embed-0.4b-bidir-langmoe") sentences = [ "Das ist eine glückliche Person", "Das ist ein glücklicher Hund", "Das ist eine sehr glückliche Person", "Heute ist ein sonniger Tag" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
library_name: sentence-transformers
pipeline_tag: sentence-similarity
license: apache-2.0
base_model: andreasmartin/apertus-v1.1-swiss-embed-0.4b-bidir
datasets:
- wikimedia/wikipedia
- eljuanina/VotingBooklets-v1
- ZurichNLP/SwissGov-RSD
language:
- de
- en
- fr
- it
- rm
- gsw
tags:
- sentence-transformers
- embeddings
- retrieval
- multilingual
- swiss
- apertus
- apertus-1.1
- bidirectional
- matryoshka
- moe
- language-moe
- sparse-routing
Apertus v1.1 Swiss Embed — Language-MoE
A standalone Sentence Transformers retrieval model based on
andreasmartin/apertus-v1.1-swiss-embed-0.4b-bidir.
The model retains the trained dense bidirectional Apertus encoder backbone and replaces the parent's single dense embedding projection with a learned Language-MoE projection head.
This release is a dense-backbone + sparse Language-MoE head, not a backbone-level native MoE model.
Architecture
plain retrieval text
↓
bidirectional Apertus encoder
↓
mean pooling
↓
Language-MoE router
↓
Top-2 of:
DE / EN / FR / IT / RM / GSW / shared
↓
1024d projection
↓
L2 normalization
| Property | Value |
|---|---|
| Total parameters | 445,696,559 (0.446B) |
| Active parameters / sentence | 440,453,679 |
| Language-MoE head parameters | 7,347,207 |
| Active head parameters / sentence | 2,104,327 |
| Router parameters | 7,175 |
| Parameters / expert | 1,048,576 |
| Experts | 7 |
| Active experts / sentence | 2 |
| Native embedding dimension | 1024 |
| Matryoshka dimensions | [1024, 768, 512, 256] |
| MoE training max length | 512 tokens |
| Inference max length | 1024 tokens |
| Sparse/reference max abs diff | 1.431e-06 |
Sentence Transformers usage
LanguageMoE is a custom Sentence Transformers module, so load the trusted repository with
trust_remote_code=True:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"andreasmartin/apertus-v1.1-swiss-embed-0.4b-bidir-langmoe",
trust_remote_code=True,
)
queries = model.encode_query([
"Welche Aufgaben hat der Bundesrat?"
])
documents = model.encode_document([
"Le Conseil fédéral est l'autorité exécutive suprême de la Confédération suisse.",
"Der Nationalrat ist die grosse Kammer der Bundesversammlung.",
])
scores = model.similarity(queries, documents)
print(scores)
For a compact Matryoshka representation:
embeddings = model.encode_query(
texts,
truncate_dim=512,
)
The stored retrieval prefixes are:
- query:
query: - document:
passage:
No chat template is used.
Lineage
| Property | Value |
|---|---|
| Original source | swiss-ai/Apertus-v1.1-0.5B |
| Dense embedding parent | andreasmartin/apertus-v1.1-swiss-embed-0.4b-bidir |
| Dense parent revision | 875c1f99d7b20e261ea43b3586fdde5f8d5648e5 |
| Apertus lineage | swiss-ai/Apertus-8B-2509 |
The repository contains the complete transformer/tokenizer/pooling/MoE/normalization pipeline. The dense parent is not fetched at inference time.
Language-MoE training
The dense embedding backbone was frozen. Its pooled hidden states were cached once and the router + expert projections were trained for the controlled MoE-head ablation.
See training_metadata.json for the exact cache/training history and
router_diagnostics.json for held-out routing counts.
Training data:
eljuanina/VotingBooklets-Diamond-v1 is held-out evaluation-only and is therefore not listed in the
training dataset metadata.
Actual training set: 31,483 triplets.
The MoE stage consumes the shared train_moe.jsonl generated together with the Dense
train.jsonl; triplets and negatives are identical. The released data_manifest.json records
the exact Hugging Face dataset commit SHAs, recipe signature and SHA256 hashes of the processed
files.
Internal retrieval diagnostics
These metrics are computed by reloading the saved standalone Sentence Transformers model. They are internal development diagnostics, not MTEB/MMTEB benchmark claims.
| Diagnostic | Dim | Accuracy@1 | Recall@10 | nDCG@10 | MRR@10 |
|---|---|---|---|---|---|
| Swiss monolingual | 1024 | 82.08% | 91.88% | 0.8678 | 0.8517 |
| Swiss monolingual | 512 | 80.00% | 91.88% | 0.8593 | 0.8404 |
| Swiss cross-lingual | 1024 | 30.63% | 96.04% | 0.6627 | 0.5635 |
| Swiss cross-lingual | 512 | 30.42% | 96.04% | 0.6602 | 0.5605 |
Monolingual evaluation queries: 480
Cross-lingual evaluation queries: 480
Router diagnostics
The held-out routing diagnostics are saved in router_diagnostics.json.
Recorded shared-expert Top-k selections in the held-out diagnostic: 0.
The declared language is used only as weak auxiliary supervision during MoE-head training; no language ID is required at inference.
Limitations
- Sparse routing is currently in the embedding projection head; the Apertus backbone remains dense.
- Wikipedia title→paragraph pairs are pseudo-retrieval supervision.
- Parallel passages provide semantic alignment rather than natural search queries.
- External MTEB/MMTEB/MIRACL and Swiss/domain-specific evaluation is required for comparative claims.
- The MoE head was trained on pooled representations up to 512 tokens. The packaged parent retains 1024-token inference capability, but quality beyond the MoE training context length has not yet been separately validated.
License and attribution
Apertus is developed by the Swiss AI Initiative. This is an independent retrieval/MoE adaptation and not an official Swiss AI Initiative embedding release.