nllb_200_distilled_600m-eng-yor-onnx
ONNX export of dsfsi/nllb_200_distilled_600m-eng-yor.
Translation direction: English (en) -> Yoruba (yo).
Attribution
The original model was trained by DSFSI (Data
Science for Social Impact research group, University of Pretoria), a
fine-tune of facebook/nllb-200-distilled-600M on the afriscience-mt
dataset. Licence: Apache-2.0, as declared by the original repository.
This repository only converts the weights to ONNX. All credit for the
model belongs to DSFSI.
Export
optimum-cli export onnx --model dsfsi/nllb_200_distilled_600m-eng-yor \
--task text2text-generation-with-past --no-post-process <outdir>
The int8 quantized files needed the .onnx.data -> .onnx_data rename and
proto location rewrite (encoder, decoder and decoder-with-past, plus one
Constant node tensor attribute per file). Verified with a real load from
a cache cleared beforehand.
Files
| Path | Precision | Size |
|---|---|---|
*.onnx + *.onnx_data (root) |
fp32 | ~7 GB |
int8/*.onnx + int8/*.onnx_data |
int8 dynamic | ~5.5 GB |
Parity
10 general-domain English sentences, PyTorch original vs ONNX, exact string
match against AutoModelForSeq2SeqLM.generate().
parity:
sample_size: 10
metric: exact_match
fp32_greedy: 1.00 # 10/10
fp32_beam4: 1.00 # 10/10
int8_greedy: 0.90 # 9/10
int8_beam4: 0.80 # 8/10
| Decoding | fp32 (n=10) | int8 (n=10) |
|---|---|---|
| greedy | 100% (10/10) | 90% (9/10) |
| beam=4 | 100% (10/10) | 80% (8/10) |
fp32 is an exact reproduction of the original model. int8 disagreements were inspected by hand and are minor word-choice/phrasing differences (e.g. ref "Ọ̀pọ̀lọpọ̀ ọ̀pẹ́ fún ìrànlọ́wọ́ ọ̀wọ́ rẹ̀." vs int8 "...fún ìrànlọ́wọ́ rẹ̀." - dropping one honorific-register word, still grammatical and understandable), not degradations.
Caveat on sample size: n=10 exact-match is a small, noisy estimate, not a FLORES-scale benchmark. A separate measurement campaign on this project found n=5 exact-match numbers for another model swing from 100% to 14% when re-measured at n=100 on FLORES-200. Treat the percentages above as an initial signal only - a proper FLORES + chrF evaluation supersedes it wherever available.
Selecting the language
This is an NLLB checkpoint, not M2M100: the target language is not
baked into forced_bos_token_id in generation_config.json - you must
pass it explicitly at generate time using NLLB's flores_code language
tags (eng_Latn, yor_Latn, not plain ISO 639-1 codes):
tokenizer.src_lang = "eng_Latn"
forced_bos_token_id = tokenizer.convert_tokens_to_ids("yor_Latn")
model.generate(**inputs, forced_bos_token_id=forced_bos_token_id, ...)
Usage
from transformers import AutoTokenizer
from optimum.onnxruntime import ORTModelForSeq2SeqLM
repo = "TigreGotico/nllb_200_distilled_600m-eng-yor-onnx"
tok = AutoTokenizer.from_pretrained(repo)
tok.src_lang = "eng_Latn"
model = ORTModelForSeq2SeqLM.from_pretrained(repo, use_cache=True, use_merged=False)
enc = tok("The weather is very nice today.", return_tensors="pt")
tgt_id = tok.convert_tokens_to_ids("yor_Latn")
out = model.generate(**enc, forced_bos_token_id=tgt_id, num_beams=4, max_new_tokens=64)
print(tok.batch_decode(out, skip_special_tokens=True)[0])
# -> Ojú-ọjọ́ dára gan-an lọ́jọ́ òní.
For the int8 build, pass subfolder="int8".
- Downloads last month
- 9
Model tree for TigreGotico/nllb_200_distilled_600m-eng-yor-onnx
Base model
facebook/nllb-200-distilled-600M