nllb_200_distilled_600m-zul-eng-onnx

ONNX export of dsfsi/nllb_200_distilled_600m-zul-eng.

Translation direction: Zulu (zu) -> English (en).

Attribution

The original model was trained by DSFSI (Data Science for Social Impact research group, University of Pretoria), a fine-tune of facebook/nllb-200-distilled-600M on the afriscience-mt dataset. Licence: Apache-2.0, as declared by the original repository. This repository only converts the weights to ONNX. All credit for the model belongs to DSFSI.

Export

optimum-cli export onnx --model dsfsi/nllb_200_distilled_600m-zul-eng \
  --task text2text-generation-with-past --no-post-process <outdir>

The int8 quantized files needed the .onnx.data -> .onnx_data rename and proto location rewrite (encoder, decoder and decoder-with-past all triggered external data this time, including one Constant node tensor attribute per file that is easy to miss since it is not a graph initializer). Verified with a real load from a cache cleared beforehand.

Files

Path Precision Size
*.onnx + *.onnx_data (root) fp32 ~7 GB
int8/*.onnx + int8/*.onnx_data int8 dynamic ~5.5 GB

Parity

10 general-domain Zulu sentences (source language for this zu->en pair), PyTorch original vs ONNX, exact string match against AutoModelForSeq2SeqLM.generate().

parity:
  sample_size: 10
  metric: exact_match
  fp32_greedy: 1.00   # 10/10
  fp32_beam4:  1.00   # 10/10
  int8_greedy: 1.00   # 10/10
  int8_beam4:  0.80   # 8/10
Decoding fp32 (n=10) int8 (n=10)
greedy 100% (10/10) 100% (10/10)
beam=4 100% (10/10) 80% (8/10)

fp32 is an exact reproduction of the original model. int8 disagreements were inspected by hand and are paraphrases (e.g. ref "It has only continued its work for three years." vs int8 "She has continued her work for three years." - a pronoun/subject choice difference, both plausible readings of the Zulu source), not degradations.

Caveat on sample size: n=10 exact-match is a small, noisy estimate, not a FLORES-scale benchmark. A separate measurement campaign on this project found n=5 exact-match numbers for another model swing from 100% to 14% when re-measured at n=100 on FLORES-200. Treat the percentages above as an initial signal only - a proper FLORES + chrF evaluation supersedes it wherever available.

Selecting the language

This is an NLLB checkpoint, not M2M100: the target language is not baked into forced_bos_token_id in generation_config.json - you must pass it explicitly at generate time using NLLB's flores_code language tags (zul_Latn, eng_Latn, not plain ISO 639-1 codes):

tokenizer.src_lang = "zul_Latn"
forced_bos_token_id = tokenizer.convert_tokens_to_ids("eng_Latn")
model.generate(**inputs, forced_bos_token_id=forced_bos_token_id, ...)

Usage

from transformers import AutoTokenizer
from optimum.onnxruntime import ORTModelForSeq2SeqLM

repo = "TigreGotico/nllb_200_distilled_600m-zul-eng-onnx"
tok = AutoTokenizer.from_pretrained(repo)
tok.src_lang = "zul_Latn"
model = ORTModelForSeq2SeqLM.from_pretrained(repo, use_cache=True, use_merged=False)

enc = tok("Isimo sezulu sihle kakhulu namuhla.", return_tensors="pt")
tgt_id = tok.convert_tokens_to_ids("eng_Latn")
out = model.generate(**enc, forced_bos_token_id=tgt_id, num_beams=4, max_new_tokens=64)
print(tok.batch_decode(out, skip_special_tokens=True)[0])
# -> The weather is very good today.

For the int8 build, pass subfolder="int8".

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TigreGotico/nllb_200_distilled_600m-zul-eng-onnx

Quantized
(1)
this model