opus-mt-en-az-onnx / README.md
Jarbas's picture
Upload README.md with huggingface_hub
4257115 verified
|
Raw
History Blame Contribute Delete
2.53 kB
metadata
license: apache-2.0
base_model: Helsinki-NLP/opus-mt-en-az
base_model_relation: quantized
tags:
  - onnx
  - translation
  - marian
library_name: transformers
pipeline_tag: translation
language:
  - en
  - az

opus-mt-en-az-onnx

ONNX export (fp32 + dynamic int8) of Helsinki-NLP/opus-mt-en-az, a Marian (en -> az) translation model from the Helsinki-NLP OPUS-MT project.

License: apache-2.0, inherited unchanged from the base model.

Contents

Files What it is
encoder_model.onnx, decoder_model.onnx, decoder_with_past_model.onnx ONNX, float32
int8/ same three graphs, dynamic int8 (QInt8, MatMul only, /lm_head/MatMul excluded)
source.spm, target.spm, vocab.json MarianTokenizer over the original sentencepiece models

fp32 size: 485 MB (three graphs) | int8 size: 295 MB

Usage

from optimum.onnxruntime import ORTModelForSeq2SeqLM
from transformers import AutoTokenizer

repo = "TigreGotico/opus-mt-en-az-onnx"
tok = AutoTokenizer.from_pretrained(repo)
model = ORTModelForSeq2SeqLM.from_pretrained(repo, use_cache=True, use_merged=False)  # fp32
# int8: ORTModelForSeq2SeqLM.from_pretrained(repo, subfolder="int8", use_cache=True, use_merged=False)
inputs = tok("The weather is very nice today.", return_tensors="pt")
out = model.generate(**inputs, num_beams=4, max_new_tokens=64)
print(tok.decode(out[0], skip_special_tokens=True))

Parity with the original PyTorch model

10 general-domain sentences, exact-string-match of generated output against MarianMTModel.generate() on the original Helsinki-NLP/opus-mt-en-az checkpoint.

parity:
  fp32_greedy: 1.00   # 10/10
  fp32_beam4:  1.00   # 10/10
  int8_greedy: 0.90   # 9/10
  int8_beam4:  0.40   # 4/10
Decoding fp32 exact match int8 exact match
greedy (num_beams=1) 10/10 (100.0%) 9/10 (90.0%)
beam=4 10/10 (100.0%) 4/10 (40.0%)

fp32 is a faithful reproduction of the original model at both decoding settings. int8 dynamic quantization noticeably degrades quality on this checkpoint under beam search - a spot check found real semantic drift on longer sentences (not just paraphrase), e.g. for "We need to discuss the budget for next quarter." the int8 beam-4 output diverged from both the reference and the fp32 ONNX output. Prefer fp32 for this pair; int8 is provided for size-constrained deployments where greedy decoding is used and some quality loss is acceptable.