Instructions to use Luigi/nemotron-asr-litert-zhtw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use Luigi/nemotron-asr-litert-zhtw with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Nemotron-3.5-ASR β LiteRT q4-mix, zh-TW fine-tuned (v2)
On-device (Android) LiteRT build of Luigi/nemotron-3.5-asr-streaming-0.6b-zhtw β the zh-TW fine-tuned Nemotron (3.2Γ better Taiwan Mandarin than base). q4-mix: INT4 encoder + fp16 decoder/joint + fp32 prompt-fusion. Bundle 663 MB.
Supersedes v1 (which packaged the zero-training warm-start).
Accuracy
| build | CV zh-TW CER (n=120) |
|---|---|
| fp32 fine-tuned model (reference) | 11.33 |
| this q4-mix LiteRT | 13.20 (+1.9 INT4 cost) |
| v1 q4-mix (base warm-start) | ~38 |
Per language-prompt slot on this build β all within ~0.7 CER, pick either:
| slot | q4-mix CER | fp32 CER |
|---|---|---|
zh-CN (4) |
13.20 | 11.57 |
auto (101) |
13.79 | 11.33 |
zh-TW (5) |
13.90 | 12.03 |
auto was the worst slot on the base model (50.58 CER) and is repaired by the
fine-tune β training used prompt_mode: unified, which trains the auto path
alongside the explicit language ID. Use auto when the speaker may switch
languages, or an explicit slot when you know it.
Other languages (fp32 FT, vs base): ko/de/ja/hi/en all improve, ar flat, fr/es +1.5 β see the source model card.
Files
| file | role | precision | size |
|---|---|---|---|
nemotron_encoder_q4.tflite |
FastConformer encoder | INT4 FC (blockwise-128) + fp32 convs | 596 MB |
nemotron_prompt_fuse_fp32.tflite |
128-slot language-prompt fusion | fp32 | 18 MB |
nemotron_decoder_fp16.tflite |
RNN-T prediction net | fp16 | 31 MB |
nemotron_joint_fp16.tflite |
RNN-T joint (β vocab 13088) | fp16 | 18 MB |
tokenizer.json / processor_config.json / config.json |
runtime | β | β |
Run
Port, runner and integration note: https://github.com/vieenrose/LiteRT/tree/nemotron/litert/samples/asr/nemotron
python -m nemotron.runner --models ./ --wav clip.wav --lang zh-CN --s2t --itn
# --lang auto also works well (see slot table); --s2t gives Traditional output
Output is Simplified Chinese β the tokenizer cannot represent many common
Traditional characters, so apply OpenCC s2t (--s2t) for Traditional.
The INT4 encoder needs the LiteRT-Next CompiledModel runtime (Android NNAPI /
XNNPACK-QD8); the classic Interpreter cannot allocate dynamic-range INT4.
Set the CPU thread count explicitly (big-core count on big.LITTLE).
Provenance
Base: nvidia/nemotron-3.5-asr-streaming-0.6b (NVIDIA Open Model License),
fine-tuned for zh-TW (see source model card), then exported with
litert_torch 0.9.1 + ai_edge_quantizer 0.7.0 (INT4 weights pre-baked to the
export grid before quantization).
- Downloads last month
- 119
Model tree for Luigi/nemotron-asr-litert-zhtw
Base model
nvidia/nemotron-3.5-asr-streaming-0.6b