Luigi's picture
Document q3_k (imatrix) near-lossless 69MB option
75d7a97 verified
|
Raw
History Blame
4.44 kB
metadata
license: apache-2.0
language:
  - zh
  - en
library_name: rapidspeech
tags:
  - automatic-speech-recognition
  - streaming
  - zipformer2
  - transducer
  - ggml
  - gguf
  - code-switching
base_model:
  - GilgameshWind/X-ASR-zh-en
pipeline_tag: automatic-speech-recognition

X-ASR zh-en streaming zipformer2 transducer — GGUF

GGUF conversions of the X-ASR zh-en streaming zipformer2 transducer (k2-fsa / sherpa-onnx export), for use with RapidSpeech.cpp (ggml backend, CPU + CUDA). Mandarin–English code-switching ASR with punctuation.

Converted with tools/convert_xasr_to_gguf.py. The encoder/decoder/joiner are fused into a single GGUF per chunk variant; streaming uses per-layer recurrent caches mirroring the ONNX state contract.

Variants

All four chunk variants share the same architecture (6 stacks / 19 layers, dims 192·256·512·768·512·256, vocab 5000); they differ only in the streaming chunk size (latency vs. accuracy trade-off).

Folder Chunk shift Encoder T Latency f16 Q4_K
160ms/ 16 frames 29 lowest ~292 MB ~85 MB
480ms/ 48 frames 61 low ~292 MB ~85 MB
960ms/ 96 frames 109 medium ~292 MB ~85 MB
1920ms/ 192 frames 205 highest accuracy ~292 MB ~86 MB

Each folder contains *-f16.gguf, *-q8_0.gguf, *-q4_k.gguf, *-q3_k.gguf (imatrix-calibrated), the imatrix-*.dat calibration file, and tokens.txt. Convolution kernels are always kept at f16 (quantizing them hurts accuracy).

Which weight format to use

Measured on the 960 ms variant (RapidSpeech CPU); accuracy is token edit-distance vs the f16 reference on a zh-en code-switch clip (0 = token-exact):

Format Size Accuracy (edit-dist) Notes
q8_0 158 MB 0 — lossless (all test clips) best quality, ~1.2× faster than f16 on CPU
q4_k 86 MB 3 near-lossless (minor casing)
q3_k (imatrix) 69 MB 0 on calib clip, ~1–2 tok on held-out smaller and more accurate than q4_k
q3_k (no imatrix) 69 MB 4 "Monday"→"MD"; use the imatrix build instead
q2_k / iq2 (imatrix) 51–57 MB 3–6 degraded even with imatrix — q3_k is the floor

The q3_k.gguf files here are imatrix-calibrated (activation-aware, AWQ): an importance matrix collected over calibration audio protects the most important weight channels, which recovers the accuracy 3-bit quantization normally loses. Generate your own with xasr-dev-test imatrix + rs-quantize --imatrix.

Recommendation: q8_0 for lossless quality, or q3_k (imatrix) for the smallest near-lossless footprint (69 MB). Quantizing the matmul weights also speeds up ggml CPU inference (less memory traffic + tuned vec-dot kernels). Below 3-bit, accuracy degrades even with an imatrix.

Parity

RapidSpeech.cpp (CPU, f16) is token-exact with sherpa-onnx (onnxruntime CPU, fp32) on the reference audio for all four variants. Q4_K matches to within occasional capitalization.

Example (10 s zh-en code-switching clip):

昨天是 Monday,today is 礼拜二,the day after tomorrow 是星期三

Benchmark (streaming, steady-state ms/chunk, warm-up excluded)

Measured on an NVIDIA GB10 host (the original Jetson Nano gen1 target was unavailable). RapidSpeech CUDA uses the FP32 non-tensor path (emulating the Nano's tensor-core-less sm_53). Numbers are relative, not Nano wall-clock.

Variant sherpa-onnx CPU RapidSpeech CPU RapidSpeech CUDA
160 ms 16.0 27.7 26.9
480 ms 23.1 48.0 38.4
960 ms 31.1 83.8 53.3
1920 ms 40.9 169.7 83.0

All configurations run faster than real time. CUDA's speedup over CPU grows with chunk size (1.0× → 2.0×) as larger GEMMs amortize per-chunk kernel-launch cost.

Usage

# RapidSpeech.cpp WebSocket streaming server
rs-xasr-ws-server -m 960ms/x-asr-zh-en-960ms-f16.gguf --port 6006

See RapidSpeech.cpp for build instructions (incl. the CUDA-10.2 / sm_53 Jetson Nano path).

License

Apache-2.0, following the upstream X-ASR model.