license: apache-2.0
language:
- zh
- en
library_name: rapidspeech
tags:
- automatic-speech-recognition
- streaming
- zipformer2
- transducer
- ggml
- gguf
- code-switching
base_model:
- GilgameshWind/X-ASR-zh-en
pipeline_tag: automatic-speech-recognition
X-ASR zh-en streaming zipformer2 transducer — GGUF
GGUF conversions of the X-ASR zh-en streaming zipformer2 transducer (k2-fsa / sherpa-onnx export), for use with RapidSpeech.cpp (ggml backend, CPU + CUDA). Mandarin–English code-switching ASR with punctuation.
Converted with tools/convert_xasr_to_gguf.py. The encoder/decoder/joiner are
fused into a single GGUF per chunk variant; streaming uses per-layer recurrent
caches mirroring the ONNX state contract.
Variants
All four chunk variants share the same architecture (6 stacks / 19 layers, dims 192·256·512·768·512·256, vocab 5000); they differ only in the streaming chunk size (latency vs. accuracy trade-off).
| Folder | Chunk shift | Encoder T | Latency | f16 | Q4_K |
|---|---|---|---|---|---|
160ms/ |
16 frames | 29 | lowest | ~292 MB | ~85 MB |
480ms/ |
48 frames | 61 | low | ~292 MB | ~85 MB |
960ms/ |
96 frames | 109 | medium | ~292 MB | ~85 MB |
1920ms/ |
192 frames | 205 | highest accuracy | ~292 MB | ~86 MB |
Each folder contains *-f16.gguf, *-q8_0.gguf, *-q4_k.gguf, *-q3_k.gguf
(imatrix-calibrated), the imatrix-*.dat calibration file, and tokens.txt.
Convolution kernels are always kept at f16 (quantizing them hurts accuracy).
Which weight format to use
Measured on the 960 ms variant (RapidSpeech CPU); accuracy is token edit-distance vs the f16 reference on a zh-en code-switch clip (0 = token-exact):
| Format | Size | Accuracy (edit-dist) | Notes |
|---|---|---|---|
| q8_0 | 158 MB | 0 — lossless (all test clips) | best quality, ~1.2× faster than f16 on CPU |
| q4_k | 86 MB | 3 | near-lossless (minor casing) |
| q3_k (imatrix) | 69 MB | 0 on calib clip, ~1–2 tok on held-out | smaller and more accurate than q4_k |
| q3_k (no imatrix) | 69 MB | 4 | "Monday"→"MD"; use the imatrix build instead |
| q2_k / iq2 (imatrix) | 51–57 MB | 3–6 | degraded even with imatrix — q3_k is the floor |
The q3_k.gguf files here are imatrix-calibrated (activation-aware, AWQ):
an importance matrix collected over calibration audio protects the most important
weight channels, which recovers the accuracy 3-bit quantization normally loses.
Generate your own with xasr-dev-test imatrix + rs-quantize --imatrix.
Recommendation: q8_0 for lossless quality, or q3_k (imatrix) for the
smallest near-lossless footprint (69 MB). Quantizing the matmul weights also
speeds up ggml CPU inference (less memory traffic + tuned vec-dot kernels).
Below 3-bit, accuracy degrades even with an imatrix.
Parity
RapidSpeech.cpp (CPU, f16) is token-exact with sherpa-onnx (onnxruntime CPU, fp32) on the reference audio for all four variants. Q4_K matches to within occasional capitalization.
Example (10 s zh-en code-switching clip):
昨天是 Monday,today is 礼拜二,the day after tomorrow 是星期三
Benchmark (streaming, steady-state ms/chunk, warm-up excluded)
Measured on an NVIDIA GB10 host (the original Jetson Nano gen1 target was unavailable). RapidSpeech CUDA uses the FP32 non-tensor path (emulating the Nano's tensor-core-less sm_53). Numbers are relative, not Nano wall-clock.
| Variant | sherpa-onnx CPU | RapidSpeech CPU | RapidSpeech CUDA |
|---|---|---|---|
| 160 ms | 16.0 | 27.7 | 26.9 |
| 480 ms | 23.1 | 48.0 | 38.4 |
| 960 ms | 31.1 | 83.8 | 53.3 |
| 1920 ms | 40.9 | 169.7 | 83.0 |
All configurations run faster than real time. CUDA's speedup over CPU grows with chunk size (1.0× → 2.0×) as larger GEMMs amortize per-chunk kernel-launch cost.
Usage
# RapidSpeech.cpp WebSocket streaming server
rs-xasr-ws-server -m 960ms/x-asr-zh-en-960ms-f16.gguf --port 6006
See RapidSpeech.cpp for build instructions (incl. the CUDA-10.2 / sm_53 Jetson Nano path).
License
Apache-2.0, following the upstream X-ASR model.