X-ASR zh-en streaming zipformer2 transducer — GGUF

GGUF conversions of the X-ASR zh-en streaming zipformer2 transducer (k2-fsa / sherpa-onnx export), for use with RapidSpeech.cpp (ggml backend, CPU + CUDA). Mandarin–English code-switching ASR with punctuation.

Converted with tools/convert_xasr_to_gguf.py. The encoder/decoder/joiner are fused into a single GGUF per chunk variant; streaming uses per-layer recurrent caches mirroring the ONNX state contract.

Variants

All four chunk variants share the same architecture (6 stacks / 19 layers, dims 192·256·512·768·512·256, vocab 5000); they differ only in the streaming chunk size (latency vs. accuracy trade-off).

Folder Chunk shift Encoder T Latency f16 Q4_K
160ms/ 16 frames 29 lowest ~292 MB ~85 MB
480ms/ 48 frames 61 low ~292 MB ~85 MB
960ms/ 96 frames 109 medium ~292 MB ~85 MB
1920ms/ 192 frames 205 highest accuracy ~292 MB ~86 MB

Each folder contains *-f16.gguf, *-q8_0.gguf, *-q4_k.gguf, *-q3_k.gguf (imatrix-calibrated), the imatrix-*.dat calibration file, and tokens.txt. The 960 ms folder additionally ships *-iq4_xs.gguf (a lossless 4-bit IQ build — the 960 ms variant was used for the full quant sweep below). Convolution kernels are always kept at f16 (quantizing them hurts accuracy).

Which weight format to use

Per-weight accuracy, measured on the 960 ms variant. Accuracy is the token edit-distance vs the f16 reference on a zh-en code-switch clip (0 = token-exact). Sizes are the actual GGUF bytes.

Format Size Edit-dist Published Notes
f16 307 MB 0 (ref) ✅ reference
q8_0 165 MB 0 — lossless ✅ best quality; ~1.2× faster than f16 on CPU
iq4_xs 87 MB 0 — lossless ✅ lossless 4-bit IQ (960 ms only)
q4_k 90 MB 3 ✅ near-lossless (minor casing: Monday→monday)
q3_k (imatrix) 72 MB 0 — lossless ✅ smallest lossless build

The q3_k.gguf files here are imatrix-calibrated (activation-aware, AWQ): an importance matrix collected over calibration audio protects the most important weight channels, recovering the accuracy 3-bit quantization normally loses (without it, q3_k scores edit-dist 4 — Monday→MD). Generate your own with xasr-dev-test imatrix + rs-quantize --imatrix.

Recommendation: q8_0 for lossless quality, or q3_k (imatrix) for the smallest lossless footprint (72 MB). Quantizing the matmul weights also speeds up ggml CPU inference (less memory traffic + tuned vec-dot kernels).

Sub-3-bit was evaluated but is not published

A full sweep below 3-bit was run on the 960 ms variant and deliberately excluded — none are useful:

Format Size Edit-dist Why excluded
q2_k (imatrix) 59 MB 3 degraded — below the 3-bit floor
iq2_s 58 MB 3–4 degraded
iq3_s 86 MB 4 dominated (bigger than q3_k-im and worse)
iq2_xxs (imatrix) 53 MB 6 degraded
iq1_s, iq2_xxs (no imatrix) 90 MB 3 fake — fell back to q4_k size, not real low-bit
iq1_m 45 MB 25–58 broken — garbage output

3-bit + imatrix is the accuracy floor. Below it, accuracy degrades (edit-dist 3–6) and 1-bit collapses entirely.

Parity

RapidSpeech.cpp (CPU, f16) is token-exact with sherpa-onnx (onnxruntime CPU, fp32) on the reference audio for all four variants. Q4_K matches to within occasional capitalization.

Example (10 s zh-en code-switching clip):

昨天是 Monday,today is 礼拜二,the day after tomorrow 是星期三

Benchmark (streaming, steady-state ms/chunk, warm-up excluded)

Measured on an NVIDIA GB10 host (the original Jetson Nano gen1 target was unavailable). RapidSpeech CUDA uses the FP32 non-tensor path (emulating the Nano's tensor-core-less sm_53). Numbers are relative, not Nano wall-clock.

Variant sherpa-onnx CPU RapidSpeech CPU RapidSpeech CUDA
160 ms 16.0 27.7 26.9
480 ms 23.1 48.0 38.4
960 ms 31.1 83.8 53.3
1920 ms 40.9 169.7 83.0

All configurations run faster than real time. CUDA's speedup over CPU grows with chunk size (1.0× → 2.0×) as larger GEMMs amortize per-chunk kernel-launch cost.

On-device Jetson Nano gen1 (sm_53) — measured

Measured on a real Jetson Nano gen1 (Tegra X1 / GM20B, sm_53, L4T R32.5.1, CUDA 10.2, MAXN, clocks unpinned → ~10% run-to-run noise), 960 ms variant, encoder ms/chunk (the fair cross-engine metric), 4 CPU threads. Full data and the CPU-thread / CUDA-core sweeps are in BENCHMARKS.md.

Engine / weights CPU (4 thr) CUDA (best) Correct?
sherpa-onnx fp32 396.9 — not runnable ✓
sherpa-onnx int8 329.8 — not runnable ✓
RapidSpeech f16 639.1 386.5 (RS_GEMM_FP16) ✓
RapidSpeech q8_0 459.2 355.3 (FP32) ✓
RapidSpeech q3_k-im 498.5 445.4 ✓
RapidSpeech iq4_xs 452.2 451.8 ✓
  • sherpa-onnx CUDA cannot run on the Nano (onnxruntime is CPU-only here; no aarch64 GPU wheel, and it would OOM 4 GB). RapidSpeech.cpp is the only way to use the Nano GPU for this model — that is the point of the port.
  • RS_GEMM_FP16=1 gives RapidSpeech CUDA a ~1.6–1.75× encoder speedup on f16 (sm_53 has native 2× FP16 throughput); it reaches parity with sherpa-onnx CPU while leaving the 4 A57 cores free. On already-quantized weights the lever is neutral — quantization and FP16 are substitute bandwidth levers.
  • sherpa-onnx CPU is the fastest engine (onnxruntime/MLAS: tuned ARM GEMM, op fusion, int8) and saturates at 3 threads; RapidSpeech CPU scales ~linearly to 4 threads but starts behind. CUDA latency is independent of CPU threads.
  • All correct except q4_k (lowercases "monday", drops a comma).

Usage

# RapidSpeech.cpp WebSocket streaming server
rs-xasr-ws-server -m 960ms/x-asr-zh-en-960ms-f16.gguf --port 6006

See RapidSpeech.cpp for build instructions (incl. the CUDA-10.2 / sm_53 Jetson Nano path).

License

Apache-2.0, following the upstream X-ASR model.

Downloads last month
110
GGUF
Model size
0.2B params
Architecture
XAsrZipformer2
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Luigi/x-asr-zh-en-streaming-zipformer2-gguf

Quantized
(3)
this model