Qwen3-4B W8A8 for RK3588 (Orange Pi 5)

Quantized Qwen/Qwen3-4B models for Rockchip RK3588 NPU using RKLLM.

Files

File hybrid_rate Description Size
Qwen3-4B-w8a8-npu.rkllm 0.0 All layers on NPU โ€” fastest throughput 4.51 GB
Qwen3-4B-w8a8-hybrid.rkllm 0.5 50% CPU A76 + 50% NPU โ€” lower NPU memory pressure 4.54 GB

Quantization Details

  • Source model: Qwen/Qwen3-4B (HuggingFace)
  • Toolkit: rkllm-toolkit 1.2.1b1
  • dtype: W8A8 (8-bit weights + 8-bit activations)
  • Algorithm: normal
  • Platform: rk3588, num_npu_core=3
  • Max context: 4096 tokens
  • Calibration: 20 representative prompts (reasoning, math, code, multilingual)

Note: W4A16 is not supported by rkllm-toolkit 1.2.1b1 for RK3588.

Usage

See the full deployment stack at: https://github.com/kamyarkazremi/orangepi5-rkllm

Includes:

  • rkllm_enhanced binary with n_keep=4 sliding-window KV cache
  • Patched API server with ChatML support and zombie recovery
  • Systemd service configuration
Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support