Qwen3.8-Flash-Next AutoRound A100 3 bpw + MTP

This is a public, mixed-precision AutoRound checkpoint of Qwen/Qwen3.8-Flash-Next, pinned to revision de4b8e4d43b917e7706784d8bb445c9af86a3540. It targets one NVIDIA A100 (SM80) with a 64 GiB GPU-resident weight budget.

Qwen3.8-Flash-Next is an experimental preview model. Runtime support is also experimental; read the compatibility section before downloading.

Quantization

  • 48 routed-expert banks: depth-stratified 2.998 effective bpw (scale/zero-point overhead included): W3A16G128 on layers 0-11 and 36-47, W2A16G64 on layers 12-35.
  • Backbone QSA/GDN major linear projections: W8A16G128.
  • Routers, shared experts, hyperconnection/control paths, PLE projections, vision tower, embeddings, and LM head: BF16.
  • MTP experts: W4A16G128 symmetric RTN; MTP QSA and dense projections: W8A16G128 symmetric RTN.
  • Backbone packing: symmetric RTN through AutoRound's shard-streamed model-free path (no calibration dataset).
  • Packing: native AutoRound auto_round:auto_gptq mixed-bit format.

The module-level precision metadata contains: {"16": 858, "3": 24, "4": 1536, "8": 175}.

Measured storage topology

  • GPU-resident weight tensors: 48.52 GiB.
  • Host-offloaded PLE n-gram table: 95.37 GiB (128 tensors).
  • MTP tensors are included in the checkpoint and counted in the GPU-resident figure.
  • Safetensor shards: 133; indexed tensors: 227702.

The 64 GiB figure is a weight budget, not a claim that every context length fits. KV cache and runtime workspaces require additional HBM. The 95 GiB n-gram table must remain in host memory; provision ample system RAM.

Runtime compatibility

At publication time, stock vLLM does not yet contain both required changes. Use a build combining:

Serve with tensor parallel size 1. Keep PLE CPU offload enabled and MTP GPU-resident. Set the runtime's GPU memory limit according to the desired KV-cache headroom; do not assume that a 64 GiB weight fit implies a 64 GiB total process fit.

Validation

The official source checkpoint passed a complete 131-shard/index inventory check before quantization. This repository passed a second safetensor/index scan, mixed-bit metadata checks, MTP presence checks, source-code exclusion, and the measured 64 GiB resident-weight gate. See source_validation.json, validation_report.json, build_info.json, and mtp_pack_plan.json for machine-readable details.

Quantization changes model outputs and may reduce quality, especially on workloads sensitive to 2-bit expert banks. No accuracy benchmark is claimed here.

License

Apache-2.0, inherited from the base model. Follow the base model card and license for acceptable use and limitations.

Downloads last month
-
Safetensors
Model size
66B params
Tensor type
BF16
·
I32
·
F16
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for klee100/Qwen3.8-Flash-Next-AutoRound-A100-3bpw-MTP

Quantized
(85)
this model