rene98c commited on
Commit
bf2add1
·
verified ·
1 Parent(s): 4878785

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +38 -0
README.md ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Qwen3.5-397B-A17B — REAP 28% Pruned, NVFP4
2
+
3
+ A personal experiment in aggressive MoE pruning. The goal: fit Qwen3.5-397B on **2× 96GB Blackwell GPUs** with usable KV cache (~90K tokens), without losing quality.
4
+
5
+ ## What this is
6
+
7
+ 28% of experts removed using [REAP](https://arxiv.org/abs/2501.02348) (Routing-Expert Activation Pruning) with a saliency × activation-count ordering, then quantized to NVFP4 using llm-compressor. Final size: **~164GB**.
8
+
9
+ Pruning is **heterogeneous** — each layer retains a different number of experts based on global importance ranking. Early layers (which carry more redundancy) are pruned more aggressively; late layers are barely touched. This maximizes quality retention for a given size budget.
10
+
11
+ ## Benchmark results (non-thinking, lm-eval-harness)
12
+
13
+ | Benchmark | This model | Nvidia NVFP4 (full, ~240GB) |
14
+ |-----------|----------:|----------:|
15
+ | IFEval | **92.55** | 91.20 |
16
+ | MMLU Redux (generative) | 90.94 | **91.24** |
17
+ | GSM8K CoT (Llama) | 96.74 | **96.80** |
18
+
19
+ 28% fewer experts, 30%+ smaller on disk, and benchmark scores within noise of the full model.
20
+
21
+ ## Requirements
22
+
23
+ - **vLLM** ≥ 0.16.1 (nightly dev builds from the cu130 index work)
24
+ - **Transformers** ≥ 5.3
25
+
26
+ ## vLLM patches required
27
+
28
+ This model uses variable expert counts per layer (not a fixed number), which stock vLLM doesn't support yet. Two files need patching:
29
+
30
+ 1. **`qwen3_next.py`** — read per-layer expert count instead of assuming a single global value
31
+ 2. **`qwen3_5.py`** — infer expert count from tensor shape during weight loading
32
+
33
+ ## Details
34
+
35
+ - **Base model:** Qwen3.5-397B-A17B
36
+ - **Pruning:** REAP with 1,188 curated calibration samples, saliency × count global ordering
37
+ - **Quantization:** NVFP4 (FP4 E2M1, group size 16, duo scaling)
38
+ - **Target hardware:** 2× RTX PRO 6000