h4rm0n1c commited on
Commit
5cdd044
·
verified ·
1 Parent(s): 033a0d4

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +72 -0
README.md ADDED
@@ -0,0 +1,72 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: gguf
4
+ base_model: sandeshrajx/Qwen3.5-24B-A3B-REAP-0.32
5
+ tags:
6
+ - qwen3.5
7
+ - moe
8
+ - gguf
9
+ - iq4_nl
10
+ - reap
11
+ - pruned
12
+ - quantzhai
13
+ - quantized
14
+ ---
15
+
16
+ # Qwen3.5-24B-A3B-REAP-0.32 IQ4_NL GGUF
17
+
18
+ GGUF quantization of [sandeshrajx/Qwen3.5-24B-A3B-REAP-0.32](https://huggingface.co/sandeshrajx/Qwen3.5-24B-A3B-REAP-0.32) — a REAP-pruned 24B A3B MoE model with **10B active parameters**.
19
+
20
+ **REAP** (Razor Edge And Pruning, [arxiv:2510.13999](https://arxiv.org/abs/2510.13999)) is a structured pruning technique that reduces the base Qwen3.5 model while preserving capability.
21
+
22
+ **Source:** [sandeshrajx/Qwen3.5-24B-A3B-REAP-0.32](https://huggingface.co/sandeshrajx/Qwen3.5-24B-A3B-REAP-0.32)
23
+ **Converted by:** [QuantZhai](https://github.com/h4rm0n1c/quantzhai) benchmark pipeline
24
+ **Quantization:** IQ4_NL with importance matrix (imatrix from sandeshrajx)
25
+
26
+ ## Model Details
27
+
28
+ | Property | Value |
29
+ |---|---|
30
+ | Architecture | Qwen3.5 MoE, REAP-pruned |
31
+ | Parameters | 24B total, ~10B active |
32
+ | Blocks | 40 |
33
+ | Experts | 175 (REAP-split), 8 active per token |
34
+ | Context length | 262144 (256K) |
35
+ | Hidden size | 3072 |
36
+ | Attention heads | 32, KV heads = 2 |
37
+ | Quantization | IQ4_NL (4.58 bpw) with imatrix |
38
+ | File size | 14.0 GB |
39
+
40
+ ## Benchmarks
41
+
42
+ Hardware: dual-GPU (RTX 3080 10GB + V100-SXM2 32GB)
43
+ Engine: llama.cpp with TurboQuant KV (q8_0 K / turbo3 V)
44
+ Perplexity: [`macvox68`](https://github.com/h4rm0n1c/macvox68) code corpus, ctx=4096, stride=512
45
+
46
+ | Metric | Cold | Warm |
47
+ |---|---|---|
48
+ | PPL | 3.0205 | **1.8298** |
49
+ | TPS | 33.6 tok/s | **45.9 tok/s** |
50
+ | TTFT | 1486 ms | 1090 ms |
51
+
52
+ ### QuantZhai Ranking
53
+
54
+ **Rank #15 of 47** — combined score 63.6 (equal-weight: TPS, PPL, convergence).
55
+ Higher than many larger dense models — REAP pruning + IQ4_NL quantization is an efficient combination.
56
+
57
+ ## Usage
58
+
59
+ ```bash
60
+ llama-cli -m Qwen3.5-24B-A3B-REAP-0.32-IQ4_NL.gguf \
61
+ -p "Write a mergesort in Python" \
62
+ -n 1024 -t 12 --temp 0.6 --top-p 0.95
63
+
64
+ llama-server -m Qwen3.5-24B-A3B-REAP-0.32-IQ4_NL.gguf \
65
+ --host 0.0.0.0 --port 8080 -ngl 99 -t 12 \
66
+ --cache-type-k q8_0 --cache-type-v turbo3
67
+ ```
68
+
69
+ ## License
70
+
71
+ Apache 2.0 (this quantization).
72
+ Source model by [sandeshrajx](https://huggingface.co/sandeshrajx) under Apache 2.0 — see [arxiv:2510.13999](https://arxiv.org/abs/2510.13999).