barozp commited on
Commit
0d150da
·
verified ·
1 Parent(s): c57c5eb

Add model card with quant list and measured speculative-decoding benchmarks

Browse files
Files changed (1) hide show
  1. README.md +59 -0
README.md ADDED
@@ -0,0 +1,59 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model_relation: quantized
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - gguf
7
+ - llama.cpp
8
+ - quantized
9
+ - moe
10
+ - qwen3
11
+ - reasoning
12
+ ---
13
+
14
+ # Qwen3.6-29B-REAP-Opus-Reasoning-Distill-GGUF
15
+
16
+ GGUF quantizations of a REAP-pruned (205/256 experts) Qwen3.6-35B-A3B MoE, merged with an Opus-reasoning-distilled LoRA adapter. This is the **plain merge — no Multi-Token Prediction (MTP) head**.
17
+
18
+ A version with an MTP head grafted on (for self-speculative decoding) is available at [barozp/Qwen3.6-29B-REAP-Opus-Reasoning-Distill-MTP-GGUF](https://huggingface.co/barozp/Qwen3.6-29B-REAP-Opus-Reasoning-Distill-MTP-GGUF). See the benchmarks below to decide which one fits your hardware.
19
+
20
+ Converted with `llama.cpp`'s `convert_hf_to_gguf.py`.
21
+
22
+ ## Files
23
+
24
+ | Quant | Size | Notes |
25
+ |---|---:|---|
26
+ | BF16 | 56.5 GB | full precision reference |
27
+ | Q8_0 | 30.1 GB | near-lossless |
28
+ | Q6_K | 23.2 GB | |
29
+ | Q5_K_M | 20.2 GB | |
30
+ | Q4_K_M | 17.3 GB | most popular K-quant |
31
+ | Q4_K_S | 16.2 GB | |
32
+ | IQ4_XS | 15.3 GB | imatrix-based, smaller & often better than Q4_K_S |
33
+ | Q3_K_M | 13.7 GB | |
34
+ | IQ3_M | 12.6 GB | |
35
+ | IQ3_XXS | 11.2 GB | |
36
+ | Q2_K | 10.6 GB | |
37
+ | IQ2_M | 9.6 GB | |
38
+
39
+ All quants ≤Q4_K_M were calibrated with an importance matrix (imatrix) built from 512 samples of [barozp/opus-reasoning-distill-train](https://huggingface.co/datasets/barozp/opus-reasoning-distill-train). K-quants (`Q*_K*`) favor broad compatibility and fast CPU inference; IQ-quants (`IQ*`) require the imatrix and give better quality per bit at ≤4-bit, at some CPU-inference speed cost.
40
+
41
+ ## Why choose this over the MTP version?
42
+
43
+ The MTP-grafted version carries an extra decoder layer used for self-speculative decoding. When that feature is left enabled on hardware that's memory-bandwidth-constrained (e.g. a small/weaker GPU with heavy CPU offload), the draft+verify overhead can compete with an already-scarce resource and result in **slower** generation than this plain model. This release removes that footgun entirely — no toggle to remember, always the same speed as the MTP file with speculative decoding off.
44
+
45
+ ## Benchmarks
46
+
47
+ Measured with `llama-cli` (`Q4_K_M`, flash attention on, greedy decoding, 5 runs per config, mean ± std) on an NVIDIA RTX PRO 6000 Blackwell Server Edition (97 GB VRAM). Full methodology and the MTP comparison are on the [MTP-GGUF model card](https://huggingface.co/barozp/Qwen3.6-29B-REAP-Opus-Reasoning-Distill-MTP-GGUF#benchmarks).
48
+
49
+ | `-ngl` | tok/s (generation) |
50
+ |---:|---:|
51
+ | 99 (full offload) | 219.06 ± 0.10 |
52
+ | 20 (partial offload) | 58.60 ± 0.52 |
53
+
54
+ This matches the MTP-GGUF file with speculative decoding disabled (`--spec-type none`) within measurement noise — confirming the MTP head, when unused, carries no VRAM/compute penalty. If your setup benefits from speculative decoding (compute-bound, strong GPU, full offload), the MTP release may be faster; see its model card for numbers (+39–67% observed in our tests).
55
+
56
+ ## Credits
57
+
58
+ - Base architecture: Qwen3.6-35B-A3B (MoE), pruned via [REAP](https://huggingface.co/papers) (Router-weighted Expert Activation Pruning) to 205/256 experts.
59
+ - Reasoning distillation: LoRA fine-tune on Opus-generated reasoning traces ([barozp/opus-reasoning-distill-train](https://huggingface.co/datasets/barozp/opus-reasoning-distill-train)).