File size: 7,458 Bytes
89096dd e341888 bfc7f3a 4652708 e341888 9677f59 4652708 e341888 89096dd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 | ---
base_model: DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
license: apache-2.0
library_name: gguf
tags:
- gguf
- rocmfp4
- qwen3.6
- deckard
- 40b
- mtp
- speculative-decoding
- frankenmerge
- strix-halo
- amd
- rocm
- vulkan
language:
- en
---
# Qwen3.6-40B-Deckard-MTP — ROCmFP4 STRIX (Q8 embeddings)
Experimental **AMD Strix Halo (gfx1151)** quant of the **40B Deckard** frankenmerge — with an MTP /
next-token-prediction head **transplanted in** — in the custom **ROCmFP4** 4-bit format, for
self-speculative decoding on a single APU.
> **⚠️ Read the lineage + status below — this is a multi-hop community derivative.** It is a
> frankenmerge with a *grafted* MTP head, then 4-bit quantized. It works, but it is several steps
> removed from any officially-released model.
> **⚠️ Ignore HuggingFace's auto-detected quant badge ("F16"/16-bit) — it's wrong.**
> HF can't read the custom ROCmFP4 tensor types and falls back to mislabeling. **These are ~4.5 bpw
> 4-bit ROCmFP4 files, not 16-bit.** Pick by filename in *Files and versions*.
> ## Requires the ROCmFP4 fork (public) — not stock llama.cpp
>
> Uses the **ROCmFP4** tensor types (`q4_0_rocmfp4`, `q4_0_rocmfp4_fast`). **Stock llama.cpp, LM
> Studio, Ollama, etc. cannot load it.** Build/run with
> **[`charlie12345/rocmfp4-llama`](https://github.com/charlie12345/rocmfp4-llama)** (`mtp-rocmfp4-strix`).
## Status — functionally verified, NOT quality-benchmarked
Honest scope of testing on a Framework Desktop / Ryzen AI Max+ 395 (gfx1151, 128 GB unified):
- ✅ **Loads and runs** on the ROCmFP4 fork (Vulkan0).
- ✅ **MTP works** — transplanted head hit **~0.95 draft acceptance** in a quick test.
- ✅ **~25 t/s** decode short-context; coherent code generation.
- ❌ **No** perplexity/KL testing, **no** coding/reasoning eval, **no** comparison vs the source.
Treat this as a "does it run with working MTP on Strix Halo" artifact, **not** a quality claim. Because
it's a frankenmerge + a *borrowed* draft head + a 4-bit quant, behavior is extra unpredictable.
Reproduce / evaluate before relying on it.
## Lineage (who made what)
This file is a stack of other people's work — credit where due:
1. **Base model:** [`DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking`](https://huggingface.co/DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking)
— a ~40B Qwen3.6 frankenmerge (Claude-Opus-distilled, uncensored thinking/coder). Source used here:
DavidAU's **NEO-CODE Q8_0** GGUF.
2. **MTP-injection method:** PiehSoft's `inject_mtp_40b.py` (from
[`PiehSoft/Qwen3.6-40B-Deckard-MTP-Q6_K`](https://huggingface.co/PiehSoft/Qwen3.6-40B-Deckard-MTP-Q6_K)),
which raw-copies a 27B `nextn` head into the 40B.
3. **MTP donor head:** the `nextn` block from a Qwen3.6-27B MTP model
([Jackrong's Qwopus-Coder](https://huggingface.co/Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF)) — a
Claude-Opus distill, so it transplants well (hence the 0.95 acceptance).
4. **ROCmFP4 quant + runtime:** [`charlie12345/rocmfp4-llama`](https://github.com/charlie12345/rocmfp4-llama).
The MTP head is **not native** to the Deckard merge — it's grafted. The draft predictions are
"borrowed" from a 27B and happen to align well; they were not trained on this model.
## Three files in this repo (pick your trade-off)
All three share the same ROCmFP4 4-bit body and Q8 embeddings. They differ on two independent knobs —
the **output head** (lm-head) and the **MTP draft head**:
| File | size | output head | MTP draft head | best for |
|---|---|---|---|---|
| `…-STRIX-embQ8.gguf` | 21.8 GB | 4-bit | 4-bit (Qwopus donor) | **fastest** — the recommended default |
| `…-STRIX-embQ8-headQ6.gguf` | 22.2 GB | **Q6_K** | 4-bit (Qwopus donor) | better output quality |
| `…-STRIX-embQ8-mtpF16.gguf` | ~22.4 GB | 4-bit | **f16** (BF16 base-27B donor) | experiment — no acceptance gain vs base |
- **Output head** (`output.weight`): 4-bit vs `Q6_K`. Q6_K is a touch more faithful at ~5–7% slower
decode (trade measured on the [27B sibling](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF)).
- **MTP draft head** (`blk.96`): the base & headQ6 variants graft a **4-bit** head from a **Qwopus-27B**
donor (a Claude-Opus distill); the **mtpF16** variant keeps the head at **f16** from a **BF16 base-27B**
donor — mirroring [PiehSoft's high-precision-head choice](https://huggingface.co/PiehSoft/Qwen3.6-40B-Deckard-MTP-Q6_K).
**Controlled warm comparison (identical 6-turn session, f16/f16 KV both):** 4-bit/Qwopus head **~0.82
sustained**, f16/base-27B head **~0.84 sustained** — a **wash** (the gap is within turn-to-turn noise; the
earlier 0.95-vs-0.77 split was just cold-vs-cold noise). **Conclusion: the f16 MTP head does not
meaningfully improve acceptance** — neither head precision nor donor moved it. So **`mtpF16` is effectively
redundant with `base`** (same speed, ~0.6 GB bigger). **Use `base` (or `headQ6` for output quality);
`mtpF16` is kept only as a documented negative-result experiment.**
- **Q8 embeddings** (not f16): the source is Q8_0, so f16 would be fake-f16 bloat — Q8 matches the
source precision exactly.
## How it was built (reproducible — incl. the gotcha)
```bash
# 1) graft a 27B nextn head into the 40B Q8_0 (raw binary copy)
python inject_mtp_40b.py \
--target Qwen3.6-40B-Deck-...-Q8_0.gguf \
--donor Qwopus3.6-27B-Coder-MTP-Q8_0.gguf \
--output deckard-40b-mtp.gguf --source-layer 64 --dest-layer 96
# 2) ** GOTCHA ** the inject script adds nextn tensors but does NOT bump block_count.
# An injected MTP model needs block_count += nextn_predict_layers, or the loader
# looks for the head one slot too low (e.g. "missing tensor blk.95.nextn..."):
python gguf_set_metadata.py deckard-40b-mtp.gguf qwen35.block_count 97 --force
# 3) quantize -> ROCmFP4 (Q8 embeddings, no imatrix). headQ6 adds --output-tensor-type q6_K
llama-quantize --allow-requantize --token-embedding-type q8_0 \
deckard-40b-mtp.gguf Qwen3.6-40B-Deckard-MTP-ROCmFP4-STRIX-embQ8.gguf Q4_0_ROCMFP4_STRIX
```
Kernels: 631 tensors on `q4_0_rocmfp4_fast`, 122 on dual-scale `q4_0_rocmfp4` (incl. the grafted head).
## Usage — coding model, try thinking OFF
It's a "…Thinking" model; for coding it tends to do better **off-thinking**. Toggle per-request with
`"chat_template_kwargs": {"enable_thinking": false}`, or launch with `--reasoning-format none`.
```bash
env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server -m Qwen3.6-40B-Deckard-MTP-ROCmFP4-STRIX-embQ8.gguf \
--alias deckard-40b-rocmfp4-mtp --host 0.0.0.0 --port 8080 \
-dev Vulkan0 -ngl 999 -fa on -c 262144 -b 2048 -ub 256 -t 16 -tb 16 \
-ctk f16 -ctv f16 -cpent 256 -ctxcp 32 --cache-reuse 256 \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0 \
--spec-type draft-mtp --spec-draft-device Vulkan0 --spec-draft-ngl all \
--spec-draft-type-k f16 --spec-draft-type-v f16 \
--spec-draft-n-max 3 --spec-draft-n-min 0 --spec-draft-p-min 0.0 --spec-draft-p-split 0.10 \
--reasoning on --reasoning-format deepseek \
--chat-template-kwargs '{"preserve_thinking": true}' --jinja --parallel 1 --metrics --no-mmap
```
## License
Derivative of the Deckard merge (Qwen3.6 lineage, Apache-2.0). Inherits the base model's terms —
verify them before redistribution/use. The MTP graft and quant add no new restrictions.
|