| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| README.md | 2.18 kB xet | 5e260149 | |
| gemma-4-26B-A4B-it-BF16-MTP.gguf | 855 MB xet | f5d09275 | |
| gemma-4-26B-A4B-it-F16-MTP.gguf | 855 MB xet | 2291cf0e | |
| gemma-4-26B-A4B-it-Q8_0-MTP.gguf | 462 MB xet | fb7a81da |
Gemma 4 26B-A4B MTP drafter
Multi-Token Prediction (MTP) drafter for unsloth/gemma-4-26B-A4B-it-GGUF. It runs as a speculative draft model that shares the target's KV cache and speeds up text generation.
Verified on a single B200 with the gemma-4-26B-A4B-it-UD-Q4_K_M.gguf target: 159 tok/s without MTP, 257 tok/s with MTP, 0.72 draft acceptance.
MTP was merged into llama.cpp on 2026-06-07 (PR ggml-org/llama.cpp#23398). You need a llama.cpp build from after that date. Older builds cannot load these (arch gemma4-assistant).
Files
For -hf auto-discovery a Q8_0 drafter sits at the repo root as mtp-gemma-4-26B-A4B-it.gguf. The same three precisions live in MTP/:
gemma-4-26B-A4B-it-Q8_0-MTP.gguf(smallest, recommended; mirrored at the repo root asmtp-gemma-4-26B-A4B-it.gguf)gemma-4-26B-A4B-it-BF16-MTP.ggufgemma-4-26B-A4B-it-F16-MTP.gguf
Build llama.cpp
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
# CUDA build. Set the arch for your GPU: 89 (RTX 4090), 90 (H100), 100 (B200).
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=90
cmake --build build --config Release -j --target llama-server
Run, the easy way
A recent llama.cpp finds the drafter automatically from the root mtp- file, so -hf is all you need. No --model-draft.
./build/bin/llama-server \
-hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M \
--spec-type draft-mtp --spec-draft-n-max 4 \
-ngl 999 -fa on
If your build is too old to auto-discover the sibling, use the explicit form below.
Run with an explicit drafter
Use this to choose a precision or point at a local file.
hf download unsloth/gemma-4-26B-A4B-it-GGUF gemma-4-26B-A4B-it-UD-Q4_K_M.gguf --local-dir .
hf download unsloth/gemma-4-26B-A4B-it-GGUF MTP/gemma-4-26B-A4B-it-Q8_0-MTP.gguf --local-dir .
./build/bin/llama-server \
-m gemma-4-26B-A4B-it-UD-Q4_K_M.gguf \
--model-draft MTP/gemma-4-26B-A4B-it-Q8_0-MTP.gguf \
--spec-type draft-mtp --spec-draft-n-max 4 \
-ngl 999 -fa on
Multi GPU: add --spec-draft-device CUDA0 -sm layer. The drafter pairs with any quant of the 26B-A4B. Quantized KV cache works.
- Total size
- 393 GB
- Files
- 34
- Last updated
- Jun 28
- Pre-warmed CDN
- US EU US EU