Usage

For now, you need to use the gg/spec-mtp-experiments branch on llama.cpp or a custom mtp fork.

You can switch to the mtp branch with git checkout gg/spec-mtp-experiments after cloning and entering the llama.cpp repository.

Add --spec-type mtp --spec-draft-n-max 5 --spec-draft-n-min 0 to your llama-server or llama-cli command.

Feel free to tweak --spec-draft-n-max and find out what works best for your setup.

Try not to push --spec-draft-n-min too far, keep it in single digits.

Credits

  • Qwen for this amazing model.
  • Unsloth for the imatrix quantization file.
  • All of the llama.cpp and ggml contributors for allowing me to run AI models locally.
Downloads last month
52
Safetensors
Model size
0.9B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EntityDeletr/Qwen3.5-0.8B-MTP-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(341)
this model