# DFlash speculative decoding for this model (MLX) > **Note:** no extra weights are needed in this repo — DFlash works with the existing quant *as-is*, plus an external drafter. This file is a usage guide. **[DFlash](https://huggingface.co/z-lab/Qwen3.6-27B-DFlash)** is a block-diffusion *speculative drafter* trained for Qwen3.6-27B targets. It drafts a 16-token block by diffusion and lets this model verify it autoregressively — so output is **lossless** (identical to plain decoding), just faster. ## Quick start ```bash pip install -U mlx_vlm # one-time: accept the gated drafter at https://huggingface.co/z-lab/Qwen3.6-27B-DFlash python3 -m mlx_vlm generate \ --model osmapi/osmQwopus-3.6-27B-V2-heretic-abliterated-uncensored-8-bit-mlx \ --draft-model z-lab/Qwen3.6-27B-DFlash --draft-kind dflash \ --prompt "" --max-tokens 256 ``` Serve it (OpenAI-compatible): ```bash python3 -m mlx_vlm server \ --model osmapi/osmQwopus-3.6-27B-V2-heretic-abliterated-uncensored-8-bit-mlx \ --draft-model z-lab/Qwen3.6-27B-DFlash --draft-kind dflash ``` ## Measured on an Apple M4 Max (128 GB) | Quant | AR baseline | + DFlash | Speedup | |---|---|---|---| | 8-bit (affine) | 16.4 tok/s | **55.3 tok/s** | **3.38×** | | bf16 (full) | 9.0 tok/s | **33.2 tok/s** | **3.67×** | Acceptance ≈ 8.95 tokens/round; +~3.9 GB peak memory for the drafter; small TTFT increase. ## Notes & limits - **Text path only** — the vision tower is not accelerated. - Speedup is workload-dependent (acceptance varies by prompt). - Larger on more memory-bound quants (bf16 > 8-bit). **Background & full benchmarks:** [https://huggingface.co/blog/junafinity/block-diffusion-on-apple-silicon-with-3-7x-speedup] **Credit:** DFlash by [z-lab](https://huggingface.co/z-lab) (arXiv:2602.06036); runtime [`mlx_vlm`](https://github.com/Blaizzy/mlx-vlm).