File size: 1,750 Bytes
a622121 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 | ---
license: apache-2.0
base_model: Qwen/Qwen3-8B
tags:
- speculative-decoding
- dflash
- sglang
- eagle
---
# Qwen3-8B DFlash speculative-decoding draft (on-policy)
A **DFlash** speculative-decoding draft head for `Qwen/Qwen3-8B`, trained with
[SpecForge](https://github.com/sgl-project/SpecForge).
## What it is
- Draft architecture: `DFlashDraftModel` (5 draft layers, hidden 4096), block size 16.
- Consumes 5 aux hidden states from the frozen target at layers **[1, 9, 17, 25, 33]**
(fc input dim = 5 x 4096 = 20480); frozen shared embedding + LM head.
## Training
- **Data: on-policy** — perfectblend prompts with responses **regenerated by
Qwen/Qwen3-8B itself** (greedy, temperature 0, no-think), ~50k conversations.
- 6 epochs, num_anchors 512, block_size 16, lr 6e-4, cosine (0.04 warmup),
loss-decay gamma 7, max_length 3072, `train_dflash.py` (hf backend).
- Final training block-accuracy ~0.6.
- **Why on-policy matters:** the DFlash training accuracy is scored against the
target's own greedy continuations; off-policy (human) text caps the achievable
accuracy at ~0.35, while on-policy data lifts it past 0.6.
## Evaluation — serving accept length (tau)
Measured via sglang `avg_spec_accept_length` on **gsm8k** (100 questions, greedy,
no-thinking, block size 16):
| draft | gsm8k tau |
|---|---|
| this checkpoint (on-policy) | **5.75** |
| off-policy baseline | 2.30 |
| DFlash paper (Qwen3-8B) | 6.54 |
| EAGLE-3 (paper baseline) | 2.96 |
## Usage (sglang)
```bash
python -m sglang.launch_server \
--model-path Qwen/Qwen3-8B --trust-remote-code \
--speculative-algorithm DFLASH \
--speculative-draft-model-path cm00cm/qwen3-8b-dflash-draft-onpolicy-perfectblend \
--speculative-dflash-block-size 16
```
|