Qwen3-8B DFlash speculative-decoding draft (on-policy)
A DFlash speculative-decoding draft head for Qwen/Qwen3-8B, trained with
SpecForge.
What it is
- Draft architecture:
DFlashDraftModel(5 draft layers, hidden 4096), block size 16. - Consumes 5 aux hidden states from the frozen target at layers [1, 9, 17, 25, 33] (fc input dim = 5 x 4096 = 20480); frozen shared embedding + LM head.
Training
- Data: on-policy โ perfectblend prompts with responses regenerated by Qwen/Qwen3-8B itself (greedy, temperature 0, no-think), ~50k conversations.
- 6 epochs, num_anchors 512, block_size 16, lr 6e-4, cosine (0.04 warmup),
loss-decay gamma 7, max_length 3072,
train_dflash.py(hf backend). - Final training block-accuracy ~0.6.
- Why on-policy matters: the DFlash training accuracy is scored against the target's own greedy continuations; off-policy (human) text caps the achievable accuracy at ~0.35, while on-policy data lifts it past 0.6.
Evaluation โ serving accept length (tau)
Measured via sglang avg_spec_accept_length on gsm8k (100 questions, greedy,
no-thinking, block size 16):
| draft | gsm8k tau |
|---|---|
| this checkpoint (on-policy) | 5.75 |
| off-policy baseline | 2.30 |
| DFlash paper (Qwen3-8B) | 6.54 |
| EAGLE-3 (paper baseline) | 2.96 |
Usage (sglang)
python -m sglang.launch_server \
--model-path Qwen/Qwen3-8B --trust-remote-code \
--speculative-algorithm DFLASH \
--speculative-draft-model-path cm00cm/qwen3-8b-dflash-draft-onpolicy-perfectblend \
--speculative-dflash-block-size 16
- Downloads last month
- 21
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support