Qwen3.6-27B-V2-abliterated-uncensored-6-bit-mlx / DFLASH_SPECULATIVE_DECODING.md
rogKesavan90's picture
Initial commit
115187f
|
Raw
History Blame Contribute Delete
1.86 kB

DFlash speculative decoding for this model (MLX)

Note: no extra weights are needed in this repo — DFlash works with the existing quant as-is, plus an external drafter. This file is a usage guide.

DFlash is a block-diffusion speculative drafter trained for Qwen3.6-27B targets. It drafts a 16-token block by diffusion and lets this model verify it autoregressively — so output is lossless (identical to plain decoding), just faster.

Quick start

pip install -U mlx_vlm
# one-time: accept the gated drafter at https://huggingface.co/z-lab/Qwen3.6-27B-DFlash
python3 -m mlx_vlm generate \
  --model osmapi/osmQwopus-3.6-27B-V2-heretic-abliterated-uncensored-6-bit-mlx \
  --draft-model z-lab/Qwen3.6-27B-DFlash --draft-kind dflash \
  --prompt "<your prompt>" --max-tokens 256

Serve it (OpenAI-compatible):

python3 -m mlx_vlm server \
  --model osmapi/osmQwopus-3.6-27B-V2-heretic-abliterated-uncensored-6-bit-mlx \
  --draft-model z-lab/Qwen3.6-27B-DFlash --draft-kind dflash

Measured on an Apple M4 Max (128 GB)

Quant AR baseline + DFlash Speedup
8-bit (affine) 16.4 tok/s 55.3 tok/s 3.38×
bf16 (full) 9.0 tok/s 33.2 tok/s 3.67×

Acceptance ≈ 8.95 tokens/round; +~3.9 GB peak memory for the drafter; small TTFT increase.

Notes & limits

  • Text path only — the vision tower is not accelerated.
  • Speedup is workload-dependent (acceptance varies by prompt).
  • Larger on more memory-bound quants (bf16 > 8-bit).

Background & full benchmarks: [https://huggingface.co/blog/junafinity/block-diffusion-on-apple-silicon-with-3-7x-speedup]

Credit: DFlash by z-lab (arXiv:2602.06036); runtime mlx_vlm.