DFlash Speculator for DeepSeek-V4-Flash (all-SWA, Muon, 50k)

A DFlash draft (speculator) model trained for speculative decoding with DeepSeek-V4-Flash as the verifier.

This is the best checkpoint selected by validation loss from a 50k-sample training run.

Key characteristics

  • Algorithm: DFlash
  • Verifier / target: deepseek-ai/DeepSeek-V4-Flash
  • Attention: Sliding-window attention (SWA) on all draft layers (sliding_window=2048)
  • Optimizer: Muon
  • Training data: 50k samples
  • Draft layers: 5 (hidden_size=4096, head_dim=256, hc_mult=4)
  • Speculative tokens: 7 (block_size 8)
  • Aux hidden-state layers: [3, 13, 23, 32, 42]
  • dtype: bfloat16

Validation metrics (best checkpoint)

Metric Value
Loss 1.2559
Full-sequence acceptance 0.3452
Position 1 acc 0.7402
Position 2 acc 0.5120
Position 3 acc 0.3689
Position 4 acc 0.2746
Position 5 acc 0.2121
Position 6 acc 0.1691
Position 7 acc 0.1373

Files

  • config.json / config.py — speculator config and model definition
  • model.safetensors — draft model weights
  • val_metrics.json — validation metrics for this checkpoint
Downloads last month
115
Safetensors
Model size
2B params
Tensor type
I64
·
BF16
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inference-optimization/dflash-DeepSeek-V4-Flash-all-swa-muon-speculators-50k

Finetuned
(18)
this model