EAGLE-3 Draft Model for Qwen/Qwen3.5-27B-FP8

EAGLE-3 draft model for accelerating Qwen/Qwen3.5-27B-FP8 inference with speculative decoding.

EAGLE-3 is lossless: the draft proposes several tokens and the base model verifies them in a single forward pass, so the output is identical to standard decoding. The quality metric is mean acceptance length (tokens accepted per base-model forward); a value > 2.0 gives a useful speedup. Wall-clock tok/s depends on hardware, batch size and context length — measure it on your own setup.

What is EAGLE-3?

EAGLE-3 is a speculative decoding method in which a small (~1B) draft model predicts several tokens ahead, while the larger base model verifies them in a single forward pass. Unlike EAGLE/EAGLE-2, EAGLE-3 predicts tokens directly (without feature prediction) and fuses features from multiple layers of the target model (low-, mid-, and high-level).

Usage — SGLang (>= 0.5.12)

Repo layout (same as the PRISM-EAGLE3 serving repos):

compressed/               # drafter weights + config (LlamaForCausalLMEagle3)
patch_sglang_eagle3.py    # one-shot patch: adds the EAGLE-3 aux-hidden capture
                          # hook to SGLang's qwen3_5.py (needed for ANY drafter
                          # on a qwen3_5-family target)
serve_sglang.sh           # launch helper
huggingface-cli download VirVen/Qwen3.5-27B-EAGLE3-v2 --local-dir .
pip install "sglang>=0.5.12"
python3 patch_sglang_eagle3.py
MODEL=Qwen/Qwen3.5-27B-FP8 DRAFTER=./compressed ./serve_sglang.sh

Or launch manually:

SGLANG_ENABLE_SPEC_V2=1 python3 -m sglang.launch_server \
    --model-path Qwen/Qwen3.5-27B-FP8 \
    --speculative-algorithm EAGLE3 \
    --speculative-draft-model-path ./compressed \
    --speculative-num-steps 3 \
    --speculative-eagle-topk 1 \
    --speculative-num-draft-tokens 4 \
    --mamba-scheduler-strategy extra_buffer \
    --mm-attention-backend sdpa \
    --dtype bfloat16 --trust-remote-code

Training Details

  • Base model: Qwen/Qwen3.5-27B-FP8 (Qwen3-32B architecture, FP8)
  • Draft model: ~1B parameters, one transformer layer, attention input size of 2×hidden_size (concatenated token embedding and fused features)
  • Feature fusion: layers 8 (low-level), 32 (mid-level), and 62 (high-level) of the target model
  • Data: ~50k examples from the Saiga dataset (Russian-language conversations)
  • Training: DeepSpeed ZeRO-2
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VirVen/Qwen3.5-27B-EAGLE3-v2

Base model

Qwen/Qwen3.5-27B
Finetuned
(1)
this model

Paper for VirVen/Qwen3.5-27B-EAGLE3-v2