Fix MTP ignore names for SGLang fused Linear layers (`qkv_proj` / `gate_up_proj`)

#9
by dived42 - opened

Summary

Hi @cyankiwi, I noticed that the MTP head is incorrectly treated as quantized when this checkpoint is loaded by SGLang, which makes NEXTN speculative decoding nearly ineffective.

SGLang fuses QKV and gate/up at module construction (self_attn.qkv_proj, mlp.gate_up_proj). The ignore check uses these runtime names, not the names produced during load_weights(). The MTP entry has no packed_modules_mapping, so the fused names are not expanded back to the original shards. Both Linears then fall back to the default INT4 scheme while the checkpoint tensors are BF16.

Fix

Add the fused runtime names to ignore:

 "ignore": [
+  "mtp.layers.0.self_attn.qkv_proj",
+  "mtp.layers.0.mlp.gate_up_proj"
 ]

## Reproduction

Hardware: 4× RTX 4090, TP=2, PD disaggregation (1 prefill + 1 decode)

```bash
python -m sglang.launch_server \
  --model-path <model> \
  --tp-size 2 \
  --speculative-algorithm NEXTN \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --speculative-attention-mode decode

Decode accept len: ~1.03 before the fix, ~3 after.

PD / HiCache / metrics / trace are orthogonal to the config issue and not required
to observe the difference.

cyankiwi org

Thanks for the PR :)

cpatonn changed pull request status to merged

Sign up or log in to comment