dived42 commited on
Commit
71ce7fd
·
verified ·
1 Parent(s): 63768c1

Fix MTP ignore names for SGLang fused Linear layers (`qkv_proj` / `gate_up_proj`)

Browse files

## Summary

Hi @cyankiwi, I noticed that the MTP head is incorrectly treated as quantized when this checkpoint is loaded by SGLang, which makes NEXTN speculative decoding nearly ineffective.

SGLang fuses QKV and gate/up at module construction (`self_attn.qkv_proj`, `mlp.gate_up_proj`). The ignore check uses these runtime names, not the names produced during `load_weights()`. The MTP entry has no `packed_modules_mapping`, so the fused names are not expanded back to the original shards. Both Linears then fall back to the default INT4 scheme while the checkpoint tensors are BF16.

## Fix

Add the fused runtime names to `ignore`:

```diff
"ignore": [
+ "mtp.layers.0.self_attn.qkv_proj",
+ "mtp.layers.0.mlp.gate_up_proj"
]

## Reproduction

Hardware: 4× RTX 4090, TP=2, PD disaggregation (1 prefill + 1 decode)

```bash
python -m sglang.launch_server \
--model-path <model> \
--tp-size 2 \
--speculative-algorithm NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--speculative-attention-mode decode
```

Decode `accept len`: ~1.03 before the fix, ~3 after.

PD / HiCache / metrics / trace are orthogonal to the config issue and not required
to observe the difference.

Files changed (1) hide show
  1. config.json +3 -1
config.json CHANGED
@@ -342,10 +342,12 @@
342
  "mtp.layers.0.mlp.down_proj",
343
  "mtp.layers.0.mlp.gate_proj",
344
  "mtp.layers.0.mlp.up_proj",
 
345
  "mtp.layers.0.self_attn.k_proj",
346
  "mtp.layers.0.self_attn.o_proj",
347
  "mtp.layers.0.self_attn.q_proj",
348
- "mtp.layers.0.self_attn.v_proj"
 
349
  ],
350
  "kv_cache_scheme": null,
351
  "quant_method": "compressed-tensors",
 
342
  "mtp.layers.0.mlp.down_proj",
343
  "mtp.layers.0.mlp.gate_proj",
344
  "mtp.layers.0.mlp.up_proj",
345
+ "mtp.layers.0.mlp.gate_up_proj",
346
  "mtp.layers.0.self_attn.k_proj",
347
  "mtp.layers.0.self_attn.o_proj",
348
  "mtp.layers.0.self_attn.q_proj",
349
+ "mtp.layers.0.self_attn.v_proj",
350
+ "mtp.layers.0.self_attn.qkv_proj"
351
  ],
352
  "kv_cache_scheme": null,
353
  "quant_method": "compressed-tensors",