MTP draft sparse-index acceptance collapses past ~index_topk under Glm5vForConditionalGeneration on DCP4 (flat glm_moe_dsa target unaffected)

#2
by tlanderso - opened

Reporting a runtime behavior in the sparkinfer v31 image referenced by this repo (verdictai/glm52-exl3-sparkinfer:v31-...), since this repo's v31 text benchmarks are the healthy control arm. Cross-posting here because the upstream fork lives on GitHub; happy to move it if there's a better channel.


sparkinfer v31: MTP draft sparse-index acceptance decorrelates with context length under Glm5vForConditionalGeneration (target unaffected)

One-line

On v31, a GLM-5.2 checkpoint wrapped as Glm5vForConditionalGeneration (MoonViT+PatchMerger graft; LLM is GlmMoeDsaForCausalLM at text_config) has its MTP draft acceptance collapse as context grows past ~`index_topk, while the **target** stays coherent to 64K. The identical LLM weights served flat as glm_moe_dsado not collapse. We isolated it to the DSA sparse-indexer/MTP path for the draft under the VLM wrapper, but the causal line is inside the compiledsparse_attn_indexer` op β€” we could not reach it from Python, and the obvious config-nesting fix does not resolve it (see Ruled out).

Reproduction (minimal)

Serve the same weights two ways on v31, same runtime/KV/knobs, and read vllm:spec_decode_num_accepted_tokens_per_pos_total (position 0) from /metrics on identical text-only prompts:

prompt tokens 2K 4K 5K 5.5K 8K 32K 64K
flat glm_moe_dsa (control) β€” 0.86 β€” β€” 0.86 0.81 0.87
glm5v wrapper 0.85 1.0 0.64 0.31 0.20 0.17 0.08
  • Wrapper collapses; flat control is flat-healthy to 64K. Image input not required (pure text triggers it; images make it worse).
  • Boundary tracks index_topk (=2048). On the older v28 base the boundary was ~2048; on v31 (with VLLM_DCP_SHARD_DRAFT=1/TOPK_OWNER_MERGE/GLOBAL_TOPK effective) it ~doubled to ~4096 β€” i.e. v31's DCP work helped but did not close it. Smooth decorrelation, not a sharp knee.
  • Target output stays coherent at all lengths; only draft acceptance falls (accept-len β†’ ~1.1).

Environment (identical across control and wrapper except the wrapper itself)

Runtime verdictai/glm52-exl3-sparkinfer:v31-gg-v20-sic3828fd-vllm0c79e41-cu132-sm120a @sha256:0433ae94665b769b78dd301f952d907508a3ba80bce47a1630ec20ade8812dff (vLLM 0.11.2.dev280+gilded.gnosis.v20.vllm0c79e41.sic3828fd), 4Γ— RTX PRO 6000 (SM120), TP4/DCP4, dcp_comm_backend=ag_rs, kv_cache_dtype=nvfp4_ds_mla, dcp_kv_cache_interleave_size=64, --attention-backend B12X_MLA_SPARSE, --moe-backend b12x, index_topk=2048, index_share_for_mtp_iteration=True, spec {"method":"mtp","num_speculative_tokens":3}, VLLM_DCP_SHARD_DRAFT=1, VLLM_DCP_TOPK_OWNER_MERGE=1, VLLM_DCP_GLOBAL_TOPK=1, VLLM_DCP_INDEXER_SHARDS=0.

Control-arm anchor: the flat glm_moe_dsa control is not a bespoke config β€” it is this repo's own published v31 preset and knobs. brandonmusic's v31 GLM-5.2-EXL3-TR3-3.0bpw release measures the flat text build on identical 4Γ—RTX PRO 6000 / TP4/DCP4 / these DCP knobs at decode C1/C4/C8 = 89.7/225.3/293.5 tok/s, GPU KV 1,132,544 tok @0.96 util, Estonia 5/5 pass. Our control arm reproduces that healthy behavior (per-position acceptance flat 0.81–0.87 to 64K); only the Glm5v wrapper differs between the healthy and collapsing arms.

Ruled out (empirically, this rig)

  • Runtime version / KV dtype: collapse present on both v28 and v31; matched nvfp4_ds_mla (KV pool 844K, β‰ˆ the flat control's 869K).
  • Image tokens: pure-text prompts collapse on the wrapper.
  • Budget size / index_topk: cannot be raised β€” index_topk=8192 fails at load: SM120 sparse MLA prefill: unsupported ... topk=8192 ... Supported: NVFP4 (GLM-family) topk in {128,512,1024,2048}. And the target holds to 64K on topk=2048, so 2048 is sufficient; the draft's failure is wrong-selection, not budget. No user-side budget lever exists.
  • Config nesting (the obvious fix β€” and it does NOT work): the DCP/indexer spec builders read num_hidden_layers and model_type from the top-level vllm_config.model_config.hf_config. Under the wrapper these are None and glm5v; flat they are 78 and glm_moe_dsa (verified offline via AutoConfig). We mirrored both onto the top-level config via hf_overrides (num_hidden_layers=78, then also model_type=glm_moe_dsa) β€” the model loaded and vision still worked, but the boundary did not move (8K stayed 0.23, 32K ~0.15). So although the nesting defect is real and statically verifiable, it is not the causal channel for this collapse (or the compiled op does not consume these top-level values). Reads are at deepseek_v2.py::_indexer_cache_dcp_shard_count / get_kv_cache_spec and mla_attention.py (L1940–1961).
  • Mirroring the index cfg onto the DRAFT's text_config (incl. index_share_for_mtp_iteration): also does NOT work, on both v28 and v31. We propagated {index_topk, index_topk_freq, index_share_for_mtp_iteration=True, index_topk_pattern, use_index_cache} from the top-level Glm5vConfig onto the draft config in SpeculativeConfig.hf_config_override (propagation confirmed in logs). Same-session A/B on v31 (nvfp4/MTP3, draft-sharding knobs live), per-position pos0 OFFβ†’ON: 8K 0.32β†’0.32, 32K 0.29β†’0.20, 64K 0.095β†’0.10 β€” boundary unchanged. So the draft does receive the sparse-index config; the collapse is not a missing-config problem at either the top level or the draft level.

Not intrinsically unsupported (existence proof)

A different engine (stock vLLM + this same MoonViT/PatchMerger graft, DCP4, sparse MLA, NVFP4 KV) sustains ~53% draft acceptance to 45K. So "sparse indexer + DCP + a VLM wrapper is unsupported, use TP-only" is not the explanation β€” the combination works elsewhere.

Where it must be

Given the above, the wrapper-specific difference reaches the draft's top-2048 selection through a path not visible from the Python config: the compiled sparse_attn_indexer op, and/or how the Glm5v wrapper feeds hidden-states / the shared MTP index (index_share_for_mtp_iteration) to the draft. The symptom β€” correct while context ≀ ~index_topk, smooth decorrelation past it, target unaffected β€” is consistent with the draft's shared/selected index being computed over the wrong context extent only under the wrapper.

Ask

A one-line log inside the draft's sparse-indexer of (a) the effective merged-index size / selected context extent and (b) the resolved dcp_world_size/dcp_replicated, for the flat glm_moe_dsa control vs the glm5v wrapper, at 4K and 8K. That single comparison should localize whether the draft is selecting over a truncated context under the wrapper. We have a standing minimal repro and can run instrumented builds.

you may want to send your agent back to this, it says " glm_moe_dsa" is part of this. that is neither in the docker compose yaml or server.sh script. the Glm5vForConditionalGeneration wrapper ( i assume you mean of a checkpoint with a vision head strapped to it) is not necessarily compatible with this. This checkpoint and the compose yaml and server.sh script were not written for that. There are other ones who have done it but it's beyond the scope of this, and i'm not sure your agent is accurate about this model, it contains some assumptions that I do not see as true.

Fair enough, and thanks for the quick reply β€” you're right on the main point. The Glm5vForConditionalGeneration wrapper (MoonViT + PatchMerger, LLM demoted to text_config) is my own graft on top of your weights, not your release. Your checkpoint is flat glm_moe_dsa and the compose/server.sh serve exactly that, so the vision-nesting behavior is genuinely outside your scope. My fault for framing it against your artifacts.

One small note just for the record: glm_moe_dsa isn't something external I assumed β€” it's the model_type / architectures (GlmMoeDsaForCausalLM) in your own config.json. I only referenced it as the LLM sitting at text_config inside my wrapper, i.e. the same arch as your flat model β€” not implying the compose or server.sh name it.

Either way it isn't a weights issue: the collapse is image-independent, wrapper-only, and kicks in exactly at index_topk=2048, which points at the sparse-index share for the MTP draft in the sparkinfer runtime (verdictai/glm52-exl3-sparkinfer:v31) under DCP + Glm5v nesting β€” not your quant. That's the layer I need to instrument, so I'll take it there. Appreciate you drawing the boundary; it actually helps localize it.

Sign up or log in to comment