name: fp8_dynamic_attnbf16 scheme: FP8_DYNAMIC # FP8 weights, FP8 dynamic per-token activations — data-free, vLLM-native engine: llmcompressor # Variant of `fp8_dynamic.yaml` that additionally keeps the ENTIRE self-attention # block in bf16, leaving only the MLPs quantised. # # Rationale: on Qwen3.5/3.6 only a quarter of the layers are `full_attention` # (16 of 64 on Qwen3.6-27B — the rest are `linear_attn`), so the whole self_attn # block is just ~1.68 B params and FP8-ing it saves only ~1.56 GiB. The MLPs are # ~17.1 B params and deliver ~15.9 GiB of the savings. Giving up 4.6% of on-disk # size buys a completely bf16 attention path. # # The specific thing this protects: `attn_output_gate: true` means the attention # OUTPUT GATE is fused into `q_proj`, which is why q_proj is [2*heads*head_dim, # hidden] = [12288, 5120] rather than [6144, 5120]. Half that tensor is a # multiplicative per-head gate on what attention writes into the residual stream, # and the 16 full-attention layers carry the long-range retrieval. Quantisation # error on a multiplicative gate is qualitatively worse than on an additive # projection. Everyone (upstream, Qwen official) quantises q_proj; this recipe # does not. # # Use `fp8_dynamic.yaml` for the standard build; use this one when multi-turn / # long-context stability matters more than 1.5 GiB. calibration: dataset: neuralmagic/calibration config: LLM split: train num_samples: 4 max_seq_length: 512 ignore: - lm_head - "re:.*visual.*" - "re:.*linear_attn.*" # Mamba/SSM block stays in bf16 — same rationale as the NVFP4A16 build - "re:.*self_attn.*" # THE VARIANT: q/k/v/o_proj too, incl. the output gate fused into q_proj - "re:.*mtp.*" # Note: Qwen3.6-27B is dense; on MoE bases also include # "re:.*mlp.gate$" and "re:.*mlp.shared_expert_gate$". export: save_compressed: true