# MTP_QUANT_MANIFEST **Label:** gptq_mtp_routed_experts_v2 **Method:** Frantar-style GPTQ with Cholesky H^-1; per-expert per-group sym INT4 group=128 **is_final_quality_preserving:** True (real GPTQ, NOT RTN) ## Calibration - tokens_processed: 473372 - per-expert token count: min=430 max=175375 mean=11095 - damp_fraction: 0.01 ## Note - MTP forward replay skipped attention (treated post-attn as identity-residual input); RMSNorms, h_proj/e_proj, gate, and per-expert MLP path are real. ## Measured draft acceptance (post-release, num_speculative_tokens=1) Weighted draft acceptance from live vLLM spec-decode metrics. Output is verifier-exact regardless — acceptance affects decode speed only. | Profile | Accepted / drafted | Weighted acceptance | Mean accepted length | |---|---:|---:|---:| | General chat, 128k | 2354/3030 | 77.7% | 1.78 | | Production, 262k | 15274/17991 | 84.9% | 1.85 | | Long-context research variant, 262k | 319/357 | 89.4% | 1.89 | | Code-heavy (ext10 gate, adapter-assisted) | — | 93–96% | — |