From: opencoti Subject: [PATCH 0086] F16 MoE prefill — route large-batch f16 mul_mat_id to MMF (#555) ggml_cuda_mul_mat_id had no fast CUDA path for f16/bf16/f32 experts at prefill batch: MMQ is quantized-only, and ggml_cuda_should_use_mmf rejected mul_mat_id whenever src1_ncols (= n_tokens) exceeded 512 (or 128 when the expert dim > 1024). At prefill (~32k tokens) that gate always failed, so f16 MoE fell into the host-side sorted fallback in ggml_cuda_mul_mat_id: a per-call cudaStreamSynchronize (which also disables CUDA graphs) + an O(n_expert*n_tokens*n_used) host sort + a per-expert cuBLAS GEMM loop. Quantized MoE dodges all of it via the single fused MMQ-id kernel. Net: f16 128e prefill ran ~12x slower than Q6_K (690 vs 8857 tok/s on Gemma-4-A4B 128e). The MMF kernel's `ncols_dst > 16` branch (ggml_cuda_mul_mat_f, this file) already performs GPU-side expert compaction (ggml_cuda_launch_mm_ids_helper, mmid.cuh) entirely on-stream and handles arbitrary n_tokens -- the gate was a conservative tuning heuristic, not a capability limit. This compiles out that mul_mat_id gate so large-batch f16/bf16/f32 MoE rides the fast GPU-compacted MMF path. Measured on a real 128e-F16 (bs2 RTX PRO 6000): prefill 690 -> 3081 tok/s (4.47x); RULER-vt@32k retrieval 100.0 (== Q6_K, correctness-clean). The residual gap to Q6_K is just the legitimate f16-vs-q6 weight-bandwidth ratio (~2.2x). Quantized MoE (MMQ path) and dense f16 mul_mat are untouched -- both are rejected/handled before this gate. The upstream gate is preserved, compiled out, behind -DOPENCOTI_MMF_LEGACY_GATE. diff --git a/llama.cpp/ggml/src/ggml-cuda/mmf.cu b/llama.cpp/ggml/src/ggml-cuda/mmf.cu --- a/llama.cpp/ggml/src/ggml-cuda/mmf.cu +++ b/llama.cpp/ggml/src/ggml-cuda/mmf.cu @@ -160,11 +160,22 @@ } if (mul_mat_id) { + // opencoti-hook: f16-moe-prefill (#555) — the upstream heuristic below rejected + // large-batch (prefill) f16/bf16/f32 mul_mat_id, sending it to the host-sync sorted + // fallback in ggml_cuda_mul_mat_id (per-call cudaStreamSynchronize + per-expert cuBLAS + // loop + disabled CUDA graphs => measured 12x slower than the quantized MMQ-id path: + // 690 vs 8857 tok/s prefill on Gemma-4-A4B 128e). The MMF kernel's `ncols_dst > 16` + // branch already performs GPU-side expert compaction (ggml_cuda_launch_mm_ids_helper, + // mmid.cuh) entirely on-stream and handles arbitrary n_tokens, so keep f16 MoE on the + // fast fused MMF path for any batch size. Upstream gate preserved (disabled) for ref; + // re-enable with -DOPENCOTI_MMF_LEGACY_GATE. +#ifdef OPENCOTI_MMF_LEGACY_GATE if (src0_ne[1] <= 1024 && src1_ncols > 512) { return false; } else if(src0_ne[1] > 1024 && src1_ncols > 128) { return false; } +#endif } else { if (GGML_CUDA_CC_IS_RDNA3_0(cc) && src1_ncols > 8) { return false;