| From: opencoti |
| Subject: [PATCH 0086] F16 MoE prefill — route large-batch f16 mul_mat_id to MMF (#555) |
|
|
| ggml_cuda_mul_mat_id had no fast CUDA path for f16/bf16/f32 experts at prefill batch: |
| MMQ is quantized-only, and ggml_cuda_should_use_mmf rejected mul_mat_id whenever |
| src1_ncols (= n_tokens) exceeded 512 (or 128 when the expert dim > 1024). At prefill |
| (~32k tokens) that gate always failed, so f16 MoE fell into the host-side sorted |
| fallback in ggml_cuda_mul_mat_id: a per-call cudaStreamSynchronize (which also disables |
| CUDA graphs) + an O(n_expert*n_tokens*n_used) host sort + a per-expert cuBLAS GEMM loop. |
| Quantized MoE dodges all of it via the single fused MMQ-id kernel. Net: f16 128e prefill |
| ran ~12x slower than Q6_K (690 vs 8857 tok/s on Gemma-4-A4B 128e). |
|
|
| The MMF kernel's `ncols_dst > 16` branch (ggml_cuda_mul_mat_f, this file) already performs |
| GPU-side expert compaction (ggml_cuda_launch_mm_ids_helper, mmid.cuh) entirely on-stream |
| and handles arbitrary n_tokens -- the gate was a conservative tuning heuristic, not a |
| capability limit. This compiles out that mul_mat_id gate so large-batch f16/bf16/f32 MoE |
| rides the fast GPU-compacted MMF path. Measured on a real 128e-F16 (bs2 RTX PRO 6000): |
| prefill 690 -> 3081 tok/s (4.47x); RULER-vt@32k retrieval 100.0 (== Q6_K, correctness-clean). |
| The residual gap to Q6_K is just the legitimate f16-vs-q6 weight-bandwidth ratio (~2.2x). |
| Quantized MoE (MMQ path) and dense f16 mul_mat are untouched -- both are rejected/handled |
| before this gate. The upstream gate is preserved, compiled out, behind |
| -DOPENCOTI_MMF_LEGACY_GATE. |
|
|
| |
| |
| |
| @@ -160,11 +160,22 @@ |
| } |
| |
| if (mul_mat_id) { |
| + // opencoti-hook: f16-moe-prefill (#555) — the upstream heuristic below rejected |
| + // large-batch (prefill) f16/bf16/f32 mul_mat_id, sending it to the host-sync sorted |
| + // fallback in ggml_cuda_mul_mat_id (per-call cudaStreamSynchronize + per-expert cuBLAS |
| + // loop + disabled CUDA graphs => measured 12x slower than the quantized MMQ-id path: |
| + // 690 vs 8857 tok/s prefill on Gemma-4-A4B 128e). The MMF kernel's `ncols_dst > 16` |
| + // branch already performs GPU-side expert compaction (ggml_cuda_launch_mm_ids_helper, |
| + // mmid.cuh) entirely on-stream and handles arbitrary n_tokens, so keep f16 MoE on the |
| + // fast fused MMF path for any batch size. Upstream gate preserved (disabled) for ref; |
| + // re-enable with -DOPENCOTI_MMF_LEGACY_GATE. |
| +#ifdef OPENCOTI_MMF_LEGACY_GATE |
| if (src0_ne[1] <= 1024 && src1_ncols > 512) { |
| return false; |
| } else if(src0_ne[1] > 1024 && src1_ncols > 128) { |
| return false; |
| } |
| +#endif |
| } else { |
| if (GGML_CUDA_CC_IS_RDNA3_0(cc) && src1_ncols > 8) { |
| return false; |
|
|