File size: 29,654 Bytes
b69d9d8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 | opencoti F5-opt W3 (#292) β op-level fused MoE up+gate+GLU (GGML_OP_MOE_FUSED_UP_GATE)
Adds a new ggml op, GGML_OP_MOE_FUSED_UP_GATE, that collapses a MoE FFN's two
separate per-expert projections (up = up_exps @ cur, gate = gate_exps @ cur)
plus the GLU activation into ONE op at decode (n_tokens == 1). The CUDA backend
dispatches it straight to the existing fused mmvq kernel (ggml_cuda_mul_mat_vec_q
with a fusion arg carrying the gate weight + glu_op), eliminating the gate_up
intermediate's HBM round-trip and the standalone GLU kernel launch per MoE layer
per decode token.
Off by default. Engaged with --fused-moe-up-gate on (server) / cparams
.fused_moe_up_gate / adapter fusedMoeUpGate. The build_moe_ffn graph hook only
emits the op when the model presents SEPARATE up/gate expert weights
(gate_exps && !gate_up_exps), same type, SILU|GELU activation, at decode β i.e.
Qwen2-MoE / Qwen3-MoE / OLMoE / Mixtral-style layouts. Models that fuse gate+up
at the weight level (a single ffn_gate_up_exps β Gemma-4 A4B, Qwen3.5-MoE)
never match the hook and run unchanged. Flag-OFF is byte-identical to the pre-W3
binary (op never emitted); flag-ON on a matching model is byte-identical to the
unfused two-matmul + GLU path. Validated on OLMoE-1B-7B-0924-Instruct-Q4_K_M:
graph-dump engagement confirmed (34 MOE_FUSED_UP_GATE nodes on CUDA0), greedy
decode byte-identical OFF vs ON, +2.4% decode tok/s (265.9 -> 272.2).
opencoti-hook: f5-opt-fmoe β see docs/features/advanced_kv.md
UPSTREAM PROVENANCE (for future re-sync)
Idea source : ik_llama.cpp op-level MoE fusion
https://github.com/ikawrakow/ik_llama.cpp
PR #229 (fused MoE mul_mat_id path), PR #520 (OOAE expert eval).
Commit context HEAD 8960c5ba5ee9db30ba838304373aa4dbec9f7cbd.
License MIT, Author Iwan Kawrakow.
Code vendored: NONE. This op is original opencoti code. ik_llama's
ggml_cuda_moe_up_gate_unary dispatches a fused mmvq-id kernel;
our tree already carries the equivalent primitive
(ggml_cuda_mul_mat_vec_q with ggml_cuda_mm_fusion_args.gate +
glu_op β ggml-cuda/mmvq.cu:555-582). W3 takes only the
op-level-fusion IDEA and routes our existing kernel through a
new ggml op + a build_moe_ffn graph hook. No ik_llama source is
copied into this patch.
ABI note : GGML_OP_COUNT 96 -> 97. The cosmocc binary AND the ggml-cuda.so
DSO must be rebuilt together and paired (asymmetric DSO cache
hygiene). The CPU backend deliberately aborts the op (CUDA-only;
CPU MoE fusion is out of scope β only the CPU embedder matters
for CPU speed, and that is the separate W2 iqk FA engine).
Op convention (locked from ggml-cuda/mmvq.cu):
out = up_proj(src0=up_exps @ src1=cur) (.) act(gate_proj=fusion.gate @ cur)
== ggml_geglu_split(gate, up); the kernel reads the GLU variant from
fusion.glu_op (NOT dst op_params). New op srcs: src0=up_exps, src1=gate_exps,
src2=cur, src3=ids; op_params[0]=glu_op. CUDA dispatch passes src0=up_exps,
src1=cur(=src2), ids=src3, fusion.gate=gate_exps(=src1), glu_op=op_params[0].
Files touched (all surgical; no new vendored files):
ggml/include/ggml.h enum GGML_OP_MOE_FUSED_UP_GATE + ggml_moe_up_gate decl
ggml/src/ggml.c NAME/SYMBOL entries, COUNT 96->97 asserts, creation fn
ggml/src/ggml-cpu/ggml-cpu.c forward-dispatch + n_tasks: CUDA-only ABORT stub
ggml/src/ggml-cuda/ggml-cuda.cu eval dispatch (reuse fused mmvq) + supports_op gate
src/llama-cparams.h bool fused_moe_up_gate
include/llama.h bool fused_moe_up_gate (llama_context_params)
src/llama-context.cpp param copy + default-params initializer
common/common.h bool fused_moe_up_gate = false
common/common.cpp cparams copy
common/arg.cpp --fused-moe-up-gate on|off (LLAMA_ARG_FUSED_MOE_UP_GATE)
src/llama-graph.cpp build_moe_ffn hook: emit op for separate-up/gate decode
---
opencoti note (2026-05-31, bug-250): hunks regenerated via snapshot-diff against the post-0040 chain so the patch strict git-applies (3 common/* hunks had whitespace context drift). Content unchanged.
diff --git a/llama.cpp/common/arg.cpp b/llama.cpp/common/arg.cpp
--- a/llama.cpp/common/arg.cpp
+++ b/llama.cpp/common/arg.cpp
@@ -1375,6 +1375,91 @@ common_params_context common_params_parser_init(common_params & params, llama_ex
params.slot_initial_ctx = value;
}
).set_env("LLAMA_ARG_SLOT_INITIAL_CTX").set_examples({LLAMA_EXAMPLE_SERVER, LLAMA_EXAMPLE_EMBEDDING, LLAMA_EXAMPLE_PARALLEL}));
+ // opencoti F5 M1 rest-kv-eviction β see docs/features/advanced_kv.md
+ add_opt(common_arg(
+ {"--rest-kv-eviction"},
+ {"--no-rest-kv-eviction"},
+ string_format("when context shift fires, evict the lowest-retention (KeyDiff key-similarity) contiguous window instead of the positional middle (default: %s)", params.rest_kv_eviction ? "enabled" : "disabled"),
+ [](common_params & params, bool value) {
+ params.rest_kv_eviction = value;
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_REST_KV_EVICTION"));
+ add_opt(common_arg(
+ {"--rest-kv-recent"}, "N",
+ string_format("rest-kv: protect the most-recent N tokens from eviction (default: %d)", params.rest_kv_recent),
+ [](common_params & params, int value) {
+ params.rest_kv_recent = value;
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_REST_KV_RECENT"));
+ add_opt(common_arg(
+ {"--rest-kv-layer"}, "N",
+ string_format("rest-kv: model layer whose keys score retention, -1 = auto/middle (default: %d)", params.rest_kv_layer),
+ [](common_params & params, int value) {
+ params.rest_kv_layer = value;
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_REST_KV_LAYER"));
+ // opencoti F5 M2 headinfer β see docs/features/advanced_kv.md
+ add_opt(common_arg(
+ {"--headinfer-gpu-heads-frac"}, "F",
+ string_format("headinfer: fraction of KV heads kept GPU-resident; the rest are offloaded to host memory and streamed back per step (default: %.2f, 1.0 = off). Only effective for GPU-offloaded layers with flash-attention.", (double) params.headinfer_gpu_heads_frac),
+ [](common_params & params, const std::string & value) {
+ params.headinfer_gpu_heads_frac = std::stof(value);
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_HEADINFER_GPU_HEADS_FRAC"));
+ // opencoti F5 M3 neo β see docs/features/advanced_kv.md
+ add_opt(common_arg(
+ {"--neo-pipeline"}, "MODE",
+ "neo (asymmetric GPU/CPU attention pipelining): off (default), on (engage when headinfer split is active for the layer), auto (reserved, same as on). Requires --headinfer-gpu-heads-frac < 1.0 to do anything.",
+ [](common_params & params, const std::string & value) {
+ if (value == "off") params.neo_pipeline_mode = 0;
+ else if (value == "on") params.neo_pipeline_mode = 1;
+ else if (value == "auto") params.neo_pipeline_mode = 2;
+ else throw std::invalid_argument("--neo-pipeline must be off|on|auto");
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_NEO_PIPELINE"));
+ // opencoti F5-opt W2 (#290) iqk CPU flash-attention β see docs/features/advanced_kv.md
+ add_opt(common_arg(
+ {"--iqk-flash-attn"}, "MODE",
+ "iqk CPU flash-attention (vendored ik_llama.cpp engine): off (default), on. "
+ "Speeds up the M2 CPU-half attention on the token-generation path; falls back "
+ "to the mainline kernel for unsupported (type, shape) cases. x86_64-only.",
+ [](common_params & params, const std::string & value) {
+ if (value == "off" || value == "false" || value == "0") params.iqk_flash_attn = false;
+ else if (value == "on" || value == "true" || value == "1") params.iqk_flash_attn = true;
+ else throw std::invalid_argument("--iqk-flash-attn must be on|off");
+ ggml_iqk_flash_attn_set_enabled(params.iqk_flash_attn);
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_IQK_FLASH_ATTN"));
+ // opencoti F5-opt W3 (#292) fused-MoE β see docs/features/advanced_kv.md
+ add_opt(common_arg(
+ {"--fused-moe-up-gate"}, "MODE",
+ "fused MoE up+gate+GLU (CUDA decode): off (default), on. Collapses the "
+ "separate-up/gate expert matmuls + GLU into one op at decode so the CUDA "
+ "backend runs the fused mmvq kernel without graph-pattern fusion. Only "
+ "engages for separate-up/gate MoE models (e.g. Gemma-4 A4B) at n_tokens==1.",
+ [](common_params & params, const std::string & value) {
+ if (value == "off" || value == "false" || value == "0") params.fused_moe_up_gate = false;
+ else if (value == "on" || value == "true" || value == "1") params.fused_moe_up_gate = true;
+ else throw std::invalid_argument("--fused-moe-up-gate must be on|off");
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_FUSED_MOE_UP_GATE"));
+ // opencoti F5-opt W1 (#293) pcie profile β see docs/features/rolling_kv.md
+ add_opt(common_arg(
+ {"--pcie-autodetect"}, "MODE",
+ "pcie profile: auto-detect host<->device bandwidth + PCIe link at boot (on (default), off). Reads the rebar-probe JSON / nvidia-smi; the result feeds M7 Rolling KV tile sizing.",
+ [](common_params & params, const std::string & value) {
+ if (value == "on" || value == "true" || value == "1") params.pcie_autodetect = true;
+ else if (value == "off" || value == "false" || value == "0") params.pcie_autodetect = false;
+ else throw std::invalid_argument("--pcie-autodetect must be on|off");
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_PCIE_AUTODETECT"));
+ add_opt(common_arg(
+ {"--pcie-bw-gbps"}, "F",
+ string_format("pcie profile: manual override of effective host->device bandwidth in GB/s (default: %.1f = autodetect). > 0 skips detection.", (double) params.pcie_bw_gbps),
+ [](common_params & params, const std::string & value) {
+ params.pcie_bw_gbps = std::stof(value);
+ }
+ ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_PCIE_BW_GBPS"));
add_opt(common_arg(
// opencoti F4 M3 Phase 3 β see docs/decisions/0001-lazy-slot-context.md
{"--slot-shrink-idle-ms"}, "N",
diff --git a/llama.cpp/common/common.cpp b/llama.cpp/common/common.cpp
--- a/llama.cpp/common/common.cpp
+++ b/llama.cpp/common/common.cpp
@@ -1636,6 +1636,8 @@ struct llama_context_params common_context_params_to_llama(const common_params &
cparams.headinfer_gpu_heads_frac = params.headinfer_gpu_heads_frac;
// opencoti F5 M3 neo β see docs/features/advanced_kv.md
cparams.neo_pipeline_mode = params.neo_pipeline_mode;
+ // opencoti F5 W3 #292 fused-MoE β see docs/features/advanced_kv.md
+ cparams.fused_moe_up_gate = params.fused_moe_up_gate;
cparams.n_rs_seq = params.speculative.need_n_rs_seq();
cparams.n_batch = params.n_batch;
cparams.n_ubatch = params.n_ubatch;
diff --git a/llama.cpp/common/common.h b/llama.cpp/common/common.h
--- a/llama.cpp/common/common.h
+++ b/llama.cpp/common/common.h
@@ -445,6 +445,41 @@ struct common_params {
// active for the layer), 2 = auto (reserved for future load-based
// engagement; same as `on` today).
uint32_t neo_pipeline_mode = 0;
+ // opencoti F5 W3 #292 fused-MoE β see docs/features/advanced_kv.md
+ // false = off (default; unfused 3-node path). true = fuse decode MoE
+ // up+gate+GLU into one GGML_OP_MOE_FUSED_UP_GATE op for the CUDA backend.
+ bool fused_moe_up_gate = false;
+ // opencoti F5-opt W1 (#293) pcie profile β see docs/features/rolling_kv.md
+ // Boot-time host<->device bandwidth + PCIe link detection, consumed by M7
+ // Rolling KV (#296) tile sizing. autodetect reads the rebar-probe JSON /
+ // nvidia-smi; bw_gbps > 0 forces a manual override (0 = autodetect/default).
+ bool pcie_autodetect = true;
+ // opencoti F5-opt W2 (#290) iqk CPU flash-attention β see docs/features/advanced_kv.md
+ // Engage the vendored ik_llama.cpp CPU Flash-Attention engine on the M2
+ // CPU-half (token-generation path). Off by default; flag-off is byte-identical.
+ // x86_64-only at runtime (the engine compiles only on the AVX2 x86_64 slice).
+ bool iqk_flash_attn = false;
+ float pcie_bw_gbps = 0.0f;
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Initial KV-cell allocation per sequence. 0 = full n_ctx_seq allocation
+ // (Phase 1 behavior β only zero-fill is deferred). Positive values cap
+ // the initial allocation and enable grow-on-demand up to n_ctx_seq.
+ int32_t slot_initial_ctx = 0;
+ // opencoti F4 M3 Phase 3 β see docs/decisions/0001-lazy-slot-context.md
+ // Idle threshold (ms) after which a grown soft cap is shrunk back to
+ // slot_initial_ctx and the over-cap pages are decommitted via
+ // posix_madvise. 0 disables shrink-on-idle (Phase 2 behavior).
+ int32_t slot_shrink_idle_ms = 0;
+ // opencoti F5 M2 headinfer β see docs/features/advanced_kv.md
+ // Fraction of KV heads kept GPU-resident; the rest are offloaded to host
+ // memory and streamed back per step. 1.0 = off (no split). Only active for
+ // GPU-offloaded layers with flash-attention (non-transposed V cache).
+ float headinfer_gpu_heads_frac = 1.0f;
+ // opencoti F5 M3 neo β see docs/features/advanced_kv.md
+ // 0 = off (default; M2 concat path runs), 1 = on (engage when split
+ // active for the layer), 2 = auto (reserved for future load-based
+ // engagement; same as `on` today).
+ uint32_t neo_pipeline_mode = 0;
// opencoti F5-opt W1 (#293) pcie profile β see docs/features/rolling_kv.md
// Boot-time host<->device bandwidth + PCIe link detection, consumed by M7
// Rolling KV (#296) tile sizing. autodetect reads the rebar-probe JSON /
diff --git a/llama.cpp/ggml/include/ggml.h b/llama.cpp/ggml/include/ggml.h
--- a/llama.cpp/ggml/include/ggml.h
+++ b/llama.cpp/ggml/include/ggml.h
@@ -583,6 +583,8 @@ extern "C" {
GGML_OP_GLU,
+ GGML_OP_MOE_FUSED_UP_GATE, // opencoti-hook: f5-opt-fmoe β W3 #292
+
GGML_OP_COUNT,
};
@@ -1437,6 +1439,22 @@ extern "C" {
struct ggml_tensor * b,
struct ggml_tensor * ids);
+ // opencoti-hook: f5-opt-fmoe β W3 #292: op-level fused MoE up+gate+GLU.
+ // build_moe_ffn otherwise emits {mul_mat_id(up), mul_mat_id(gate), glu}
+ // as 3 nodes that the CUDA backend can only fuse via the fragile
+ // ggml_can_fuse_subgraph graph-pattern path (which does not engage at
+ // decode for Gemma-4). This single op lets the backend run the already
+ // existing fused mmvq kernel directly. Output is GLU(gate@cur) (.) (up@cur),
+ // matching ggml_geglu_split(gate, up). up_exps/gate_exps: expert weight
+ // stacks (same type); cur: input activations; ids: expert selection.
+ GGML_API struct ggml_tensor * ggml_moe_up_gate(
+ struct ggml_context * ctx,
+ struct ggml_tensor * up_exps,
+ struct ggml_tensor * gate_exps,
+ struct ggml_tensor * cur,
+ struct ggml_tensor * ids,
+ enum ggml_glu_op glu_op);
+
// A: m columns, n rows,
// B: p columns, n rows,
// result is m columns, p rows
diff --git a/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c b/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c
--- a/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c
+++ b/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c
@@ -2066,6 +2066,13 @@ static void ggml_compute_forward(struct ggml_compute_params * params, struct ggm
{
ggml_compute_forward_glu(params, tensor);
} break;
+ case GGML_OP_MOE_FUSED_UP_GATE:
+ {
+ // opencoti-hook: f5-opt-fmoe β W3 #292. CUDA-only op. The
+ // build_moe_ffn graph hook emits it only when the CUDA backend
+ // owns the FFN nodes, so the CPU scheduler never reaches here.
+ GGML_ABORT("MOE_FUSED_UP_GATE has no CPU implementation (CUDA-only op)");
+ } break;
case GGML_OP_GET_REL_POS:
{
ggml_compute_forward_get_rel_pos(params, tensor);
@@ -2333,6 +2340,12 @@ static int ggml_get_n_tasks(struct ggml_tensor * node, int n_threads) {
GGML_ABORT("fatal error");
}
break;
+ case GGML_OP_MOE_FUSED_UP_GATE:
+ {
+ // opencoti-hook: f5-opt-fmoe β W3 #292. CUDA-only; never
+ // scheduled on CPU. Single task keeps the planner well-formed.
+ n_tasks = 1;
+ } break;
case GGML_OP_SILU_BACK:
case GGML_OP_MUL:
case GGML_OP_DIV:
diff --git a/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu b/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu
--- a/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu
+++ b/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu
@@ -3005,6 +3005,25 @@ static bool ggml_cuda_compute_forward(ggml_backend_cuda_context & ctx, struct gg
case GGML_OP_MUL_MAT_ID:
ggml_cuda_mul_mat_id(ctx, dst);
break;
+ case GGML_OP_MOE_FUSED_UP_GATE:
+ {
+ // opencoti-hook: f5-opt-fmoe β W3 #292. Drive the existing fused
+ // mmvq-id kernel directly: up = src0 @ src1, gate = fusion.gate @
+ // src1, out = act(gate) (.) up. Byte-identical to the graph-pattern
+ // fusion path (see the { MUL_MAT_ID, MUL_MAT_ID, GLU } branch in
+ // the eval fusion loop). Decode-only (ne[2]==1): the build_moe_ffn
+ // hook only emits this at n_tokens==1 and supports_op enforces it.
+ const ggml_tensor * up_exps = dst->src[0];
+ const ggml_tensor * gate_exps = dst->src[1];
+ const ggml_tensor * cur = dst->src[2];
+ const ggml_tensor * ids = dst->src[3];
+ GGML_ASSERT(dst->ne[2] == 1 && "MOE_FUSED_UP_GATE is decode-only (n_tokens==1)");
+ ggml_cuda_mm_fusion_args_host fusion_data{};
+ fusion_data.gate = gate_exps;
+ fusion_data.glu_op = (ggml_glu_op) dst->op_params[0];
+ ggml_cuda_mul_mat_vec_q(ctx, up_exps, cur, ids, dst, &fusion_data);
+ }
+ break;
case GGML_OP_OUT_PROD:
ggml_cuda_out_prod(ctx, dst);
break;
@@ -5210,6 +5229,44 @@ static bool GGML_CALL ggml_backend_cuda_device_supports_op(ggml_backend_dev_t de
return false;
}
break;
+ case GGML_OP_MOE_FUSED_UP_GATE:
+ {
+ // opencoti-hook: f5-opt-fmoe β W3 #292. Supported only in the
+ // exact shape the fused mmvq-id decode kernel handles: quantized
+ // up/gate expert stacks of identical type, F32 activations, I32
+ // ids, single-token (decode) output, non-split weights, and a GLU
+ // op the kernel implements explicitly (GEGLU/SWIGLU/SWIGLU_OAI β
+ // others silently degrade to a plain product). build_moe_ffn only
+ // emits the op when all of these hold, so supports_op and the hook
+ // stay consistent (the scheduler never strands it on the CPU).
+ const ggml_tensor * up_exps = op->src[0];
+ const ggml_tensor * gate_exps = op->src[1];
+ const ggml_tensor * cur = op->src[2];
+ const ggml_tensor * ids = op->src[3];
+ if (!up_exps || !gate_exps || !cur || !ids) {
+ return false;
+ }
+ if (up_exps->buffer && ggml_backend_buft_is_cuda_split(up_exps->buffer->buft)) {
+ return false;
+ }
+ if (!ggml_is_quantized(up_exps->type) || up_exps->type != gate_exps->type) {
+ return false;
+ }
+ if (cur->type != GGML_TYPE_F32 || ids->type != GGML_TYPE_I32) {
+ return false;
+ }
+ if (op->ne[2] != 1) {
+ return false;
+ }
+ switch ((ggml_glu_op) op->op_params[0]) {
+ case GGML_GLU_OP_GEGLU:
+ case GGML_GLU_OP_SWIGLU:
+ case GGML_GLU_OP_SWIGLU_OAI:
+ return true;
+ default:
+ return false;
+ }
+ }
case GGML_OP_MUL_MAT:
case GGML_OP_MUL_MAT_ID:
{
diff --git a/llama.cpp/ggml/src/ggml.c b/llama.cpp/ggml/src/ggml.c
--- a/llama.cpp/ggml/src/ggml.c
+++ b/llama.cpp/ggml/src/ggml.c
@@ -1078,9 +1078,10 @@ static const char * GGML_OP_NAME[GGML_OP_COUNT] = {
"OPT_STEP_SGD",
"GLU",
+ "MOE_FUSED_UP_GATE",
};
-static_assert(GGML_OP_COUNT == 96, "GGML_OP_COUNT != 96");
+static_assert(GGML_OP_COUNT == 97, "GGML_OP_COUNT != 97");
static const char * GGML_OP_SYMBOL[GGML_OP_COUNT] = {
"none",
@@ -1188,9 +1189,10 @@ static const char * GGML_OP_SYMBOL[GGML_OP_COUNT] = {
"sgd(x)",
"glu(x)",
+ "moe_up_gate(up,gate,x)",
};
-static_assert(GGML_OP_COUNT == 96, "GGML_OP_COUNT != 96");
+static_assert(GGML_OP_COUNT == 97, "GGML_OP_COUNT != 97");
static_assert(GGML_OP_POOL_COUNT == 2, "GGML_OP_POOL_COUNT != 2");
@@ -3314,6 +3316,53 @@ struct ggml_tensor * ggml_mul_mat_id(
return result;
}
+// ggml_moe_up_gate
+// opencoti-hook: f5-opt-fmoe β W3 #292. Op-level fusion of the two parallel
+// expert matmuls (up = up_exps @ cur, gate = gate_exps @ cur) plus the GLU
+// activation that build_moe_ffn otherwise emits as 3 separate graph nodes.
+// Output == ggml_geglu_split(gate, up) == act(gate) (.) up, where the CUDA
+// backend runs the existing fused mmvq kernel (act on the gate projection,
+// multiply by the up projection β see ggml-cuda/mmvq.cu). Mirrors
+// ggml_mul_mat_id's shape contract: up_exps is the primary "as" weight, cur
+// the "b" activations shared by both projections, ids the expert selection.
+
+struct ggml_tensor * ggml_moe_up_gate(
+ struct ggml_context * ctx,
+ struct ggml_tensor * up_exps,
+ struct ggml_tensor * gate_exps,
+ struct ggml_tensor * cur,
+ struct ggml_tensor * ids,
+ enum ggml_glu_op glu_op) {
+ GGML_ASSERT(!ggml_is_transposed(up_exps));
+ GGML_ASSERT(!ggml_is_transposed(gate_exps));
+ GGML_ASSERT(ids->type == GGML_TYPE_I32);
+
+ // up and gate are parallel expert stacks: identical shape + type so one
+ // shared-activation fused mmvq can drive both weight sets.
+ GGML_ASSERT(ggml_are_same_shape(up_exps, gate_exps));
+ GGML_ASSERT(up_exps->type == gate_exps->type);
+
+ GGML_ASSERT(up_exps->ne[3] == 1); // 3d (one matrix per expert)
+ GGML_ASSERT(cur->ne[3] == 1); // cur is 3d
+ GGML_ASSERT(ids->ne[2] == 1 && ids->ne[3] == 1); // ids is 2d
+ GGML_ASSERT(ids->ne[1] == cur->ne[2]); // one expert list per cur row
+ GGML_ASSERT(up_exps->ne[0] == cur->ne[0]); // can_mul_mat
+ GGML_ASSERT(ids->ne[0] % cur->ne[1] == 0); // can broadcast
+
+ const int64_t ne[4] = { up_exps->ne[1], ids->ne[0], cur->ne[2], 1 };
+ struct ggml_tensor * result = ggml_new_tensor(ctx, GGML_TYPE_F32, 4, ne);
+
+ ggml_set_op_params_i32(result, 0, (int32_t) glu_op);
+
+ result->op = GGML_OP_MOE_FUSED_UP_GATE;
+ result->src[0] = up_exps;
+ result->src[1] = gate_exps;
+ result->src[2] = cur;
+ result->src[3] = ids;
+
+ return result;
+}
+
// ggml_out_prod
static inline bool ggml_can_out_prod(const struct ggml_tensor * t0, const struct ggml_tensor * t1) {
diff --git a/llama.cpp/include/llama.h b/llama.cpp/include/llama.h
--- a/llama.cpp/include/llama.h
+++ b/llama.cpp/include/llama.h
@@ -362,6 +362,9 @@ extern "C" {
// 0 = off (default), 1 = on, 2 = auto. Effective only when the
// M2 head split is active for a given layer.
uint32_t neo_pipeline_mode;
+ // opencoti F5 W3 #292 fused-MoE β see docs/features/advanced_kv.md
+ // true = fuse decode MoE up+gate+GLU into one op. false = off (default).
+ bool fused_moe_up_gate;
uint32_t n_rs_seq; // number of recurrent-state snapshots per seq for rollback (0 = no rollback) [EXPERIMENTAL]
int32_t n_threads; // number of threads to use for generation
int32_t n_threads_batch; // number of threads to use for batch processing
diff --git a/llama.cpp/src/llama-context.cpp b/llama.cpp/src/llama-context.cpp
--- a/llama.cpp/src/llama-context.cpp
+++ b/llama.cpp/src/llama-context.cpp
@@ -66,6 +66,8 @@ llama_context::llama_context(
cparams.headinfer_gpu_heads_frac = params.headinfer_gpu_heads_frac;
// opencoti F5 M3 neo β see docs/features/advanced_kv.md
cparams.neo_pipeline_mode = params.neo_pipeline_mode;
+ // opencoti F5 W3 #292 fused-MoE β see docs/features/advanced_kv.md
+ cparams.fused_moe_up_gate = params.fused_moe_up_gate;
cparams.n_threads = params.n_threads;
cparams.n_threads_batch = params.n_threads_batch;
@@ -3359,6 +3361,7 @@ llama_context_params llama_context_default_params() {
/*.slot_shrink_idle_ms =*/ 0,
/*.headinfer_gpu_heads_frac =*/ 1.0f,
/*.neo_pipeline_mode =*/ 0, // opencoti F5 M3 neo β off by default
+ /*.fused_moe_up_gate =*/ false, // opencoti F5 W3 #292 β off by default
/*.n_rs_seq =*/ 0,
/*.n_threads =*/ GGML_DEFAULT_N_THREADS, // TODO: better default
/*.n_threads_batch =*/ GGML_DEFAULT_N_THREADS,
diff --git a/llama.cpp/src/llama-cparams.h b/llama.cpp/src/llama-cparams.h
--- a/llama.cpp/src/llama-cparams.h
+++ b/llama.cpp/src/llama-cparams.h
@@ -32,6 +32,10 @@ struct llama_cparams {
// 2 = auto (same engagement as `on` for now β reserved for future
// load-based engagement heuristics).
uint32_t neo_pipeline_mode;
+ // opencoti F5 W3 #292 fused-MoE β see docs/features/advanced_kv.md.
+ // true = build_moe_ffn collapses the up/gate expert matmuls + GLU into one
+ // GGML_OP_MOE_FUSED_UP_GATE op at decode (n_tokens==1). false = unfused path.
+ bool fused_moe_up_gate;
uint32_t n_rs_seq; // number of recurrent-state snapshots per seq for rollback
int32_t n_threads; // number of threads to use for generation
int32_t n_threads_batch; // number of threads to use for batch processing
diff --git a/llama.cpp/src/llama-graph.cpp b/llama.cpp/src/llama-graph.cpp
--- a/llama.cpp/src/llama-graph.cpp
+++ b/llama.cpp/src/llama-graph.cpp
@@ -1534,7 +1534,28 @@ ggml_tensor * llm_graph_context::build_moe_ffn(
ggml_tensor * up = nullptr;
ggml_tensor * experts = nullptr;
- if (gate_up_exps) {
+ // opencoti-hook: f5-opt-fmoe β W3 #292: collapse the separate-up/gate
+ // expert matmuls + the GLU activation into ONE GGML_OP_MOE_FUSED_UP_GATE
+ // op at decode, so the CUDA backend runs the existing fused mmvq kernel
+ // without relying on graph-pattern fusion (ggml_can_fuse_subgraph, which
+ // does not engage for Gemma-4 at decode). Guards mirror supports_op:
+ // decode only (n_tokens==1), separate up/gate of identical type, no
+ // bias/scale, and GELUβGEGLU or SILUβSWIGLU. Off path is byte-identical.
+ bool fused_up_gate = false;
+ if (cparams.fused_moe_up_gate && gate_exps && !gate_up_exps &&
+ !up_exps_b && !gate_exps_b && !up_exps_s && !gate_exps_s &&
+ up_exps->type == gate_exps->type && n_tokens == 1 &&
+ (type_op == LLM_FFN_GELU || type_op == LLM_FFN_SILU)) {
+ const enum ggml_glu_op glu_op = (type_op == LLM_FFN_GELU)
+ ? GGML_GLU_OP_GEGLU : GGML_GLU_OP_SWIGLU;
+ cur = ggml_moe_up_gate(ctx0, up_exps, gate_exps, cur, selected_experts, glu_op);
+ cb(cur, "ffn_moe_fused_up_gate", il);
+ fused_up_gate = true;
+ }
+
+ if (fused_up_gate) {
+ // cur already holds act(gate) (.) up from the fused op
+ } else if (gate_up_exps) {
// merged gate_up path: one mul_mat_id, then split into gate and up views
ggml_tensor * gate_up = build_lora_mm_id(gate_up_exps, cur, selected_experts); // [n_ff*2, n_expert_used, n_tokens]
cb(gate_up, "ffn_moe_gate_up", il);
@@ -1601,6 +1622,9 @@ ggml_tensor * llm_graph_context::build_moe_ffn(
const bool has_gate = gate_exps || gate_up_exps;
+ // opencoti-hook: f5-opt-fmoe β W3 #292: the fused op already applied the
+ // GLU activation; skip the separate activation pass.
+ if (!fused_up_gate)
switch (type_op) {
case LLM_FFN_SILU:
if (gate_exps) {
|