File size: 29,654 Bytes
b69d9d8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
opencoti F5-opt W3 (#292) β€” op-level fused MoE up+gate+GLU (GGML_OP_MOE_FUSED_UP_GATE)

Adds a new ggml op, GGML_OP_MOE_FUSED_UP_GATE, that collapses a MoE FFN's two
separate per-expert projections (up = up_exps @ cur, gate = gate_exps @ cur)
plus the GLU activation into ONE op at decode (n_tokens == 1). The CUDA backend
dispatches it straight to the existing fused mmvq kernel (ggml_cuda_mul_mat_vec_q
with a fusion arg carrying the gate weight + glu_op), eliminating the gate_up
intermediate's HBM round-trip and the standalone GLU kernel launch per MoE layer
per decode token.

Off by default. Engaged with --fused-moe-up-gate on (server) / cparams
.fused_moe_up_gate / adapter fusedMoeUpGate. The build_moe_ffn graph hook only
emits the op when the model presents SEPARATE up/gate expert weights
(gate_exps && !gate_up_exps), same type, SILU|GELU activation, at decode β€” i.e.
Qwen2-MoE / Qwen3-MoE / OLMoE / Mixtral-style layouts. Models that fuse gate+up
at the weight level (a single ffn_gate_up_exps β€” Gemma-4 A4B, Qwen3.5-MoE)
never match the hook and run unchanged. Flag-OFF is byte-identical to the pre-W3
binary (op never emitted); flag-ON on a matching model is byte-identical to the
unfused two-matmul + GLU path. Validated on OLMoE-1B-7B-0924-Instruct-Q4_K_M:
graph-dump engagement confirmed (34 MOE_FUSED_UP_GATE nodes on CUDA0), greedy
decode byte-identical OFF vs ON, +2.4% decode tok/s (265.9 -> 272.2).

  opencoti-hook: f5-opt-fmoe β€” see docs/features/advanced_kv.md

UPSTREAM PROVENANCE (for future re-sync)
  Idea source : ik_llama.cpp op-level MoE fusion
                https://github.com/ikawrakow/ik_llama.cpp
                PR #229 (fused MoE mul_mat_id path), PR #520 (OOAE expert eval).
                Commit context HEAD 8960c5ba5ee9db30ba838304373aa4dbec9f7cbd.
                License MIT, Author Iwan Kawrakow.
  Code vendored: NONE. This op is original opencoti code. ik_llama's
                ggml_cuda_moe_up_gate_unary dispatches a fused mmvq-id kernel;
                our tree already carries the equivalent primitive
                (ggml_cuda_mul_mat_vec_q with ggml_cuda_mm_fusion_args.gate +
                glu_op β€” ggml-cuda/mmvq.cu:555-582). W3 takes only the
                op-level-fusion IDEA and routes our existing kernel through a
                new ggml op + a build_moe_ffn graph hook. No ik_llama source is
                copied into this patch.
  ABI note    : GGML_OP_COUNT 96 -> 97. The cosmocc binary AND the ggml-cuda.so
                DSO must be rebuilt together and paired (asymmetric DSO cache
                hygiene). The CPU backend deliberately aborts the op (CUDA-only;
                CPU MoE fusion is out of scope β€” only the CPU embedder matters
                for CPU speed, and that is the separate W2 iqk FA engine).
  Op convention (locked from ggml-cuda/mmvq.cu):
    out = up_proj(src0=up_exps @ src1=cur) (.) act(gate_proj=fusion.gate @ cur)
        == ggml_geglu_split(gate, up); the kernel reads the GLU variant from
    fusion.glu_op (NOT dst op_params). New op srcs: src0=up_exps, src1=gate_exps,
    src2=cur, src3=ids; op_params[0]=glu_op. CUDA dispatch passes src0=up_exps,
    src1=cur(=src2), ids=src3, fusion.gate=gate_exps(=src1), glu_op=op_params[0].

Files touched (all surgical; no new vendored files):
  ggml/include/ggml.h            enum GGML_OP_MOE_FUSED_UP_GATE + ggml_moe_up_gate decl
  ggml/src/ggml.c                NAME/SYMBOL entries, COUNT 96->97 asserts, creation fn
  ggml/src/ggml-cpu/ggml-cpu.c   forward-dispatch + n_tasks: CUDA-only ABORT stub
  ggml/src/ggml-cuda/ggml-cuda.cu eval dispatch (reuse fused mmvq) + supports_op gate
  src/llama-cparams.h            bool fused_moe_up_gate
  include/llama.h                bool fused_moe_up_gate (llama_context_params)
  src/llama-context.cpp          param copy + default-params initializer
  common/common.h                bool fused_moe_up_gate = false
  common/common.cpp              cparams copy
  common/arg.cpp                 --fused-moe-up-gate on|off (LLAMA_ARG_FUSED_MOE_UP_GATE)
  src/llama-graph.cpp            build_moe_ffn hook: emit op for separate-up/gate decode

---
opencoti note (2026-05-31, bug-250): hunks regenerated via snapshot-diff against the post-0040 chain so the patch strict git-applies (3 common/* hunks had whitespace context drift). Content unchanged.
diff --git a/llama.cpp/common/arg.cpp b/llama.cpp/common/arg.cpp
--- a/llama.cpp/common/arg.cpp
+++ b/llama.cpp/common/arg.cpp
@@ -1375,6 +1375,91 @@ common_params_context common_params_parser_init(common_params & params, llama_ex
             params.slot_initial_ctx = value;
         }
     ).set_env("LLAMA_ARG_SLOT_INITIAL_CTX").set_examples({LLAMA_EXAMPLE_SERVER, LLAMA_EXAMPLE_EMBEDDING, LLAMA_EXAMPLE_PARALLEL}));
+    // opencoti F5 M1 rest-kv-eviction β€” see docs/features/advanced_kv.md
+    add_opt(common_arg(
+        {"--rest-kv-eviction"},
+        {"--no-rest-kv-eviction"},
+        string_format("when context shift fires, evict the lowest-retention (KeyDiff key-similarity) contiguous window instead of the positional middle (default: %s)", params.rest_kv_eviction ? "enabled" : "disabled"),
+        [](common_params & params, bool value) {
+            params.rest_kv_eviction = value;
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_REST_KV_EVICTION"));
+    add_opt(common_arg(
+        {"--rest-kv-recent"}, "N",
+        string_format("rest-kv: protect the most-recent N tokens from eviction (default: %d)", params.rest_kv_recent),
+        [](common_params & params, int value) {
+            params.rest_kv_recent = value;
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_REST_KV_RECENT"));
+    add_opt(common_arg(
+        {"--rest-kv-layer"}, "N",
+        string_format("rest-kv: model layer whose keys score retention, -1 = auto/middle (default: %d)", params.rest_kv_layer),
+        [](common_params & params, int value) {
+            params.rest_kv_layer = value;
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_REST_KV_LAYER"));
+    // opencoti F5 M2 headinfer β€” see docs/features/advanced_kv.md
+    add_opt(common_arg(
+        {"--headinfer-gpu-heads-frac"}, "F",
+        string_format("headinfer: fraction of KV heads kept GPU-resident; the rest are offloaded to host memory and streamed back per step (default: %.2f, 1.0 = off). Only effective for GPU-offloaded layers with flash-attention.", (double) params.headinfer_gpu_heads_frac),
+        [](common_params & params, const std::string & value) {
+            params.headinfer_gpu_heads_frac = std::stof(value);
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_HEADINFER_GPU_HEADS_FRAC"));
+    // opencoti F5 M3 neo β€” see docs/features/advanced_kv.md
+    add_opt(common_arg(
+        {"--neo-pipeline"}, "MODE",
+        "neo (asymmetric GPU/CPU attention pipelining): off (default), on (engage when headinfer split is active for the layer), auto (reserved, same as on). Requires --headinfer-gpu-heads-frac < 1.0 to do anything.",
+        [](common_params & params, const std::string & value) {
+            if      (value == "off")  params.neo_pipeline_mode = 0;
+            else if (value == "on")   params.neo_pipeline_mode = 1;
+            else if (value == "auto") params.neo_pipeline_mode = 2;
+            else throw std::invalid_argument("--neo-pipeline must be off|on|auto");
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_NEO_PIPELINE"));
+    // opencoti F5-opt W2 (#290) iqk CPU flash-attention β€” see docs/features/advanced_kv.md
+    add_opt(common_arg(
+        {"--iqk-flash-attn"}, "MODE",
+        "iqk CPU flash-attention (vendored ik_llama.cpp engine): off (default), on. "
+        "Speeds up the M2 CPU-half attention on the token-generation path; falls back "
+        "to the mainline kernel for unsupported (type, shape) cases. x86_64-only.",
+        [](common_params & params, const std::string & value) {
+            if      (value == "off" || value == "false" || value == "0") params.iqk_flash_attn = false;
+            else if (value == "on"  || value == "true"  || value == "1") params.iqk_flash_attn = true;
+            else throw std::invalid_argument("--iqk-flash-attn must be on|off");
+            ggml_iqk_flash_attn_set_enabled(params.iqk_flash_attn);
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_IQK_FLASH_ATTN"));
+    // opencoti F5-opt W3 (#292) fused-MoE β€” see docs/features/advanced_kv.md
+    add_opt(common_arg(
+        {"--fused-moe-up-gate"}, "MODE",
+        "fused MoE up+gate+GLU (CUDA decode): off (default), on. Collapses the "
+        "separate-up/gate expert matmuls + GLU into one op at decode so the CUDA "
+        "backend runs the fused mmvq kernel without graph-pattern fusion. Only "
+        "engages for separate-up/gate MoE models (e.g. Gemma-4 A4B) at n_tokens==1.",
+        [](common_params & params, const std::string & value) {
+            if      (value == "off" || value == "false" || value == "0") params.fused_moe_up_gate = false;
+            else if (value == "on"  || value == "true"  || value == "1") params.fused_moe_up_gate = true;
+            else throw std::invalid_argument("--fused-moe-up-gate must be on|off");
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_FUSED_MOE_UP_GATE"));
+    // opencoti F5-opt W1 (#293) pcie profile β€” see docs/features/rolling_kv.md
+    add_opt(common_arg(
+        {"--pcie-autodetect"}, "MODE",
+        "pcie profile: auto-detect host<->device bandwidth + PCIe link at boot (on (default), off). Reads the rebar-probe JSON / nvidia-smi; the result feeds M7 Rolling KV tile sizing.",
+        [](common_params & params, const std::string & value) {
+            if      (value == "on"  || value == "true"  || value == "1") params.pcie_autodetect = true;
+            else if (value == "off" || value == "false" || value == "0") params.pcie_autodetect = false;
+            else throw std::invalid_argument("--pcie-autodetect must be on|off");
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_PCIE_AUTODETECT"));
+    add_opt(common_arg(
+        {"--pcie-bw-gbps"}, "F",
+        string_format("pcie profile: manual override of effective host->device bandwidth in GB/s (default: %.1f = autodetect). > 0 skips detection.", (double) params.pcie_bw_gbps),
+        [](common_params & params, const std::string & value) {
+            params.pcie_bw_gbps = std::stof(value);
+        }
+    ).set_examples({LLAMA_EXAMPLE_SERVER}).set_env("LLAMA_ARG_PCIE_BW_GBPS"));
     add_opt(common_arg(
         // opencoti F4 M3 Phase 3 β€” see docs/decisions/0001-lazy-slot-context.md
         {"--slot-shrink-idle-ms"}, "N",
diff --git a/llama.cpp/common/common.cpp b/llama.cpp/common/common.cpp
--- a/llama.cpp/common/common.cpp
+++ b/llama.cpp/common/common.cpp
@@ -1636,6 +1636,8 @@ struct llama_context_params common_context_params_to_llama(const common_params &
     cparams.headinfer_gpu_heads_frac = params.headinfer_gpu_heads_frac;
     // opencoti F5 M3 neo β€” see docs/features/advanced_kv.md
     cparams.neo_pipeline_mode = params.neo_pipeline_mode;
+    // opencoti F5 W3 #292 fused-MoE β€” see docs/features/advanced_kv.md
+    cparams.fused_moe_up_gate = params.fused_moe_up_gate;
     cparams.n_rs_seq          = params.speculative.need_n_rs_seq();
     cparams.n_batch           = params.n_batch;
     cparams.n_ubatch          = params.n_ubatch;
diff --git a/llama.cpp/common/common.h b/llama.cpp/common/common.h
--- a/llama.cpp/common/common.h
+++ b/llama.cpp/common/common.h
@@ -445,6 +445,41 @@ struct common_params {
     // active for the layer), 2 = auto (reserved for future load-based
     // engagement; same as `on` today).
     uint32_t neo_pipeline_mode    =     0;
+    // opencoti F5 W3 #292 fused-MoE β€” see docs/features/advanced_kv.md
+    // false = off (default; unfused 3-node path). true = fuse decode MoE
+    // up+gate+GLU into one GGML_OP_MOE_FUSED_UP_GATE op for the CUDA backend.
+    bool     fused_moe_up_gate    = false;
+    // opencoti F5-opt W1 (#293) pcie profile β€” see docs/features/rolling_kv.md
+    // Boot-time host<->device bandwidth + PCIe link detection, consumed by M7
+    // Rolling KV (#296) tile sizing. autodetect reads the rebar-probe JSON /
+    // nvidia-smi; bw_gbps > 0 forces a manual override (0 = autodetect/default).
+    bool    pcie_autodetect       =  true;
+    // opencoti F5-opt W2 (#290) iqk CPU flash-attention β€” see docs/features/advanced_kv.md
+    // Engage the vendored ik_llama.cpp CPU Flash-Attention engine on the M2
+    // CPU-half (token-generation path). Off by default; flag-off is byte-identical.
+    // x86_64-only at runtime (the engine compiles only on the AVX2 x86_64 slice).
+    bool    iqk_flash_attn        = false;
+    float   pcie_bw_gbps          =  0.0f;
+    // opencoti F4 M3 Phase 2 β€” see docs/decisions/0001-lazy-slot-context.md
+    // Initial KV-cell allocation per sequence. 0 = full n_ctx_seq allocation
+    // (Phase 1 behavior β€” only zero-fill is deferred). Positive values cap
+    // the initial allocation and enable grow-on-demand up to n_ctx_seq.
+    int32_t slot_initial_ctx      =     0;
+    // opencoti F4 M3 Phase 3 β€” see docs/decisions/0001-lazy-slot-context.md
+    // Idle threshold (ms) after which a grown soft cap is shrunk back to
+    // slot_initial_ctx and the over-cap pages are decommitted via
+    // posix_madvise. 0 disables shrink-on-idle (Phase 2 behavior).
+    int32_t slot_shrink_idle_ms   =     0;
+    // opencoti F5 M2 headinfer β€” see docs/features/advanced_kv.md
+    // Fraction of KV heads kept GPU-resident; the rest are offloaded to host
+    // memory and streamed back per step. 1.0 = off (no split). Only active for
+    // GPU-offloaded layers with flash-attention (non-transposed V cache).
+    float   headinfer_gpu_heads_frac = 1.0f;
+    // opencoti F5 M3 neo β€” see docs/features/advanced_kv.md
+    // 0 = off (default; M2 concat path runs), 1 = on (engage when split
+    // active for the layer), 2 = auto (reserved for future load-based
+    // engagement; same as `on` today).
+    uint32_t neo_pipeline_mode    =     0;
     // opencoti F5-opt W1 (#293) pcie profile β€” see docs/features/rolling_kv.md
     // Boot-time host<->device bandwidth + PCIe link detection, consumed by M7
     // Rolling KV (#296) tile sizing. autodetect reads the rebar-probe JSON /
diff --git a/llama.cpp/ggml/include/ggml.h b/llama.cpp/ggml/include/ggml.h
--- a/llama.cpp/ggml/include/ggml.h
+++ b/llama.cpp/ggml/include/ggml.h
@@ -583,6 +583,8 @@ extern "C" {
 
         GGML_OP_GLU,
 
+        GGML_OP_MOE_FUSED_UP_GATE, // opencoti-hook: f5-opt-fmoe β€” W3 #292
+
         GGML_OP_COUNT,
     };
 
@@ -1437,6 +1439,22 @@ extern "C" {
             struct ggml_tensor  * b,
             struct ggml_tensor  * ids);
 
+    // opencoti-hook: f5-opt-fmoe β€” W3 #292: op-level fused MoE up+gate+GLU.
+    // build_moe_ffn otherwise emits {mul_mat_id(up), mul_mat_id(gate), glu}
+    // as 3 nodes that the CUDA backend can only fuse via the fragile
+    // ggml_can_fuse_subgraph graph-pattern path (which does not engage at
+    // decode for Gemma-4). This single op lets the backend run the already
+    // existing fused mmvq kernel directly. Output is GLU(gate@cur) (.) (up@cur),
+    // matching ggml_geglu_split(gate, up). up_exps/gate_exps: expert weight
+    // stacks (same type); cur: input activations; ids: expert selection.
+    GGML_API struct ggml_tensor * ggml_moe_up_gate(
+            struct ggml_context * ctx,
+            struct ggml_tensor  * up_exps,
+            struct ggml_tensor  * gate_exps,
+            struct ggml_tensor  * cur,
+            struct ggml_tensor  * ids,
+            enum   ggml_glu_op    glu_op);
+
     // A: m columns, n rows,
     // B: p columns, n rows,
     // result is m columns, p rows
diff --git a/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c b/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c
--- a/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c
+++ b/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c
@@ -2066,6 +2066,13 @@ static void ggml_compute_forward(struct ggml_compute_params * params, struct ggm
             {
                 ggml_compute_forward_glu(params, tensor);
             } break;
+        case GGML_OP_MOE_FUSED_UP_GATE:
+            {
+                // opencoti-hook: f5-opt-fmoe β€” W3 #292. CUDA-only op. The
+                // build_moe_ffn graph hook emits it only when the CUDA backend
+                // owns the FFN nodes, so the CPU scheduler never reaches here.
+                GGML_ABORT("MOE_FUSED_UP_GATE has no CPU implementation (CUDA-only op)");
+            } break;
         case GGML_OP_GET_REL_POS:
             {
                 ggml_compute_forward_get_rel_pos(params, tensor);
@@ -2333,6 +2340,12 @@ static int ggml_get_n_tasks(struct ggml_tensor * node, int n_threads) {
                     GGML_ABORT("fatal error");
             }
             break;
+        case GGML_OP_MOE_FUSED_UP_GATE:
+            {
+                // opencoti-hook: f5-opt-fmoe β€” W3 #292. CUDA-only; never
+                // scheduled on CPU. Single task keeps the planner well-formed.
+                n_tasks = 1;
+            } break;
         case GGML_OP_SILU_BACK:
         case GGML_OP_MUL:
         case GGML_OP_DIV:
diff --git a/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu b/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu
--- a/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu
+++ b/llama.cpp/ggml/src/ggml-cuda/ggml-cuda.cu
@@ -3005,6 +3005,25 @@ static bool ggml_cuda_compute_forward(ggml_backend_cuda_context & ctx, struct gg
         case GGML_OP_MUL_MAT_ID:
             ggml_cuda_mul_mat_id(ctx, dst);
             break;
+        case GGML_OP_MOE_FUSED_UP_GATE:
+            {
+                // opencoti-hook: f5-opt-fmoe β€” W3 #292. Drive the existing fused
+                // mmvq-id kernel directly: up = src0 @ src1, gate = fusion.gate @
+                // src1, out = act(gate) (.) up. Byte-identical to the graph-pattern
+                // fusion path (see the { MUL_MAT_ID, MUL_MAT_ID, GLU } branch in
+                // the eval fusion loop). Decode-only (ne[2]==1): the build_moe_ffn
+                // hook only emits this at n_tokens==1 and supports_op enforces it.
+                const ggml_tensor * up_exps   = dst->src[0];
+                const ggml_tensor * gate_exps = dst->src[1];
+                const ggml_tensor * cur       = dst->src[2];
+                const ggml_tensor * ids       = dst->src[3];
+                GGML_ASSERT(dst->ne[2] == 1 && "MOE_FUSED_UP_GATE is decode-only (n_tokens==1)");
+                ggml_cuda_mm_fusion_args_host fusion_data{};
+                fusion_data.gate   = gate_exps;
+                fusion_data.glu_op = (ggml_glu_op) dst->op_params[0];
+                ggml_cuda_mul_mat_vec_q(ctx, up_exps, cur, ids, dst, &fusion_data);
+            }
+            break;
         case GGML_OP_OUT_PROD:
             ggml_cuda_out_prod(ctx, dst);
             break;
@@ -5210,6 +5229,44 @@ static bool GGML_CALL ggml_backend_cuda_device_supports_op(ggml_backend_dev_t de
                     return false;
             }
             break;
+        case GGML_OP_MOE_FUSED_UP_GATE:
+            {
+                // opencoti-hook: f5-opt-fmoe β€” W3 #292. Supported only in the
+                // exact shape the fused mmvq-id decode kernel handles: quantized
+                // up/gate expert stacks of identical type, F32 activations, I32
+                // ids, single-token (decode) output, non-split weights, and a GLU
+                // op the kernel implements explicitly (GEGLU/SWIGLU/SWIGLU_OAI β€”
+                // others silently degrade to a plain product). build_moe_ffn only
+                // emits the op when all of these hold, so supports_op and the hook
+                // stay consistent (the scheduler never strands it on the CPU).
+                const ggml_tensor * up_exps   = op->src[0];
+                const ggml_tensor * gate_exps = op->src[1];
+                const ggml_tensor * cur       = op->src[2];
+                const ggml_tensor * ids       = op->src[3];
+                if (!up_exps || !gate_exps || !cur || !ids) {
+                    return false;
+                }
+                if (up_exps->buffer && ggml_backend_buft_is_cuda_split(up_exps->buffer->buft)) {
+                    return false;
+                }
+                if (!ggml_is_quantized(up_exps->type) || up_exps->type != gate_exps->type) {
+                    return false;
+                }
+                if (cur->type != GGML_TYPE_F32 || ids->type != GGML_TYPE_I32) {
+                    return false;
+                }
+                if (op->ne[2] != 1) {
+                    return false;
+                }
+                switch ((ggml_glu_op) op->op_params[0]) {
+                    case GGML_GLU_OP_GEGLU:
+                    case GGML_GLU_OP_SWIGLU:
+                    case GGML_GLU_OP_SWIGLU_OAI:
+                        return true;
+                    default:
+                        return false;
+                }
+            }
         case GGML_OP_MUL_MAT:
         case GGML_OP_MUL_MAT_ID:
             {
diff --git a/llama.cpp/ggml/src/ggml.c b/llama.cpp/ggml/src/ggml.c
--- a/llama.cpp/ggml/src/ggml.c
+++ b/llama.cpp/ggml/src/ggml.c
@@ -1078,9 +1078,10 @@ static const char * GGML_OP_NAME[GGML_OP_COUNT] = {
     "OPT_STEP_SGD",
 
     "GLU",
+    "MOE_FUSED_UP_GATE",
 };
 
-static_assert(GGML_OP_COUNT == 96, "GGML_OP_COUNT != 96");
+static_assert(GGML_OP_COUNT == 97, "GGML_OP_COUNT != 97");
 
 static const char * GGML_OP_SYMBOL[GGML_OP_COUNT] = {
     "none",
@@ -1188,9 +1189,10 @@ static const char * GGML_OP_SYMBOL[GGML_OP_COUNT] = {
     "sgd(x)",
 
     "glu(x)",
+    "moe_up_gate(up,gate,x)",
 };
 
-static_assert(GGML_OP_COUNT == 96, "GGML_OP_COUNT != 96");
+static_assert(GGML_OP_COUNT == 97, "GGML_OP_COUNT != 97");
 
 static_assert(GGML_OP_POOL_COUNT == 2, "GGML_OP_POOL_COUNT != 2");
 
@@ -3314,6 +3316,53 @@ struct ggml_tensor * ggml_mul_mat_id(
     return result;
 }
 
+// ggml_moe_up_gate
+// opencoti-hook: f5-opt-fmoe β€” W3 #292. Op-level fusion of the two parallel
+// expert matmuls (up = up_exps @ cur, gate = gate_exps @ cur) plus the GLU
+// activation that build_moe_ffn otherwise emits as 3 separate graph nodes.
+// Output == ggml_geglu_split(gate, up) == act(gate) (.) up, where the CUDA
+// backend runs the existing fused mmvq kernel (act on the gate projection,
+// multiply by the up projection β€” see ggml-cuda/mmvq.cu). Mirrors
+// ggml_mul_mat_id's shape contract: up_exps is the primary "as" weight, cur
+// the "b" activations shared by both projections, ids the expert selection.
+
+struct ggml_tensor * ggml_moe_up_gate(
+        struct ggml_context * ctx,
+        struct ggml_tensor  * up_exps,
+        struct ggml_tensor  * gate_exps,
+        struct ggml_tensor  * cur,
+        struct ggml_tensor  * ids,
+        enum   ggml_glu_op    glu_op) {
+    GGML_ASSERT(!ggml_is_transposed(up_exps));
+    GGML_ASSERT(!ggml_is_transposed(gate_exps));
+    GGML_ASSERT(ids->type == GGML_TYPE_I32);
+
+    // up and gate are parallel expert stacks: identical shape + type so one
+    // shared-activation fused mmvq can drive both weight sets.
+    GGML_ASSERT(ggml_are_same_shape(up_exps, gate_exps));
+    GGML_ASSERT(up_exps->type == gate_exps->type);
+
+    GGML_ASSERT(up_exps->ne[3] == 1);                 // 3d (one matrix per expert)
+    GGML_ASSERT(cur->ne[3] == 1);                     // cur is 3d
+    GGML_ASSERT(ids->ne[2] == 1 && ids->ne[3] == 1);  // ids is 2d
+    GGML_ASSERT(ids->ne[1] == cur->ne[2]);            // one expert list per cur row
+    GGML_ASSERT(up_exps->ne[0] == cur->ne[0]);        // can_mul_mat
+    GGML_ASSERT(ids->ne[0] % cur->ne[1] == 0);        // can broadcast
+
+    const int64_t ne[4] = { up_exps->ne[1], ids->ne[0], cur->ne[2], 1 };
+    struct ggml_tensor * result = ggml_new_tensor(ctx, GGML_TYPE_F32, 4, ne);
+
+    ggml_set_op_params_i32(result, 0, (int32_t) glu_op);
+
+    result->op     = GGML_OP_MOE_FUSED_UP_GATE;
+    result->src[0] = up_exps;
+    result->src[1] = gate_exps;
+    result->src[2] = cur;
+    result->src[3] = ids;
+
+    return result;
+}
+
 // ggml_out_prod
 
 static inline bool ggml_can_out_prod(const struct ggml_tensor * t0, const struct ggml_tensor * t1) {
diff --git a/llama.cpp/include/llama.h b/llama.cpp/include/llama.h
--- a/llama.cpp/include/llama.h
+++ b/llama.cpp/include/llama.h
@@ -362,6 +362,9 @@ extern "C" {
         // 0 = off (default), 1 = on, 2 = auto. Effective only when the
         // M2 head split is active for a given layer.
         uint32_t neo_pipeline_mode;
+        // opencoti F5 W3 #292 fused-MoE β€” see docs/features/advanced_kv.md
+        // true = fuse decode MoE up+gate+GLU into one op. false = off (default).
+        bool     fused_moe_up_gate;
         uint32_t n_rs_seq;          // number of recurrent-state snapshots per seq for rollback (0 = no rollback) [EXPERIMENTAL]
         int32_t  n_threads;         // number of threads to use for generation
         int32_t  n_threads_batch;   // number of threads to use for batch processing
diff --git a/llama.cpp/src/llama-context.cpp b/llama.cpp/src/llama-context.cpp
--- a/llama.cpp/src/llama-context.cpp
+++ b/llama.cpp/src/llama-context.cpp
@@ -66,6 +66,8 @@ llama_context::llama_context(
     cparams.headinfer_gpu_heads_frac = params.headinfer_gpu_heads_frac;
     // opencoti F5 M3 neo β€” see docs/features/advanced_kv.md
     cparams.neo_pipeline_mode = params.neo_pipeline_mode;
+    // opencoti F5 W3 #292 fused-MoE β€” see docs/features/advanced_kv.md
+    cparams.fused_moe_up_gate = params.fused_moe_up_gate;
 
     cparams.n_threads        = params.n_threads;
     cparams.n_threads_batch  = params.n_threads_batch;
@@ -3359,6 +3361,7 @@ llama_context_params llama_context_default_params() {
         /*.slot_shrink_idle_ms         =*/ 0,
         /*.headinfer_gpu_heads_frac    =*/ 1.0f,
         /*.neo_pipeline_mode            =*/ 0,  // opencoti F5 M3 neo β€” off by default
+        /*.fused_moe_up_gate           =*/ false, // opencoti F5 W3 #292 β€” off by default
         /*.n_rs_seq                    =*/ 0,
         /*.n_threads                   =*/ GGML_DEFAULT_N_THREADS, // TODO: better default
         /*.n_threads_batch             =*/ GGML_DEFAULT_N_THREADS,
diff --git a/llama.cpp/src/llama-cparams.h b/llama.cpp/src/llama-cparams.h
--- a/llama.cpp/src/llama-cparams.h
+++ b/llama.cpp/src/llama-cparams.h
@@ -32,6 +32,10 @@ struct llama_cparams {
     // 2 = auto (same engagement as `on` for now β€” reserved for future
     //         load-based engagement heuristics).
     uint32_t neo_pipeline_mode;
+    // opencoti F5 W3 #292 fused-MoE β€” see docs/features/advanced_kv.md.
+    // true = build_moe_ffn collapses the up/gate expert matmuls + GLU into one
+    // GGML_OP_MOE_FUSED_UP_GATE op at decode (n_tokens==1). false = unfused path.
+    bool     fused_moe_up_gate;
     uint32_t n_rs_seq;        // number of recurrent-state snapshots per seq for rollback
     int32_t  n_threads;       // number of threads to use for generation
     int32_t  n_threads_batch; // number of threads to use for batch processing
diff --git a/llama.cpp/src/llama-graph.cpp b/llama.cpp/src/llama-graph.cpp
--- a/llama.cpp/src/llama-graph.cpp
+++ b/llama.cpp/src/llama-graph.cpp
@@ -1534,7 +1534,28 @@ ggml_tensor * llm_graph_context::build_moe_ffn(
     ggml_tensor * up = nullptr;
     ggml_tensor * experts = nullptr;
 
-    if (gate_up_exps) {
+    // opencoti-hook: f5-opt-fmoe β€” W3 #292: collapse the separate-up/gate
+    // expert matmuls + the GLU activation into ONE GGML_OP_MOE_FUSED_UP_GATE
+    // op at decode, so the CUDA backend runs the existing fused mmvq kernel
+    // without relying on graph-pattern fusion (ggml_can_fuse_subgraph, which
+    // does not engage for Gemma-4 at decode). Guards mirror supports_op:
+    // decode only (n_tokens==1), separate up/gate of identical type, no
+    // bias/scale, and GELU→GEGLU or SILU→SWIGLU. Off path is byte-identical.
+    bool fused_up_gate = false;
+    if (cparams.fused_moe_up_gate && gate_exps && !gate_up_exps &&
+            !up_exps_b && !gate_exps_b && !up_exps_s && !gate_exps_s &&
+            up_exps->type == gate_exps->type && n_tokens == 1 &&
+            (type_op == LLM_FFN_GELU || type_op == LLM_FFN_SILU)) {
+        const enum ggml_glu_op glu_op = (type_op == LLM_FFN_GELU)
+            ? GGML_GLU_OP_GEGLU : GGML_GLU_OP_SWIGLU;
+        cur = ggml_moe_up_gate(ctx0, up_exps, gate_exps, cur, selected_experts, glu_op);
+        cb(cur, "ffn_moe_fused_up_gate", il);
+        fused_up_gate = true;
+    }
+
+    if (fused_up_gate) {
+        // cur already holds act(gate) (.) up from the fused op
+    } else if (gate_up_exps) {
         // merged gate_up path: one mul_mat_id, then split into gate and up views
         ggml_tensor * gate_up = build_lora_mm_id(gate_up_exps, cur, selected_experts); // [n_ff*2, n_expert_used, n_tokens]
         cb(gate_up, "ffn_moe_gate_up", il);
@@ -1601,6 +1622,9 @@ ggml_tensor * llm_graph_context::build_moe_ffn(
 
     const bool has_gate = gate_exps || gate_up_exps;
 
+    // opencoti-hook: f5-opt-fmoe β€” W3 #292: the fused op already applied the
+    // GLU activation; skip the separate activation pass.
+    if (!fused_up_gate)
     switch (type_op) {
         case LLM_FFN_SILU:
             if (gate_exps) {