Decode-throughput lever map β turbo / turbo_tcq / scalar (A4B, RTX 3090)
Source: nsys per-kernel decode profiles (.opencoti/nsys/lever_{scalar,turbo3,tcq}.nsys-rep),
A4B-98e-v6-coder Q4_K_M, ctx 4096, n_predict 96, DSO 60b17cb1 (post-#535 vectorized K-read).
Cross-config absolute totals are NOT comparable (different #decode-steps captured); the comparable
views are per-phase % of that config's own total and per-call avg Β΅s.
SCOPE (WS2 / #629). Every "turbo/TCQ costs decode vs scalar" number in this document is noDCA, short-ctx (ctx 4096), f16-lift β the regime where the centroid-unpack / codebook lookup is pure overhead. With DCA-on at long ctx the picture inverts: the WS2 in-register turbo-V kernel (DEFAULT-ON) makes turbo/TCQ-V a decode WIN of +68β95% @256k over the f16-lift ref, because it kills the context-growing whole-cache
cpy_turboN_f16materialize that dominates the long-ctx step. The gaps below are the short-ctx floor, not the serving verdict β see context.md Β§5x.
Measured baseline (decode tps, A4B)
scalar q8_0 68.7 Β· turbo3 58.2 Β· turbo3_tcq 55.6. Gap turbo3βscalar = β15% β noDCA / short-ctx only; DCA-on long-ctx inverts (Β§5x, WS2 in-register turbo-V is +68β95%).
Per-call FA-VEC K-read (config-independent)
| D | scalar q8_0 | turbo3 | turbo3_tcq |
|---|---|---|---|
| 256 (local layers) | 21.1 Β΅s | 63.0 Β΅s (3.0Γ) | 49.1 Β΅s (2.3Γ) |
| 512 (global layers) | β | 270.4 Β΅s | 210.6 Β΅s |
mul_mat_vec_q experts = 17.8 Β΅s/call for ALL THREE (KV-type-independent β confirms FFN is shared).
Phase fractions (% of each config's decode kernel time)
| Phase | scalar | turbo3 | tcq |
|---|---|---|---|
| FFN/MoE experts (mul_mat_vec_q) | 54.3 | 44.4 | 41.6 |
| NORM/softmax/act/elementwise | 15.9 | 13.0 | 12.2 |
| Turbo/TCQ K-read (FA-VEC) | (4.5 scalar) | 20.5 | 15.0 |
| TCQ KV-write (Viterbi enc/set_rows) | 2.5 | 2.3 | 13.6 |
| KVwrite cpy/quantize | 4.6 | 3.7 | 3.5 |
| WHT rotate (turbo) | 2.0 | 2.9 | 2.7 |
| embed/get_rows/argsort | 3.7 | 3.0 | 2.8 |
| ROPE | 1.1 | 0.9 | 0.9 |
| stream-k fixup/combine | 0.9 | 0.6 | 0.5 |
| FFN other matmul | 3.2 | 3.9 | 3.0 |
| OTHER (launch slack etc.) | 7.2 | 4.8 | 4.2 |
LEVERS by phase β ranked by (impact Γ tractability)
L1 β SPECULATIVE DECODE (cross-cutting MULTIPLIER) β biggest
MTP draft head already ships 1.46Γ on f16. If turbo/tcq COMPOSE with MTP, that ~1.4Γ multiplies the WHOLE decode β bigger than any single-kernel tweak. STATUS: unverified for turbo/tcq KV. Effort: gate-only (no kernel) if it already composes; medium if KV-write-on-accept needs work. Verify first.
L2 β FFN/MoE experts (mul_mat_vec_q) β 42β54%, SHARED, biggest absolute
17.8 Β΅s/call, dominates ALL configs. Improving it raises turbo AND scalar (doesn't close the gap, but the user's goal is "improve turbo" in absolute terms). Levers: (a) fused up+gate op (GGML_OP_MOE_FUSED_UP_GATE exists β verify it's ON the A4B decode path, not just prefill); (b) mmvq nrows batching across the 2 active experts; (c) ik_llama fused-MoE (#292 OOAE) decode path. Effort: medium-high. Risk: numerics (cosine gate).
(a) MEASURED 2026-06-23 β INAPPLICABLE TO A4B (op never engages); do NOT default-flip. --fused-moe-up-gate
is a structural no-op on Gemma A4B: the boot -v log shows its experts load as a single combined
blk.N.ffn_gate_up_exps.weight tensor (up and gate projections already share one matmul). The graph hook
fires the fused op only when fused_moe_up_gate && gate_exps && !gate_up_exps && n_tokens==1 && (GELU|SILU);
A4B has gate_up_exps, so !gate_up_exps is false β the op is never inserted, flipping the flag leaves the
graph identical (bit-identical output). 2Γ2 proof {graphs on/off}Γ{fmug on/off}, A4B Q4_K_M decode tok/s:
| fmug=OFF | fmug=ON | fused-gain | |
|---|---|---|---|
| graphs ON | 92.70 | 92.45 | β0.3% |
| graphs OFF | 78.25 | 78.33 | +0.1% |
Fused-gain β 0 in both rows β this is NOT subsumption-by-graphs (that would leave the graphs-OFF gain
visible); it's the op never firing because the experts are pre-fused. (Graphs reconfirmed orthogonally:
+18.5% at fmug-off.) So L2(a) is only meaningful for GGUFs with separate ffn_gate_exps+ffn_up_exps
experts β check the target model's tensor names first. For pre-fused-expert models like A4B, the up/gate fusion
is already baked into the weights. What remains in L2 is the matmul itself (b/c) β a genuine compute-bound
win, but medium-high effort with numeric risk; deferred unless the FFN matmul becomes the proven bottleneck
post-L1/L6. NOTE: on a model with separate experts AND graphs on, expect fused-MoE to be subsumed (graphs
already kill its launch saving) β the residual would only be the saved intermediate-buffer traffic.
(a) CONTRAST β Qwen DOES have separate experts, op fires. Qwen3.6-35B-A3B (qwen35moe, 256 experts) loads
separate ffn_gate_exps+ffn_up_exps (vs A4B's combined ffn_gate_up_exps), so the fused op engages.
2Γ2 {graphs}Γ{fmug}, f16 decode tok/s:
| fmug=OFF | fmug=ON | fused-gain | |
|---|---|---|---|
| graphs ON | 89.17 | 91.31 | +2.4% |
| graphs OFF | 76.97 | 76.26 | β0.9% (noise) |
So fused-MoE is model-dependent: inert on A4B (pre-fused weights), a small real win on Qwen-style separate-
expert MoEs β and, interestingly, the gain shows up with graphs ON (+2.4%), not off (the fewer graph nodes
help replay; off-graph it's neutral/noise). CLOSED 2026-06-24 (greedy-identity gate): on Qwen3.6-35B-A3B
(qwen35moe, 256 experts, separate ffn_gate/up_exps) fused-MoE off-vs-on is greedy BIT-IDENTICAL at
temp=0 for BOTH f16 and turbo3_tcq KV (764B==764B, 661B==661B) β numerically safe. But under the SHIPPING
config (graphs ON) the re-measured gain is only +0.7β0.9%, not +2.4%: the +2.4% was a graphs-OFF baseline
artifact (graphs already amortize the launch overhead fused-MoE targets, leaving a sub-1% residual once graphs
are on). Verdict: safe but sub-1% under the shipping config β NOT worth a default-flip + the validated-model-
surface expansion; the flag stays opt-in (default OFF). fused-MoE is a minor, situational lever, NOT a
headline like L6/L1.
(b) "mmvq nrows-batch across active experts" β ALREADY IMPLEMENTED (nothing to build), 2026-06-23.
MMVQ_MAX_BATCH_SIZE = 8 (mmvq.cuh:3) and both targets use n_expert_used = 8, so ne2=8 β€ 8 takes the
batched MUL_MAT_ID path at ggml-cuda.cu:2677-2681 β ggml_cuda_mul_mat_vec_q(src0, src1, ids, dst): one
launch with all 8 active experts mapped onto the grid's nchannels_dst axis (mmvq.cu:673, grid
(nblocks, nchannels_dst, ntokens)). There is no per-expert-launch loop to collapse. (Only >8-expert-used
models fall to the slower MMQ path β not our targets.) Reduces no bytes-read; moot.
(c) "ik_llama OOAE + fused-MoE" β fused-MoE shipped (0042), OOAE irrelevant to the matmul, 2026-06-23.
The fused up+gate op is already the 0042 path (ggml-cuda.cu:3030). ik_llama's other piece is topk-moe
(fused softmax+top-k expert selection); our tree uses standard argsort_top_k. topk-moe speeds selection,
not the expert matmul, and reads the same expert weights β orthogonal to the 42β54% FFN cost. (#292 was logged
"OOAE + fused-MoE pair" but only the fused-MoE half landed; the topk-moe half is absent and would not move
decode tps.)
ROOT REASON b AND c ARE DEAD ENDS (verified): the expert matmul is a GEMV over quantized weights β
arithmetic intensity β3.5 FLOP/byte βͺ the RTX 3090 roofline ridge (~37) β hard bandwidth-bound on
reading the expert weights. Neither b nor c reduces those bytes. The only levers that attack a bandwidth-
bound weight read: L1/MTP (amortizes each weight read across K draft tokens per verify step β the proven
+55%) or lower-bit experts (quant choice). A faster GEMV kernel would help only if mmvq leaves bandwidth
on the table β a separate #397-class effort to be justified by an ncu DRAM-throughput measurement first, NOT
b/c. Conclusion: L2(b)/(c) require no work; closed.
#397 GEMV-kernel verdict β MEASURED (ncu, 2026-06-23, A4B Q4_K_M decode): the time-dominant expert
mul_mat_vec_q<12,β¦> launches run at 77β81 % of the 3090's peak DRAM bandwidth (time-weighted 61.6 %
across all decode mmvq launches; SM 41 %). The heavy kernels are already near the memory-bound practical
ceiling (85β90 %), so a faster GEMV buys at most ~1.1β1.25Γ on the dominant slice and less overall.
#397 (int8-dp4a turbo KQ / faster-GEMV) is NOT worth pursuing for decode tps β the real levers remain
L1/MTP and lower-bit experts, which cut bytes read, not read speed. Method (reusable): the cosmopolitan
APE binary can't be execve'd by ncu/nsys directly ("Unknown executable format") β copy it and
./copy --assimilate to a native ELF (cosmo runtime intercepts the flag pre-main), then
ncu --kernel-name regex:mul_mat_vec --launch-skip N --launch-count M --metrics gpu__dram_throughput.avg.pct_of_peak_sustained_elapsed,β¦ (set GGML_CUDA_DISABLE_GRAPHS=1 so each launch is
individually filterable; mul_mat_vec_q is decode-only β prefill uses the batched mul_mat path).
L3 β TCQ decode-time Viterbi encode/set_rows β 13.6% of tcq, TCQ-SPECIFIC β β (a) DONE (#541); (b) REJECTED (#545); (c) UNATTEMPTED
The trellis encode runs on EVERY decode KV-write. 13.6% of tcq decode = the single biggest tcq-vs-turbo3
penalty. Levers: (a) codebook-hoist in the ENCODE kernel β DONE (#541, DSO b4ccda34): hoisted
d_turbo3_tcq_codebook[16*lane+i] out of the 128-step Viterbi DP into a per-lane cb[16] register array
(turbo2: cb[8]). Byte-identical CSE (encode REG:60/STACK:0, occupancy 3/SM held). Result turbo3_tcq decode
59.7 β 60.4 (+1.2%); scalar 87.6 / turbo3 76.2 unchanged. 27B KLD healthy on both edited kernels
(turbo3_tcq 0.034606 sametop 95.5%, turbo2_tcq 0.069940 PASS sametop 92.6%). VERDICT: the codebook reads were
NOT the encode bottleneck β the 128-step serial DP is β so the hoist only buys ~1%. The big tcq win is (b).
(b) cheaper 1-token decode-time encode (the Viterbi DP is NOT amortizable to O(8): the 128-step trellis spans
the 128 HEAD-DIM elements of the single new token's vector, all new each step; the candidate lever was a
greedy walk over the 128 element-steps, O(128Β·8) vs O(128Β·512)). REJECTED 2026-06-24 (#545, P1 spike,
flag-gated greedy encode on DSO 9403e25f, then reverted to L3 byte-exact). Same-DSO 27B fidelity-KLD:
greedy-tcq 1.0575 (sametop 61.6%) vs Viterbi-tcq 0.0346 (95.5%) vs plain turbo3 0.0470 (94.7%). Greedy
does not merely regress toward turbo3 β it collapses ~30Γ past Viterbi and ~23Γ past plain turbo3: the
512-entry codebook is jointly trained for the Viterbi-optimal paths, so a greedy walk lands on mismatched
centroids and each locally-best symbol paints the next step into a corner. The trellis's MSE-optimal
noise-shaping is irreducible β the ~8% decode-time encode cost is the price of tcq's quality edge, not a
removable inefficiency. Since plain turbo3 is both faster AND (at 0.047) far better than greedy-tcq (1.057),
the answer where greedy was the goal is simply "ship turbo3". (c) skip re-encode via an f16 ring + lazy
batched TCQ-encode β UNATTEMPTED; the only remaining idea, but it changes write semantics (a token's KV is
written f16 then re-encoded in a later batch), so it is a structural rework gated on the same KLD bar, not a
quick win. Effort: high. Risk: quality (KLD gate) + correctness (deferred-encode state IO).
L3 round verdict (closes the decode-tps investigation): the only banked win is L3(a) +1.2% on tcq. The
floor is real and now evidence-backed on both axes β turbo3 vs scalar q8_0 β β10% (centroid unpack is a
ShiftβMaskβLUTβScale dependency chain vs int8*scale), turbo3_tcq vs turbo3 β β8% to β13% (512-entry LUT +
9-bit unaligned extract). No free lever remains; further tcq speedup requires either format change (smaller
LUT β retrain, P2, quality-gated) or the (c) deferred-encode rework. β noDCA / short-ctx only; with DCA-on
at long ctx the WS2 in-register turbo-V path is a decode WIN (+68β95%), see context.md Β§5x.
L4 β NORM/softmax/elementwise β 12β16%, SHARED, many small kernels
rms_norm + soft_max + scale/add/residual = lots of tiny launch-bound kernels. Lever: fuse (norm+scale, residual-add into the matmul epilogue, RoPE+norm). Also a CUDA-graph candidate. Effort: medium. Risk: low (byte-identical-ish).
L5 β Turbo K-read FA-VEC β 20.5% turbo3 / 15% tcq, TURBO-SPECIFIC β β DONE (#535b, 2026-06-23)
Vectorized loads (#535) took D256 70β63 Β΅s (+1.6%). Then #535b smem-codebook port (DSO 47b0bca5): the
plain turbo K-read/V-dequant indexed FATTN_TURBO{2,3,4}_CENTROIDS in __constant__, which serializes the
data-dependent per-element lookups up to 32-way/warp; the TCQ tiers already staged to per-block smem (#534).
Porting that staging to turbo2/3/4 is BYTE-IDENTICAL and gave:
- D256 turbo3 K-read 63.0 β 48.4 Β΅s (β23%) β now beats turbo3_tcq's 49.1 (8-entry LUT < 512-entry codebook).
- D512 turbo3 K-read 270.4 β 135.2 Β΅s (β50%).
- decode tps: turbo3 58.2 β 62.08 (+6.7%), turbo4 β 63.5, turbo2 β 66.5; scalar 68.96 unchanged.
- KLD turbo3βf16 = 0.349646 == the #535 value (byte-identical); turbo4 0.109 sane.
- turbo2 quality stays bug-583 K-side degenerate (in-register 2-bit; pre-existing, reproduced byte-identically: turbo2K/f16V=7.39, f16K/turbo2V=0.043) β use turbo2 via DCA dequant-on-lift, never this in-register reader. LESSON β cerebrum 2026-06-23: small data-dependent LUTs in hot decode loops belong in smem, not constant.
L6 β CUDA graphs β β DONE (#533/#539, 2026-06-23, ships as 0084-cuda-graphs)
Decode launches ~hundreds of tiny kernels/token; at 58 tps each token is ~17 ms but GPU SM throughput is
~4.5% (latency/launch-bound). Root cause found: CUDA graph capture/replay was COMPILED OUT of the
llamafile build β USE_CUDA_GRAPH is gated #if defined(GGML_CUDA_USE_GRAPHS) (common.cuh:1218) and the
llamafile cuda.sh never defined it (upstream CMake does, by default). So #ifdef USE_CUDA_GRAPH
(ggml-cuda.cu:3318/4478/4515/4593/4698) was absent and every token paid full per-launch CPU overhead.
Fix = one build flag -DGGML_CUDA_USE_GRAPHS (0084; OPENCOTI_NO_CUDA_GRAPHS=1 to omit). MoE does NOT
disable it (the MUL_MAT_ID quantized-expert path stays graph-compatible; NEO head-split force-disables to
eager). Measured (A4B Q4_K_M, airtight same-DSO GGML_CUDA_DISABLE_GRAPHS toggle):
- f16 decode 77.4 β 92.4 (+19.4%), turbo3_tcq 53.6 β 60.6 (+13.1%) β shared, lifts every KV tier.
- Stacks with MTP (L1): turbo3_tcq graphs+MTP β 89 tok/s (+69% over the 52.9 base; exceeds plain f16 77.7, ~75% of f16-with-same-levers 119.7).
- Numerically inert β greedy bit-identical graphs-on vs -off across f16/turbo3/turbo3_tcq/q8_0/q4_0 + 8K ctx (graphs replay the identical kernel sequence; not a compute change).
- GENERALIZES to Qwen (model-agnostic). Qwen3.6-35B-A3B (qwen35moe, 256 experts) decode graphs on vs off:
f16 76.97β89.17 (+15.9%), turbo3_tcq 68.52β78.49 (+14.6%) β same magnitude as Gemma. The lever is
not arch-specific; it lifts every model's latency-bound decode. (turbo3_tcq also confirmed working on Qwen
head_dim 128.)
LESSON β cerebrum 2026-06-23: opencoti inherited llamafile's
cuda.sh, which omits upstream-default build defines β auditcuda.shvs upstream CMakeGGML_CUDA_*options for other free wins (graphs were one).
L7 β WHT rotate (turbo) β 2β3%, turbo-specific, small
k_turbo_wht runs per KV-write. Minor. Fuse into set_rows. Low priority.
Recommended sequence (phase-by-phase)
- L1 verify turbo/tcq Γ MTP composition (cheap, possibly huge).
- L3 TCQ decode-encode smem/cheap-encode (closes tcq's 13.6% β the clearest tcq win).
- L5b V-read vectorization (cheap byte-identical, finishes the turbo K/V read).
- L2 FFN/MoE fused-up-gate verify + decode-path (biggest absolute, shared).
- L4 / L6 norm fusion + CUDA graphs (shared launch-bound reduction).