File size: 9,010 Bytes
5ef4cc5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | # Adaptive compute β depth-skip & dynamic-precision feasibility
> Status: **TRAINING-FREE VARIANTS CLOSED** (2026-06-30). Three cheap PyTorch forward-hook probes
> (S1/S1b/S2) on Gemma-4 A4B establish that *static, predictor-free* compute-skipping and
> dynamic weight-precision have **no headroom** on this model as deployed. One variant is not yet
> ruled out β a **trained** Mixture-of-Depths router (Β§6) β and is being prototyped. This note is the
> record so the line is not re-litigated from scratch.
This is the *compute-axis* companion to [decode-levers.md](decode-levers.md) (which makes the compute
opencoti *does* run faster) and [poly_kv.md] / [context.md](context.md) (the *KV/memory* axis). Where
those find wins, this documents a **negative result**: you cannot cheaply *remove* compute on A4B.
The motivating idea (user, 2026-06-30): *"I have q8 weights; if I can tell in advance that q4 would
not change an activation meaningfully, I skip the high-precision compute and only activate the neurons
that change something."* Plus the depth-skip cousin: skip whole layers / MLP sub-blocks that contribute
little. The user explicitly distrusted the usual predictor premise (*"'10% of neurons carry 90% of the
mass' β I don't think it's ever true"*) β and the data backs that distrust.
---
## 1. Prior art (field scan)
| Method | Mechanism | Training? | Batch-friendly? | Note |
|---|---|---|---|---|
| DejaVu | contextual sparsity + low-rank predictor | no (predictor) | **no** | union of active neurons grows with batch |
| PowerInfer | hot/cold neuron split | needs ReLU model | **no** | same batch erosion |
| Mixture-of-Depths (MoD) | router picks top-k *tokens* per layer | **yes** | **yes** (capacity-based) | Router-Tuning = cheap post-hoc variant β Β§6 |
| LayerSkip / SkipDecode | early-exit / per-token depth | **yes** (trained) | SkipDecode yes | needs finetune |
| Any-Precision LLM | MSB-first bitplane nested quant | no (training-free) | yes | exact truncation; no GGUF path |
| DeltaLLM / temporal-delta | reuse previous-step state | no | no | KV-side; no reported wall-clock decode win |
| ProSparse / TurboSparse / CATS | ReLUfication | finetune (CATS near-free) | yes | changes activation fn |
Takeaway going in: the **training-free + batch-friendly** quadrant (what opencoti wants for
[multi-session serving](#)) was empty in the literature. S1βS2 tested whether A4B's own structure
opens it anyway. It does not.
---
## 2. Method
PyTorch forward-hooks on the real **Gemma-4 26B-A4B-it** (bf16, `/srv/ml/models/base/β¦` on bs2 GPU0;
30 decoder layers, d_model 2816; 3 realistic coding prompts, greedy decode). Architecture-level
property β bf16 weights are representative. Probes are in `.opencoti/dynq/{s1_probe,s1b_precision,
s2_skip}.py`; raw JSON `{s1,s1b,s2}_out.json`. Correctness gate = mean **KL(bf16 β condition)** on
gen-position logits (greedy token-match is FP-unstable and inadmissible β see bug-270/356).
Notation: `r_layer = βh_outβh_inβ/βh_inβ` (whole-block residual contribution); `r_attn`, `r_mlp` =
sub-block-output / layer-input norm proxies; `cosΞh` = cosine of consecutive decode tokens' residual
input; `Ξh-sparsity` = fraction of dims for 90% of βΞhβΒ².
---
## 3. S1 β is there a lazy band?
| band | r_attn (p50) | r_mlp (p50) | reading |
|---|---|---|---|
| L0 | 1.76 | 11.9 | embedding bootstrap |
| L3β8 | 0.4β1.3 | 0.19β0.76 | both active |
| **L9β19** | 0.30β0.57 | **0.011β0.051** | MLP near-dead, attention alive |
| L23β28 | 0.53β0.96 | 0.82 β **7.14** | late MLP does the heavy lifting |
- `r_layer` p50 never drops below ~0.26 β **no skippable whole-block dead band.**
- The only lazy region is the **MLP (MoE) sub-block of mid layers ~9β19** (1β5% of residual). Compute is
*back-loaded*: late MLPs explode. This *looked* like a batch-friendly static-skip target β Β§5 kills it.
- **Temporal-delta (#4-A) β half-refuted.** `cosΞh` is high (0.75β0.85) at L10β19, but `Ξh-sparsity` is
sparse (2β5% dims) only in *early* layers (L1β9, where cos is low 0.13β0.65) and densifies to 20β40%
exactly where it's stable. The two conditions temporal-delta needs (stability *and* sparse delta)
**anti-correlate**; no general "compute only WΒ·Ξh" win. Batch-fragile besides.
---
## 4. S1b β can the lazy band run at lower weight precision? (the precision-guard, #4-B)
Mean KL(bf16 β cond), MLP weights blockwise-quantized (q4_0/q6_0-style, block 32, symmetric absmax):
| condition | layers | mean KL |
|---|---|---|
| q6_all | 30 | **0.0042** |
| q4_early (0β8) | 9 | 0.0040 |
| q4_late (23β29) | 7 | 0.0135 |
| **q4_mid (9β19)** | 11 | **0.0188** |
| q4_all | 30 | 0.0359 |
Per-layer q4 MLP rel-err *rises* 0.07 (early) β 0.15β0.24 (mid/late).
**Refuted.** q4 on the *lazy* mid-band costs **more** than on the *heavy* late-band β the small MLPs are
precision-*sensitive*. Small contribution β precision-robust; you cannot use "this contributes little"
as a guard for "q4 is safe here." **Premise-breaker:** production opencoti already ships **Q4_K_M** β the
idea assumed q8 weights as the expensive baseline to skip down from; there is no q8 compute to elide.
(q6 is ~9Γ cheaper than q4 in KL, but that points the wrong way β more memory, and prod is already q4.)
---
## 5. S2 β can the lazy band be skipped? (MoD-static, #3)
Mean KL(bf16 β cond), MLP output zeroed via hook (no weight mutation):
- **Single-layer skip:** most mid layers **0.13β0.23** (β10Γ worse than q4-ing *all* 11 at 0.019). Cheap
singles are scattered (L01 0.028, L04 0.04, L13 0.054). Critical layers: L02 **2.0**, L29 **2.1**,
L26/L28 ~1.4 β skipping these alone wrecks the model.
- **Cumulative bands EXPLODE:** mid_9_19 **12.4**, late_23_29 **19.0**, early_0_8 **12.7**, and even the
**10 individually-cheapest layers β 7.27**. Errors **compound super-linearly.**
**Refuted.** The 1β5% mid-MLP contributions are tiny per-step but essential in aggregate; they cannot be
skipped, individually barely and in any band catastrophically. The S1 "batch-friendly mid-MLP static
skip" hope is dead.
---
## 6. Verdict + the one open angle
**The unifying law:** on A4B, a sublayer's small per-step contribution predicts **neither** skippability
(S2) **nor** precision-robustness (S1b). The training-free / predictor-free adaptive-compute family β
DejaVu-style sparsity, depth-skip, dynamic weight-precision, temporal-delta β has **no headroom** here.
This closes the *compute axis* for "more concurrent agentic sessions per card"; the real headroom stays
on the **KV/memory axis** (turbo/TCQ KV-quant, DCA, PolyKV/SharedKVPool prefix-sharing) + MTP for decode
throughput β all shipped or in-flight.
**Caveats (scope of the closure):** greedy-KL on 3 short coding prompts (not RULER/niah) β but cumulative
KL 7β19 is far past any rescue threshold; A4B is MoE (the compounding logic is general); the probes test
**static / training-free** skip+precision only.
## 7. Trained MoD router (the open angle) β router-only FAILS
`.opencoti/dynq/mod_router.py`: per-layer linear gate on band L9β22, **base frozen**, trained by
self-distillation KL(teacher_full β student) with **hard top-k (CAP 0.5) selection in the loop** (kept
tokens scaled by the router sigmoid for gradient, dropped β 0). 48 teacher continuations, 250 steps,
LR 1e-2. Eval = mean KL at hard CAP 0.5.
| config | KL | keep |
|---|---|---|
| pre-train (untrained gate) | 7.00 | 0.50 |
| random 50%-skip control | 4.38 | 0.50 |
| **trained router** | **3.72** | 0.50 |
| (S2 static all-skip mid-band) | 12.4 | β |
**Router-only MoD does not work on A4B.** A trained router beats random selection by only ~15% (3.72 vs
4.38) and the absolute KL stays catastrophic (~3.7 β ~200Γ the q4_mid precision baseline of 0.019; any
usable bar is <~0.1). Training KL barely moved (7.3β~3.7, noisy/stuck). **There is no 50%-skippable token
subset the frozen network tolerates** β *which* tokens you skip barely matters, exactly as S1b/S2 predict
(distributed, compounding contribution).
**Why, and what's left:** classic MoD (Raposo et al.) trains the router *jointly with the whole network
from the start*, so the model learns to route around skips. A post-hoc router on a frozen pretrained base
has no such adaptation, and lands far outside the model's operating distribution. Making MoD work here
would require **full-model (or heavy-LoRA) joint finetuning** β the expensive retrain the whole idea was
meant to avoid, and the S1b/S2 compounding evidence makes the payoff doubtful. **Verdict: the entire
training-cheap adaptive-compute line is closed on A4B.** A full MoD finetune remains theoretically open
but is a different, costly project, not pursued.
Other later angles (parked): Any-Precision MSB-first bitplane weights (training-free but no GGUF path);
CATS-style near-training-free activation sparsity (changes the activation function).
|