Adaptive compute β depth-skip & dynamic-precision feasibility
Status: TRAINING-FREE VARIANTS CLOSED (2026-06-30). Three cheap PyTorch forward-hook probes (S1/S1b/S2) on Gemma-4 A4B establish that static, predictor-free compute-skipping and dynamic weight-precision have no headroom on this model as deployed. One variant is not yet ruled out β a trained Mixture-of-Depths router (Β§6) β and is being prototyped. This note is the record so the line is not re-litigated from scratch.
This is the compute-axis companion to decode-levers.md (which makes the compute opencoti does run faster) and [poly_kv.md] / context.md (the KV/memory axis). Where those find wins, this documents a negative result: you cannot cheaply remove compute on A4B.
The motivating idea (user, 2026-06-30): "I have q8 weights; if I can tell in advance that q4 would not change an activation meaningfully, I skip the high-precision compute and only activate the neurons that change something." Plus the depth-skip cousin: skip whole layers / MLP sub-blocks that contribute little. The user explicitly distrusted the usual predictor premise ("'10% of neurons carry 90% of the mass' β I don't think it's ever true") β and the data backs that distrust.
1. Prior art (field scan)
| Method | Mechanism | Training? | Batch-friendly? | Note |
|---|---|---|---|---|
| DejaVu | contextual sparsity + low-rank predictor | no (predictor) | no | union of active neurons grows with batch |
| PowerInfer | hot/cold neuron split | needs ReLU model | no | same batch erosion |
| Mixture-of-Depths (MoD) | router picks top-k tokens per layer | yes | yes (capacity-based) | Router-Tuning = cheap post-hoc variant β Β§6 |
| LayerSkip / SkipDecode | early-exit / per-token depth | yes (trained) | SkipDecode yes | needs finetune |
| Any-Precision LLM | MSB-first bitplane nested quant | no (training-free) | yes | exact truncation; no GGUF path |
| DeltaLLM / temporal-delta | reuse previous-step state | no | no | KV-side; no reported wall-clock decode win |
| ProSparse / TurboSparse / CATS | ReLUfication | finetune (CATS near-free) | yes | changes activation fn |
Takeaway going in: the training-free + batch-friendly quadrant (what opencoti wants for multi-session serving) was empty in the literature. S1βS2 tested whether A4B's own structure opens it anyway. It does not.
2. Method
PyTorch forward-hooks on the real Gemma-4 26B-A4B-it (bf16, /srv/ml/models/base/β¦ on bs2 GPU0;
30 decoder layers, d_model 2816; 3 realistic coding prompts, greedy decode). Architecture-level
property β bf16 weights are representative. Probes are in .opencoti/dynq/{s1_probe,s1b_precision, s2_skip}.py; raw JSON {s1,s1b,s2}_out.json. Correctness gate = mean KL(bf16 β condition) on
gen-position logits (greedy token-match is FP-unstable and inadmissible β see bug-270/356).
Notation: r_layer = βh_outβh_inβ/βh_inβ (whole-block residual contribution); r_attn, r_mlp =
sub-block-output / layer-input norm proxies; cosΞh = cosine of consecutive decode tokens' residual
input; Ξh-sparsity = fraction of dims for 90% of βΞhβΒ².
3. S1 β is there a lazy band?
| band | r_attn (p50) | r_mlp (p50) | reading |
|---|---|---|---|
| L0 | 1.76 | 11.9 | embedding bootstrap |
| L3β8 | 0.4β1.3 | 0.19β0.76 | both active |
| L9β19 | 0.30β0.57 | 0.011β0.051 | MLP near-dead, attention alive |
| L23β28 | 0.53β0.96 | 0.82 β 7.14 | late MLP does the heavy lifting |
r_layerp50 never drops below ~0.26 β no skippable whole-block dead band.- The only lazy region is the MLP (MoE) sub-block of mid layers ~9β19 (1β5% of residual). Compute is back-loaded: late MLPs explode. This looked like a batch-friendly static-skip target β Β§5 kills it.
- Temporal-delta (#4-A) β half-refuted.
cosΞhis high (0.75β0.85) at L10β19, butΞh-sparsityis sparse (2β5% dims) only in early layers (L1β9, where cos is low 0.13β0.65) and densifies to 20β40% exactly where it's stable. The two conditions temporal-delta needs (stability and sparse delta) anti-correlate; no general "compute only WΒ·Ξh" win. Batch-fragile besides.
4. S1b β can the lazy band run at lower weight precision? (the precision-guard, #4-B)
Mean KL(bf16 β cond), MLP weights blockwise-quantized (q4_0/q6_0-style, block 32, symmetric absmax):
| condition | layers | mean KL |
|---|---|---|
| q6_all | 30 | 0.0042 |
| q4_early (0β8) | 9 | 0.0040 |
| q4_late (23β29) | 7 | 0.0135 |
| q4_mid (9β19) | 11 | 0.0188 |
| q4_all | 30 | 0.0359 |
Per-layer q4 MLP rel-err rises 0.07 (early) β 0.15β0.24 (mid/late).
Refuted. q4 on the lazy mid-band costs more than on the heavy late-band β the small MLPs are precision-sensitive. Small contribution β precision-robust; you cannot use "this contributes little" as a guard for "q4 is safe here." Premise-breaker: production opencoti already ships Q4_K_M β the idea assumed q8 weights as the expensive baseline to skip down from; there is no q8 compute to elide. (q6 is ~9Γ cheaper than q4 in KL, but that points the wrong way β more memory, and prod is already q4.)
5. S2 β can the lazy band be skipped? (MoD-static, #3)
Mean KL(bf16 β cond), MLP output zeroed via hook (no weight mutation):
- Single-layer skip: most mid layers 0.13β0.23 (β10Γ worse than q4-ing all 11 at 0.019). Cheap singles are scattered (L01 0.028, L04 0.04, L13 0.054). Critical layers: L02 2.0, L29 2.1, L26/L28 ~1.4 β skipping these alone wrecks the model.
- Cumulative bands EXPLODE: mid_9_19 12.4, late_23_29 19.0, early_0_8 12.7, and even the 10 individually-cheapest layers β 7.27. Errors compound super-linearly.
Refuted. The 1β5% mid-MLP contributions are tiny per-step but essential in aggregate; they cannot be skipped, individually barely and in any band catastrophically. The S1 "batch-friendly mid-MLP static skip" hope is dead.
6. Verdict + the one open angle
The unifying law: on A4B, a sublayer's small per-step contribution predicts neither skippability (S2) nor precision-robustness (S1b). The training-free / predictor-free adaptive-compute family β DejaVu-style sparsity, depth-skip, dynamic weight-precision, temporal-delta β has no headroom here. This closes the compute axis for "more concurrent agentic sessions per card"; the real headroom stays on the KV/memory axis (turbo/TCQ KV-quant, DCA, PolyKV/SharedKVPool prefix-sharing) + MTP for decode throughput β all shipped or in-flight.
Caveats (scope of the closure): greedy-KL on 3 short coding prompts (not RULER/niah) β but cumulative KL 7β19 is far past any rescue threshold; A4B is MoE (the compounding logic is general); the probes test static / training-free skip+precision only.
7. Trained MoD router (the open angle) β router-only FAILS
.opencoti/dynq/mod_router.py: per-layer linear gate on band L9β22, base frozen, trained by
self-distillation KL(teacher_full β student) with hard top-k (CAP 0.5) selection in the loop (kept
tokens scaled by the router sigmoid for gradient, dropped β 0). 48 teacher continuations, 250 steps,
LR 1e-2. Eval = mean KL at hard CAP 0.5.
| config | KL | keep |
|---|---|---|
| pre-train (untrained gate) | 7.00 | 0.50 |
| random 50%-skip control | 4.38 | 0.50 |
| trained router | 3.72 | 0.50 |
| (S2 static all-skip mid-band) | 12.4 | β |
Router-only MoD does not work on A4B. A trained router beats random selection by only 15% (3.72 vs
4.38) and the absolute KL stays catastrophic (3.7 β 200Γ the q4_mid precision baseline of 0.019; any
usable bar is <0.1). Training KL barely moved (7.3β~3.7, noisy/stuck). There is no 50%-skippable token
subset the frozen network tolerates β which tokens you skip barely matters, exactly as S1b/S2 predict
(distributed, compounding contribution).
Why, and what's left: classic MoD (Raposo et al.) trains the router jointly with the whole network from the start, so the model learns to route around skips. A post-hoc router on a frozen pretrained base has no such adaptation, and lands far outside the model's operating distribution. Making MoD work here would require full-model (or heavy-LoRA) joint finetuning β the expensive retrain the whole idea was meant to avoid, and the S1b/S2 compounding evidence makes the payoff doubtful. Verdict: the entire training-cheap adaptive-compute line is closed on A4B. A full MoD finetune remains theoretically open but is a different, costly project, not pursued.
Other later angles (parked): Any-Precision MSB-first bitplane weights (training-free but no GGUF path); CATS-style near-training-free activation sparsity (changes the activation function).