# Adaptive compute — depth-skip & dynamic-precision feasibility > Status: **TRAINING-FREE VARIANTS CLOSED** (2026-06-30). Three cheap PyTorch forward-hook probes > (S1/S1b/S2) on Gemma-4 A4B establish that *static, predictor-free* compute-skipping and > dynamic weight-precision have **no headroom** on this model as deployed. One variant is not yet > ruled out — a **trained** Mixture-of-Depths router (§6) — and is being prototyped. This note is the > record so the line is not re-litigated from scratch. This is the *compute-axis* companion to [decode-levers.md](decode-levers.md) (which makes the compute opencoti *does* run faster) and [poly_kv.md] / [context.md](context.md) (the *KV/memory* axis). Where those find wins, this documents a **negative result**: you cannot cheaply *remove* compute on A4B. The motivating idea (user, 2026-06-30): *"I have q8 weights; if I can tell in advance that q4 would not change an activation meaningfully, I skip the high-precision compute and only activate the neurons that change something."* Plus the depth-skip cousin: skip whole layers / MLP sub-blocks that contribute little. The user explicitly distrusted the usual predictor premise (*"'10% of neurons carry 90% of the mass' — I don't think it's ever true"*) — and the data backs that distrust. --- ## 1. Prior art (field scan) | Method | Mechanism | Training? | Batch-friendly? | Note | |---|---|---|---|---| | DejaVu | contextual sparsity + low-rank predictor | no (predictor) | **no** | union of active neurons grows with batch | | PowerInfer | hot/cold neuron split | needs ReLU model | **no** | same batch erosion | | Mixture-of-Depths (MoD) | router picks top-k *tokens* per layer | **yes** | **yes** (capacity-based) | Router-Tuning = cheap post-hoc variant → §6 | | LayerSkip / SkipDecode | early-exit / per-token depth | **yes** (trained) | SkipDecode yes | needs finetune | | Any-Precision LLM | MSB-first bitplane nested quant | no (training-free) | yes | exact truncation; no GGUF path | | DeltaLLM / temporal-delta | reuse previous-step state | no | no | KV-side; no reported wall-clock decode win | | ProSparse / TurboSparse / CATS | ReLUfication | finetune (CATS near-free) | yes | changes activation fn | Takeaway going in: the **training-free + batch-friendly** quadrant (what opencoti wants for [multi-session serving](#)) was empty in the literature. S1–S2 tested whether A4B's own structure opens it anyway. It does not. --- ## 2. Method PyTorch forward-hooks on the real **Gemma-4 26B-A4B-it** (bf16, `/srv/ml/models/base/…` on bs2 GPU0; 30 decoder layers, d_model 2816; 3 realistic coding prompts, greedy decode). Architecture-level property → bf16 weights are representative. Probes are in `.opencoti/dynq/{s1_probe,s1b_precision, s2_skip}.py`; raw JSON `{s1,s1b,s2}_out.json`. Correctness gate = mean **KL(bf16 ‖ condition)** on gen-position logits (greedy token-match is FP-unstable and inadmissible — see bug-270/356). Notation: `r_layer = ‖h_out−h_in‖/‖h_in‖` (whole-block residual contribution); `r_attn`, `r_mlp` = sub-block-output / layer-input norm proxies; `cosΔh` = cosine of consecutive decode tokens' residual input; `Δh-sparsity` = fraction of dims for 90% of ‖Δh‖². --- ## 3. S1 — is there a lazy band? | band | r_attn (p50) | r_mlp (p50) | reading | |---|---|---|---| | L0 | 1.76 | 11.9 | embedding bootstrap | | L3–8 | 0.4–1.3 | 0.19–0.76 | both active | | **L9–19** | 0.30–0.57 | **0.011–0.051** | MLP near-dead, attention alive | | L23–28 | 0.53–0.96 | 0.82 → **7.14** | late MLP does the heavy lifting | - `r_layer` p50 never drops below ~0.26 → **no skippable whole-block dead band.** - The only lazy region is the **MLP (MoE) sub-block of mid layers ~9–19** (1–5% of residual). Compute is *back-loaded*: late MLPs explode. This *looked* like a batch-friendly static-skip target — §5 kills it. - **Temporal-delta (#4-A) — half-refuted.** `cosΔh` is high (0.75–0.85) at L10–19, but `Δh-sparsity` is sparse (2–5% dims) only in *early* layers (L1–9, where cos is low 0.13–0.65) and densifies to 20–40% exactly where it's stable. The two conditions temporal-delta needs (stability *and* sparse delta) **anti-correlate**; no general "compute only W·Δh" win. Batch-fragile besides. --- ## 4. S1b — can the lazy band run at lower weight precision? (the precision-guard, #4-B) Mean KL(bf16 ‖ cond), MLP weights blockwise-quantized (q4_0/q6_0-style, block 32, symmetric absmax): | condition | layers | mean KL | |---|---|---| | q6_all | 30 | **0.0042** | | q4_early (0–8) | 9 | 0.0040 | | q4_late (23–29) | 7 | 0.0135 | | **q4_mid (9–19)** | 11 | **0.0188** | | q4_all | 30 | 0.0359 | Per-layer q4 MLP rel-err *rises* 0.07 (early) → 0.15–0.24 (mid/late). **Refuted.** q4 on the *lazy* mid-band costs **more** than on the *heavy* late-band — the small MLPs are precision-*sensitive*. Small contribution ⇏ precision-robust; you cannot use "this contributes little" as a guard for "q4 is safe here." **Premise-breaker:** production opencoti already ships **Q4_K_M** — the idea assumed q8 weights as the expensive baseline to skip down from; there is no q8 compute to elide. (q6 is ~9× cheaper than q4 in KL, but that points the wrong way — more memory, and prod is already q4.) --- ## 5. S2 — can the lazy band be skipped? (MoD-static, #3) Mean KL(bf16 ‖ cond), MLP output zeroed via hook (no weight mutation): - **Single-layer skip:** most mid layers **0.13–0.23** (≈10× worse than q4-ing *all* 11 at 0.019). Cheap singles are scattered (L01 0.028, L04 0.04, L13 0.054). Critical layers: L02 **2.0**, L29 **2.1**, L26/L28 ~1.4 — skipping these alone wrecks the model. - **Cumulative bands EXPLODE:** mid_9_19 **12.4**, late_23_29 **19.0**, early_0_8 **12.7**, and even the **10 individually-cheapest layers → 7.27**. Errors **compound super-linearly.** **Refuted.** The 1–5% mid-MLP contributions are tiny per-step but essential in aggregate; they cannot be skipped, individually barely and in any band catastrophically. The S1 "batch-friendly mid-MLP static skip" hope is dead. --- ## 6. Verdict + the one open angle **The unifying law:** on A4B, a sublayer's small per-step contribution predicts **neither** skippability (S2) **nor** precision-robustness (S1b). The training-free / predictor-free adaptive-compute family — DejaVu-style sparsity, depth-skip, dynamic weight-precision, temporal-delta — has **no headroom** here. This closes the *compute axis* for "more concurrent agentic sessions per card"; the real headroom stays on the **KV/memory axis** (turbo/TCQ KV-quant, DCA, PolyKV/SharedKVPool prefix-sharing) + MTP for decode throughput — all shipped or in-flight. **Caveats (scope of the closure):** greedy-KL on 3 short coding prompts (not RULER/niah) — but cumulative KL 7–19 is far past any rescue threshold; A4B is MoE (the compounding logic is general); the probes test **static / training-free** skip+precision only. ## 7. Trained MoD router (the open angle) — router-only FAILS `.opencoti/dynq/mod_router.py`: per-layer linear gate on band L9–22, **base frozen**, trained by self-distillation KL(teacher_full ‖ student) with **hard top-k (CAP 0.5) selection in the loop** (kept tokens scaled by the router sigmoid for gradient, dropped → 0). 48 teacher continuations, 250 steps, LR 1e-2. Eval = mean KL at hard CAP 0.5. | config | KL | keep | |---|---|---| | pre-train (untrained gate) | 7.00 | 0.50 | | random 50%-skip control | 4.38 | 0.50 | | **trained router** | **3.72** | 0.50 | | (S2 static all-skip mid-band) | 12.4 | — | **Router-only MoD does not work on A4B.** A trained router beats random selection by only ~15% (3.72 vs 4.38) and the absolute KL stays catastrophic (~3.7 — ~200× the q4_mid precision baseline of 0.019; any usable bar is <~0.1). Training KL barely moved (7.3→~3.7, noisy/stuck). **There is no 50%-skippable token subset the frozen network tolerates** — *which* tokens you skip barely matters, exactly as S1b/S2 predict (distributed, compounding contribution). **Why, and what's left:** classic MoD (Raposo et al.) trains the router *jointly with the whole network from the start*, so the model learns to route around skips. A post-hoc router on a frozen pretrained base has no such adaptation, and lands far outside the model's operating distribution. Making MoD work here would require **full-model (or heavy-LoRA) joint finetuning** — the expensive retrain the whole idea was meant to avoid, and the S1b/S2 compounding evidence makes the payoff doubtful. **Verdict: the entire training-cheap adaptive-compute line is closed on A4B.** A full MoD finetune remains theoretically open but is a different, costly project, not pursued. Other later angles (parked): Any-Precision MSB-first bitplane weights (training-free but no GGUF path); CATS-style near-training-free activation sparsity (changes the activation function).