| # Adaptive compute β depth-skip & dynamic-precision feasibility |
|
|
| > Status: **TRAINING-FREE VARIANTS CLOSED** (2026-06-30). Three cheap PyTorch forward-hook probes |
| > (S1/S1b/S2) on Gemma-4 A4B establish that *static, predictor-free* compute-skipping and |
| > dynamic weight-precision have **no headroom** on this model as deployed. One variant is not yet |
| > ruled out β a **trained** Mixture-of-Depths router (Β§6) β and is being prototyped. This note is the |
| > record so the line is not re-litigated from scratch. |
|
|
| This is the *compute-axis* companion to [decode-levers.md](decode-levers.md) (which makes the compute |
| opencoti *does* run faster) and [poly_kv.md] / [context.md](context.md) (the *KV/memory* axis). Where |
| those find wins, this documents a **negative result**: you cannot cheaply *remove* compute on A4B. |
|
|
| The motivating idea (user, 2026-06-30): *"I have q8 weights; if I can tell in advance that q4 would |
| not change an activation meaningfully, I skip the high-precision compute and only activate the neurons |
| that change something."* Plus the depth-skip cousin: skip whole layers / MLP sub-blocks that contribute |
| little. The user explicitly distrusted the usual predictor premise (*"'10% of neurons carry 90% of the |
| mass' β I don't think it's ever true"*) β and the data backs that distrust. |
|
|
| --- |
|
|
| ## 1. Prior art (field scan) |
|
|
| | Method | Mechanism | Training? | Batch-friendly? | Note | |
| |---|---|---|---|---| |
| | DejaVu | contextual sparsity + low-rank predictor | no (predictor) | **no** | union of active neurons grows with batch | |
| | PowerInfer | hot/cold neuron split | needs ReLU model | **no** | same batch erosion | |
| | Mixture-of-Depths (MoD) | router picks top-k *tokens* per layer | **yes** | **yes** (capacity-based) | Router-Tuning = cheap post-hoc variant β Β§6 | |
| | LayerSkip / SkipDecode | early-exit / per-token depth | **yes** (trained) | SkipDecode yes | needs finetune | |
| | Any-Precision LLM | MSB-first bitplane nested quant | no (training-free) | yes | exact truncation; no GGUF path | |
| | DeltaLLM / temporal-delta | reuse previous-step state | no | no | KV-side; no reported wall-clock decode win | |
| | ProSparse / TurboSparse / CATS | ReLUfication | finetune (CATS near-free) | yes | changes activation fn | |
|
|
| Takeaway going in: the **training-free + batch-friendly** quadrant (what opencoti wants for |
| [multi-session serving](#)) was empty in the literature. S1βS2 tested whether A4B's own structure |
| opens it anyway. It does not. |
|
|
| --- |
|
|
| ## 2. Method |
|
|
| PyTorch forward-hooks on the real **Gemma-4 26B-A4B-it** (bf16, `/srv/ml/models/base/β¦` on bs2 GPU0; |
| 30 decoder layers, d_model 2816; 3 realistic coding prompts, greedy decode). Architecture-level |
| property β bf16 weights are representative. Probes are in `.opencoti/dynq/{s1_probe,s1b_precision, |
| s2_skip}.py`; raw JSON `{s1,s1b,s2}_out.json`. Correctness gate = mean **KL(bf16 β condition)** on |
| gen-position logits (greedy token-match is FP-unstable and inadmissible β see bug-270/356). |
| |
| Notation: `r_layer = βh_outβh_inβ/βh_inβ` (whole-block residual contribution); `r_attn`, `r_mlp` = |
| sub-block-output / layer-input norm proxies; `cosΞh` = cosine of consecutive decode tokens' residual |
| input; `Ξh-sparsity` = fraction of dims for 90% of βΞhβΒ². |
| |
| --- |
| |
| ## 3. S1 β is there a lazy band? |
| |
| | band | r_attn (p50) | r_mlp (p50) | reading | |
| |---|---|---|---| |
| | L0 | 1.76 | 11.9 | embedding bootstrap | |
| | L3β8 | 0.4β1.3 | 0.19β0.76 | both active | |
| | **L9β19** | 0.30β0.57 | **0.011β0.051** | MLP near-dead, attention alive | |
| | L23β28 | 0.53β0.96 | 0.82 β **7.14** | late MLP does the heavy lifting | |
| |
| - `r_layer` p50 never drops below ~0.26 β **no skippable whole-block dead band.** |
| - The only lazy region is the **MLP (MoE) sub-block of mid layers ~9β19** (1β5% of residual). Compute is |
| *back-loaded*: late MLPs explode. This *looked* like a batch-friendly static-skip target β Β§5 kills it. |
| - **Temporal-delta (#4-A) β half-refuted.** `cosΞh` is high (0.75β0.85) at L10β19, but `Ξh-sparsity` is |
| sparse (2β5% dims) only in *early* layers (L1β9, where cos is low 0.13β0.65) and densifies to 20β40% |
| exactly where it's stable. The two conditions temporal-delta needs (stability *and* sparse delta) |
| **anti-correlate**; no general "compute only WΒ·Ξh" win. Batch-fragile besides. |
|
|
| --- |
|
|
| ## 4. S1b β can the lazy band run at lower weight precision? (the precision-guard, #4-B) |
|
|
| Mean KL(bf16 β cond), MLP weights blockwise-quantized (q4_0/q6_0-style, block 32, symmetric absmax): |
|
|
| | condition | layers | mean KL | |
| |---|---|---| |
| | q6_all | 30 | **0.0042** | |
| | q4_early (0β8) | 9 | 0.0040 | |
| | q4_late (23β29) | 7 | 0.0135 | |
| | **q4_mid (9β19)** | 11 | **0.0188** | |
| | q4_all | 30 | 0.0359 | |
|
|
| Per-layer q4 MLP rel-err *rises* 0.07 (early) β 0.15β0.24 (mid/late). |
|
|
| **Refuted.** q4 on the *lazy* mid-band costs **more** than on the *heavy* late-band β the small MLPs are |
| precision-*sensitive*. Small contribution β precision-robust; you cannot use "this contributes little" |
| as a guard for "q4 is safe here." **Premise-breaker:** production opencoti already ships **Q4_K_M** β the |
| idea assumed q8 weights as the expensive baseline to skip down from; there is no q8 compute to elide. |
| (q6 is ~9Γ cheaper than q4 in KL, but that points the wrong way β more memory, and prod is already q4.) |
|
|
| --- |
|
|
| ## 5. S2 β can the lazy band be skipped? (MoD-static, #3) |
|
|
| Mean KL(bf16 β cond), MLP output zeroed via hook (no weight mutation): |
|
|
| - **Single-layer skip:** most mid layers **0.13β0.23** (β10Γ worse than q4-ing *all* 11 at 0.019). Cheap |
| singles are scattered (L01 0.028, L04 0.04, L13 0.054). Critical layers: L02 **2.0**, L29 **2.1**, |
| L26/L28 ~1.4 β skipping these alone wrecks the model. |
| - **Cumulative bands EXPLODE:** mid_9_19 **12.4**, late_23_29 **19.0**, early_0_8 **12.7**, and even the |
| **10 individually-cheapest layers β 7.27**. Errors **compound super-linearly.** |
|
|
| **Refuted.** The 1β5% mid-MLP contributions are tiny per-step but essential in aggregate; they cannot be |
| skipped, individually barely and in any band catastrophically. The S1 "batch-friendly mid-MLP static |
| skip" hope is dead. |
|
|
| --- |
|
|
| ## 6. Verdict + the one open angle |
|
|
| **The unifying law:** on A4B, a sublayer's small per-step contribution predicts **neither** skippability |
| (S2) **nor** precision-robustness (S1b). The training-free / predictor-free adaptive-compute family β |
| DejaVu-style sparsity, depth-skip, dynamic weight-precision, temporal-delta β has **no headroom** here. |
| This closes the *compute axis* for "more concurrent agentic sessions per card"; the real headroom stays |
| on the **KV/memory axis** (turbo/TCQ KV-quant, DCA, PolyKV/SharedKVPool prefix-sharing) + MTP for decode |
| throughput β all shipped or in-flight. |
|
|
| **Caveats (scope of the closure):** greedy-KL on 3 short coding prompts (not RULER/niah) β but cumulative |
| KL 7β19 is far past any rescue threshold; A4B is MoE (the compounding logic is general); the probes test |
| **static / training-free** skip+precision only. |
|
|
| ## 7. Trained MoD router (the open angle) β router-only FAILS |
|
|
| `.opencoti/dynq/mod_router.py`: per-layer linear gate on band L9β22, **base frozen**, trained by |
| self-distillation KL(teacher_full β student) with **hard top-k (CAP 0.5) selection in the loop** (kept |
| tokens scaled by the router sigmoid for gradient, dropped β 0). 48 teacher continuations, 250 steps, |
| LR 1e-2. Eval = mean KL at hard CAP 0.5. |
| |
| | config | KL | keep | |
| |---|---|---| |
| | pre-train (untrained gate) | 7.00 | 0.50 | |
| | random 50%-skip control | 4.38 | 0.50 | |
| | **trained router** | **3.72** | 0.50 | |
| | (S2 static all-skip mid-band) | 12.4 | β | |
| |
| **Router-only MoD does not work on A4B.** A trained router beats random selection by only ~15% (3.72 vs |
| 4.38) and the absolute KL stays catastrophic (~3.7 β ~200Γ the q4_mid precision baseline of 0.019; any |
| usable bar is <~0.1). Training KL barely moved (7.3β~3.7, noisy/stuck). **There is no 50%-skippable token |
| subset the frozen network tolerates** β *which* tokens you skip barely matters, exactly as S1b/S2 predict |
| (distributed, compounding contribution). |
|
|
| **Why, and what's left:** classic MoD (Raposo et al.) trains the router *jointly with the whole network |
| from the start*, so the model learns to route around skips. A post-hoc router on a frozen pretrained base |
| has no such adaptation, and lands far outside the model's operating distribution. Making MoD work here |
| would require **full-model (or heavy-LoRA) joint finetuning** β the expensive retrain the whole idea was |
| meant to avoid, and the S1b/S2 compounding evidence makes the payoff doubtful. **Verdict: the entire |
| training-cheap adaptive-compute line is closed on A4B.** A full MoD finetune remains theoretically open |
| but is a different, costly project, not pursued. |
|
|
| Other later angles (parked): Any-Precision MSB-first bitplane weights (training-free but no GGUF path); |
| CATS-style near-training-free activation sparsity (changes the activation function). |
|
|