# Real-Weights Demo: Kimi K2.6 Layer 0 Attention Distribution This walkthrough runs the eviction policy on **real Kimi K2.6 attention distributions**, not the synthetic distributions used in the multi-step validation in [02_validation_results.md](02_validation_results.md). It loads the actual layer-0 weights from the published Kimi K2.6 checkpoint, runs a single full-prefix forward over 256 synthetic input tokens, captures the real attention score distribution, and applies the H2O heavy-hitter eviction policy to it. The point: validate the eviction policy makes sensible decisions when fed actual Kimi attention scores (vs. random distributions). ## What the demo does 1. Loads the canonical Kimi K2.6 architecture config (61L, 64H, MLA with kv_lora_rank=512, qk_dim=192, v_dim=128). 2. Instantiates a single `transformers.models.deepseek_v3.modeling_deepseek_v3.DeepseekV3Attention` module (layer index 0). 3. Loads the actual layer-0 attention weights from `model-00001-of-000064.safetensors` of the published Kimi K2.6 checkpoint. All 7 weights match cleanly: `q_a_proj`, `q_b_proj`, `q_a_layernorm`, `kv_a_proj_with_mqa`, `kv_b_proj`, `kv_a_layernorm`, `o_proj`. 4. Runs one forward pass with `seq_len=256` synthetic input embeddings, captures the `attention_weights` tensor of shape `[1, 64 heads, 256, 256]`. 5. Computes per-kv-token cumulative attention mass (sum across heads and across all queries that attend to that kv position). 6. Applies the H2O policy: keep `n_sink=4` start tokens, `n_recent=32` end tokens, plus the top `budget=64` heavy hitters from the middle. Mark all others as evicted. 7. Reports the score distribution and the kept/evicted ratio. ## Hardware NVIDIA TITAN RTX (24 GB), BF16 compute. The single-layer forward over 256 tokens completes in **0.07 seconds**. ## Configuration ``` config: 61L hidden=7168 heads=64 qk_dim=192 v_dim=128 kv_lora_rank=512 layer params: 101,124,096 seq_len: 256 budget: 64 (heavy-hitter slots in the middle) n_sink: 4 (always kept, indices 0..3) n_recent: 32 (always kept, last 32) ``` ## Results ``` attn_out shape: torch.Size([1, 256, 7168]) attn_weights shape: torch.Size([1, 64, 256, 256]) score per token: shape=(256,) score range: [0.336, 350.870] score mean: 64.000 score std: 64.217 H2O eviction policy applied: kept 100 of 256 tokens (39.1%) sinks 4 (indices 0..3) heavy 64 of 220 middle tokens chosen recent 32 (last 32) evicted 156 (60.9%) top 10 heavy-hitter scores: 319.18 277.95 257.06 238.43 237.26 222.64 210.78 208.90 207.39 189.85 mean score of heavy-hitters kept: 141.082 mean score of evicted tokens: 38.583 heavy / evicted score ratio: 3.66x ``` ## Interpretation The attention distribution from real Kimi K2.6 layer 0 is **strongly heavy-tailed**: - Score spread: ~1000x between min (0.336) and max (350.87). - Top heavy-hitter scores are roughly 10x the mean (319 vs. mean of 64). - The standard deviation (64) is comparable to the mean — wide variance, lots of structure for the eviction policy to exploit. The H2O policy correctly identifies the high-attention tokens: tokens it keeps as heavy-hitters score on average **3.66 times higher** than tokens it evicts. This is a meaningful gap — the policy is not just shuffling random selections. This validates that on a real frontier-scale MLA attention layer, the H2O recipe (top-k by accumulated attention mass + sinks + recent window) makes sensible decisions about which tokens to keep when cache pressure forces eviction. ## What this does NOT prove - **Output quality:** we use random input embeddings, so the attention layer's output is not meaningful text. We're only validating the *attention distribution* and the *eviction decision*, not generation quality. - **Multi-layer consistency:** layer 0's attention pattern may differ from layers 30 or 60. A full-model evaluation would average across all 61 layers; we only loaded 1. - **End-to-end with cache eviction:** we apply the policy to a fully-populated 256-token cache after the forward; we do not exercise the cache-management plumbing (the `install_kv_eviction` patch on transformers 5.x DynamicCache is still on the roadmap — see README). ## Reproducing The Kimi K2.6 weights are published by Moonshot AI on HuggingFace: [`moonshotai/Kimi-K2.6`](https://huggingface.co/moonshotai/Kimi-K2.6). The script in `scripts/kimi_layer_eviction_demo.py` runs as-is on a host with: - transformers >= 5.0 (DeepseekV3 model class) - safetensors - torch + CUDA-capable GPU with >= 4 GB VRAM (single-layer fits comfortably) - ~1 GB of read access to `model-00001-of-000064.safetensors` ```bash # After cloning Kimi K2.6 to /path/to/Kimi-K2.6/ git clone https://huggingface.co/GenomaLabs-com/kv-cache-eviction-mla cd kv-cache-eviction-mla # Edit KIMI_PATH in scripts/kimi_layer_eviction_demo.py to point at your checkpoint python scripts/kimi_layer_eviction_demo.py ``` Output: `results/kimi_layer_eviction_demo.csv` (256 rows: token_idx, attention_score, kept, category). ## Per-token CSV available `results/kimi_layer_eviction_demo.csv` contains per-token data for downstream analysis: | Column | Description | |---|---| | `token_idx` | Position in the sequence (0..255) | | `attention_score` | Cumulative attention mass received from all queries / heads | | `kept` | True if the H2O policy retains this token | | `category` | `sink` / `recent` / `heavy` / `evicted` | This CSV can be plotted (score vs index, colored by category) to visualize the eviction decision against the actual attention distribution.