kv-cache-eviction-mla / notebooks /03_kimi_real_weights_demo.md
GENOMA LABS / research
B3a-pivot: real Kimi K2.6 weights demo - eviction policy validated on actual MLA attention distribution
aac6c2e
|
Raw History Blame Contribute Delete
5.71 kB

Real-Weights Demo: Kimi K2.6 Layer 0 Attention Distribution

This walkthrough runs the eviction policy on real Kimi K2.6 attention distributions, not the synthetic distributions used in the multi-step validation in 02_validation_results.md. It loads the actual layer-0 weights from the published Kimi K2.6 checkpoint, runs a single full-prefix forward over 256 synthetic input tokens, captures the real attention score distribution, and applies the H2O heavy-hitter eviction policy to it.

The point: validate the eviction policy makes sensible decisions when fed actual Kimi attention scores (vs. random distributions).

What the demo does

  1. Loads the canonical Kimi K2.6 architecture config (61L, 64H, MLA with kv_lora_rank=512, qk_dim=192, v_dim=128).
  2. Instantiates a single transformers.models.deepseek_v3.modeling_deepseek_v3.DeepseekV3Attention module (layer index 0).
  3. Loads the actual layer-0 attention weights from model-00001-of-000064.safetensors of the published Kimi K2.6 checkpoint. All 7 weights match cleanly: q_a_proj, q_b_proj, q_a_layernorm, kv_a_proj_with_mqa, kv_b_proj, kv_a_layernorm, o_proj.
  4. Runs one forward pass with seq_len=256 synthetic input embeddings, captures the attention_weights tensor of shape [1, 64 heads, 256, 256].
  5. Computes per-kv-token cumulative attention mass (sum across heads and across all queries that attend to that kv position).
  6. Applies the H2O policy: keep n_sink=4 start tokens, n_recent=32 end tokens, plus the top budget=64 heavy hitters from the middle. Mark all others as evicted.
  7. Reports the score distribution and the kept/evicted ratio.

Hardware

NVIDIA TITAN RTX (24 GB), BF16 compute. The single-layer forward over 256 tokens completes in 0.07 seconds.

Configuration

config:        61L  hidden=7168  heads=64
               qk_dim=192  v_dim=128  kv_lora_rank=512
layer params:  101,124,096
seq_len:       256
budget:        64    (heavy-hitter slots in the middle)
n_sink:        4     (always kept, indices 0..3)
n_recent:      32    (always kept, last 32)

Results

attn_out shape:        torch.Size([1, 256, 7168])
attn_weights shape:    torch.Size([1, 64, 256, 256])
score per token:       shape=(256,)
  score range:         [0.336, 350.870]
  score mean:          64.000
  score std:           64.217

H2O eviction policy applied:
  kept       100 of 256 tokens (39.1%)
  sinks       4 (indices 0..3)
  heavy      64 of 220 middle tokens chosen
  recent     32 (last 32)
  evicted    156 (60.9%)

top 10 heavy-hitter scores:
  319.18  277.95  257.06  238.43  237.26
  222.64  210.78  208.90  207.39  189.85

mean score of heavy-hitters kept:  141.082
mean score of evicted tokens:        38.583
heavy / evicted score ratio:          3.66x

Interpretation

The attention distribution from real Kimi K2.6 layer 0 is strongly heavy-tailed:

  • Score spread: ~1000x between min (0.336) and max (350.87).
  • Top heavy-hitter scores are roughly 10x the mean (319 vs. mean of 64).
  • The standard deviation (64) is comparable to the mean — wide variance, lots of structure for the eviction policy to exploit.

The H2O policy correctly identifies the high-attention tokens: tokens it keeps as heavy-hitters score on average 3.66 times higher than tokens it evicts. This is a meaningful gap — the policy is not just shuffling random selections.

This validates that on a real frontier-scale MLA attention layer, the H2O recipe (top-k by accumulated attention mass + sinks + recent window) makes sensible decisions about which tokens to keep when cache pressure forces eviction.

What this does NOT prove

  • Output quality: we use random input embeddings, so the attention layer's output is not meaningful text. We're only validating the attention distribution and the eviction decision, not generation quality.
  • Multi-layer consistency: layer 0's attention pattern may differ from layers 30 or 60. A full-model evaluation would average across all 61 layers; we only loaded 1.
  • End-to-end with cache eviction: we apply the policy to a fully-populated 256-token cache after the forward; we do not exercise the cache-management plumbing (the install_kv_eviction patch on transformers 5.x DynamicCache is still on the roadmap — see README).

Reproducing

The Kimi K2.6 weights are published by Moonshot AI on HuggingFace: moonshotai/Kimi-K2.6. The script in scripts/kimi_layer_eviction_demo.py runs as-is on a host with:

  • transformers >= 5.0 (DeepseekV3 model class)
  • safetensors
  • torch + CUDA-capable GPU with >= 4 GB VRAM (single-layer fits comfortably)
  • ~1 GB of read access to model-00001-of-000064.safetensors
# After cloning Kimi K2.6 to /path/to/Kimi-K2.6/
git clone https://huggingface.co/GenomaLabs-com/kv-cache-eviction-mla
cd kv-cache-eviction-mla
# Edit KIMI_PATH in scripts/kimi_layer_eviction_demo.py to point at your checkpoint
python scripts/kimi_layer_eviction_demo.py

Output: results/kimi_layer_eviction_demo.csv (256 rows: token_idx, attention_score, kept, category).

Per-token CSV available

results/kimi_layer_eviction_demo.csv contains per-token data for downstream analysis:

Column Description
token_idx Position in the sequence (0..255)
attention_score Cumulative attention mass received from all queries / heads
kept True if the H2O policy retains this token
category sink / recent / heavy / evicted

This CSV can be plotted (score vs index, colored by category) to visualize the eviction decision against the actual attention distribution.