Instructions to use GenomaLabs-com/kv-cache-eviction-mla with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GenomaLabs-com/kv-cache-eviction-mla with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("GenomaLabs-com/kv-cache-eviction-mla", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download notebooks/03_kimi_real_weights_demo.md from GenomaLabs-com/kv-cache-eviction-mla: direct link, hf CLI and curl.
- Browser
- Download file 5.71 kB
-
https://huggingface.co/GenomaLabs-com/kv-cache-eviction-mla/resolve/main/notebooks/03_kimi_real_weights_demo.md
- Command line
-
hf download hf://GenomaLabs-com/kv-cache-eviction-mla/notebooks/03_kimi_real_weights_demo.md
-
curl -L -o 03_kimi_real_weights_demo.md https://huggingface.co/GenomaLabs-com/kv-cache-eviction-mla/resolve/main/notebooks/03_kimi_real_weights_demo.md
Real-Weights Demo: Kimi K2.6 Layer 0 Attention Distribution
This walkthrough runs the eviction policy on real Kimi K2.6 attention distributions, not the synthetic distributions used in the multi-step validation in 02_validation_results.md. It loads the actual layer-0 weights from the published Kimi K2.6 checkpoint, runs a single full-prefix forward over 256 synthetic input tokens, captures the real attention score distribution, and applies the H2O heavy-hitter eviction policy to it.
The point: validate the eviction policy makes sensible decisions when fed actual Kimi attention scores (vs. random distributions).
What the demo does
- Loads the canonical Kimi K2.6 architecture config (61L, 64H, MLA with kv_lora_rank=512, qk_dim=192, v_dim=128).
- Instantiates a single
transformers.models.deepseek_v3.modeling_deepseek_v3.DeepseekV3Attentionmodule (layer index 0). - Loads the actual layer-0 attention weights from
model-00001-of-000064.safetensorsof the published Kimi K2.6 checkpoint. All 7 weights match cleanly:q_a_proj,q_b_proj,q_a_layernorm,kv_a_proj_with_mqa,kv_b_proj,kv_a_layernorm,o_proj. - Runs one forward pass with
seq_len=256synthetic input embeddings, captures theattention_weightstensor of shape[1, 64 heads, 256, 256]. - Computes per-kv-token cumulative attention mass (sum across heads and across all queries that attend to that kv position).
- Applies the H2O policy: keep
n_sink=4start tokens,n_recent=32end tokens, plus the topbudget=64heavy hitters from the middle. Mark all others as evicted. - Reports the score distribution and the kept/evicted ratio.
Hardware
NVIDIA TITAN RTX (24 GB), BF16 compute. The single-layer forward over 256 tokens completes in 0.07 seconds.
Configuration
config: 61L hidden=7168 heads=64
qk_dim=192 v_dim=128 kv_lora_rank=512
layer params: 101,124,096
seq_len: 256
budget: 64 (heavy-hitter slots in the middle)
n_sink: 4 (always kept, indices 0..3)
n_recent: 32 (always kept, last 32)
Results
attn_out shape: torch.Size([1, 256, 7168])
attn_weights shape: torch.Size([1, 64, 256, 256])
score per token: shape=(256,)
score range: [0.336, 350.870]
score mean: 64.000
score std: 64.217
H2O eviction policy applied:
kept 100 of 256 tokens (39.1%)
sinks 4 (indices 0..3)
heavy 64 of 220 middle tokens chosen
recent 32 (last 32)
evicted 156 (60.9%)
top 10 heavy-hitter scores:
319.18 277.95 257.06 238.43 237.26
222.64 210.78 208.90 207.39 189.85
mean score of heavy-hitters kept: 141.082
mean score of evicted tokens: 38.583
heavy / evicted score ratio: 3.66x
Interpretation
The attention distribution from real Kimi K2.6 layer 0 is strongly heavy-tailed:
- Score spread: ~1000x between min (0.336) and max (350.87).
- Top heavy-hitter scores are roughly 10x the mean (319 vs. mean of 64).
- The standard deviation (64) is comparable to the mean — wide variance, lots of structure for the eviction policy to exploit.
The H2O policy correctly identifies the high-attention tokens: tokens it keeps as heavy-hitters score on average 3.66 times higher than tokens it evicts. This is a meaningful gap — the policy is not just shuffling random selections.
This validates that on a real frontier-scale MLA attention layer, the H2O recipe (top-k by accumulated attention mass + sinks + recent window) makes sensible decisions about which tokens to keep when cache pressure forces eviction.
What this does NOT prove
- Output quality: we use random input embeddings, so the attention layer's output is not meaningful text. We're only validating the attention distribution and the eviction decision, not generation quality.
- Multi-layer consistency: layer 0's attention pattern may differ from layers 30 or 60. A full-model evaluation would average across all 61 layers; we only loaded 1.
- End-to-end with cache eviction: we apply the policy to a fully-populated 256-token cache after the forward; we do not exercise the cache-management plumbing (the
install_kv_evictionpatch on transformers 5.x DynamicCache is still on the roadmap — see README).
Reproducing
The Kimi K2.6 weights are published by Moonshot AI on HuggingFace: moonshotai/Kimi-K2.6. The script in scripts/kimi_layer_eviction_demo.py runs as-is on a host with:
- transformers >= 5.0 (DeepseekV3 model class)
- safetensors
- torch + CUDA-capable GPU with >= 4 GB VRAM (single-layer fits comfortably)
- ~1 GB of read access to
model-00001-of-000064.safetensors
# After cloning Kimi K2.6 to /path/to/Kimi-K2.6/
git clone https://huggingface.co/GenomaLabs-com/kv-cache-eviction-mla
cd kv-cache-eviction-mla
# Edit KIMI_PATH in scripts/kimi_layer_eviction_demo.py to point at your checkpoint
python scripts/kimi_layer_eviction_demo.py
Output: results/kimi_layer_eviction_demo.csv (256 rows: token_idx, attention_score, kept, category).
Per-token CSV available
results/kimi_layer_eviction_demo.csv contains per-token data for downstream analysis:
| Column | Description |
|---|---|
token_idx |
Position in the sequence (0..255) |
attention_score |
Cumulative attention mass received from all queries / heads |
kept |
True if the H2O policy retains this token |
category |
sink / recent / heavy / evicted |
This CSV can be plotted (score vs index, colored by category) to visualize the eviction decision against the actual attention distribution.