Instructions to use GenomaLabs-com/kv-cache-eviction-mla with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GenomaLabs-com/kv-cache-eviction-mla with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("GenomaLabs-com/kv-cache-eviction-mla", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download notebooks/01_smoke_test_walkthrough.md from GenomaLabs-com/kv-cache-eviction-mla: direct link, hf CLI and curl.
- Browser
- Download file 5.19 kB
-
https://huggingface.co/GenomaLabs-com/kv-cache-eviction-mla/resolve/main/notebooks/01_smoke_test_walkthrough.md
- Command line
-
hf download hf://GenomaLabs-com/kv-cache-eviction-mla/notebooks/01_smoke_test_walkthrough.md
-
curl -L -o 01_smoke_test_walkthrough.md https://huggingface.co/GenomaLabs-com/kv-cache-eviction-mla/resolve/main/notebooks/01_smoke_test_walkthrough.md
Smoke Test Walkthrough
This walkthrough explains the smoke test built into src/kv_eviction_mla.py. It runs without a GPU and without downloading any real model — it constructs a minimal fake attention cache and exercises the eviction code path so you can verify the logic before patching a real model.
What the smoke test does
state = _EvictionState(budget=8, n_sink=2, n_recent=2, evict_every=1)
state.score = torch.rand(1, 20) # 20 tokens currently in the cache
cache = _FakeCache(1, 4, 20, 192, 128, "cpu")
_maybe_evict(cache, 0, state)
kept = cache.key_cache[0].shape[2]
assert kept == 2 + 8 + 2, f"Expected 12 kept, got {kept}"
print(f"OK - evicted 20 -> {kept} tokens (2 sinks + 8 heavy-hitters + 2 recent)")
What's happening, step by step:
- Set up an eviction state with a tiny budget.
budget=8means we keep the top-8 heavy-hitters;n_sink=2keeps the first 2 tokens;n_recent=2keeps the last 2 tokens. - Fake the per-token attention scores. In a real model, these accumulate as the eviction-aware forward runs across generation steps. Here we just put random numbers in, simulating a model that has run for 20 steps.
- Construct a fake KV cache with the canonical DeepseekV3 dimensions (1 batch, 4 heads, 20 tokens already cached, qk_dim=192, v_dim=128). Note this is a stripped-down
_FakeCacheclass for testing — the real code path operates ontransformers.cache_utils.DynamicCacheor compatible. - Call
_maybe_evict, which is the internal function that the patched forward calls after every step. - Assert the result. With a 20-token cache and a budget of
2 + 8 + 2 = 12, exactly 8 tokens should have been evicted, leaving 12.
If you run python src/kv_eviction_mla.py directly, the smoke test runs and prints:
Smoke test: checking patch logic on mock attention layer...
OK - evicted 20 -> 12 tokens (2 sinks + 8 heavy-hitters + 2 recent)
Usage example:
from kv_eviction_mla import install_kv_eviction, reset_eviction_scores
install_kv_eviction(model, budget=4096, n_sink=4, n_recent=512)
# generate ...
Why this matters before you patch a real model
The eviction logic is small (~30 lines of actual eviction; the rest is hook plumbing) but the failure modes are nasty if it goes wrong:
- If you evict the sinks, the model collapses immediately to producing repetitive garbage. The first 4 tokens carry disproportionate attention mass and the model gets dramatically uncalibrated without them.
- If you keep the wrong heavy hitters (e.g. by accumulating attention scores per-head instead of cross-head, with no normalization), you may keep tokens that one head cares about but no other head ever queries. The cache shrinks but quality degrades. See the discussion in
practice/03_kv_cache_techniques.mdof the GENOMA LABS handbook for details. - If you evict the recent window, generation stalls or produces topical drift, since the model loses its local context.
The smoke test confirms all three boundary conditions are respected for at least one set of inputs. It does not catch every bug; before deploying on a production workload, you should run a perplexity sweep at the deploy budget on representative prompts (this requires a real model and GPU; not in this notebook).
Memory model
The KV cache size for the canonical DeepseekV3 layout (61 layers, 64 heads per layer, qk_dim=192, v_dim=128, FP16) is:
cache_bytes = layers * seq_len * heads * (qk_dim + v_dim) * 2
= 61 * seq_len * 64 * (192 + 128) * 2
= 2,498,560 * seq_len bytes per token
~= 40,960 KB per layer per 1K tokens
Plugging in:
| seq_len | full cache | budget=4096 cache | savings |
|---|---|---|---|
| 8K | ~20 GB | ~10.2 GB | 2x |
| 16K | ~41 GB | ~10.2 GB | 4x |
| 32K | ~82 GB | ~10.2 GB | 8x |
| 64K | ~164 GB | ~10.2 GB | 16x |
| 128K | ~328 GB | ~10.2 GB | 32x |
| 1M | ~2.5 TB | ~10.2 GB | 250x |
The fixed-size budget is the whole point: cache memory does not grow with input length once the budget is hit. Eviction does add a small per-step overhead (the argpartition to find the bottom-k heavy hitters and the masked tensor copy), but in practice this is dominated by the model's per-step compute and is not a noticeable bottleneck.
What to validate next on real hardware
Before deploying, do at minimum:
- Perplexity sweep at full cache vs your deploy budget on a sample of representative prompts. Plot perplexity as a function of budget; look for the elbow.
- Task-level evaluation at the deploy budget. Perplexity is a coarse proxy.
- Worst-case benchmark — run RULER NIAH at the deploy budget across a range of context lengths. NIAH is the canonical worst case for cache eviction; if your workload has any retrieval-of-specific-facts pattern, the NIAH score predicts your worst-case behavior under eviction.
A reproducible RULER 128K benchmark on a public DeepSeek model with this eviction code will be added to this repository in a follow-up release. Until then, the smoke test verifies the logic; production validation is on you.