# Smoke Test Walkthrough This walkthrough explains the smoke test built into `src/kv_eviction_mla.py`. It runs without a GPU and without downloading any real model — it constructs a minimal fake attention cache and exercises the eviction code path so you can verify the logic before patching a real model. ## What the smoke test does ```python state = _EvictionState(budget=8, n_sink=2, n_recent=2, evict_every=1) state.score = torch.rand(1, 20) # 20 tokens currently in the cache cache = _FakeCache(1, 4, 20, 192, 128, "cpu") _maybe_evict(cache, 0, state) kept = cache.key_cache[0].shape[2] assert kept == 2 + 8 + 2, f"Expected 12 kept, got {kept}" print(f"OK - evicted 20 -> {kept} tokens (2 sinks + 8 heavy-hitters + 2 recent)") ``` What's happening, step by step: 1. **Set up an eviction state** with a tiny budget. `budget=8` means we keep the top-8 heavy-hitters; `n_sink=2` keeps the first 2 tokens; `n_recent=2` keeps the last 2 tokens. 2. **Fake the per-token attention scores.** In a real model, these accumulate as the eviction-aware forward runs across generation steps. Here we just put random numbers in, simulating a model that has run for 20 steps. 3. **Construct a fake KV cache** with the canonical DeepseekV3 dimensions (1 batch, 4 heads, 20 tokens already cached, qk_dim=192, v_dim=128). Note this is a stripped-down `_FakeCache` class for testing — the real code path operates on `transformers.cache_utils.DynamicCache` or compatible. 4. **Call `_maybe_evict`**, which is the internal function that the patched forward calls after every step. 5. **Assert the result.** With a 20-token cache and a budget of `2 + 8 + 2 = 12`, exactly 8 tokens should have been evicted, leaving 12. If you run `python src/kv_eviction_mla.py` directly, the smoke test runs and prints: ``` Smoke test: checking patch logic on mock attention layer... OK - evicted 20 -> 12 tokens (2 sinks + 8 heavy-hitters + 2 recent) Usage example: from kv_eviction_mla import install_kv_eviction, reset_eviction_scores install_kv_eviction(model, budget=4096, n_sink=4, n_recent=512) # generate ... ``` ## Why this matters before you patch a real model The eviction logic is small (~30 lines of actual eviction; the rest is hook plumbing) but the failure modes are nasty if it goes wrong: - **If you evict the sinks**, the model collapses immediately to producing repetitive garbage. The first 4 tokens carry disproportionate attention mass and the model gets dramatically uncalibrated without them. - **If you keep the wrong heavy hitters** (e.g. by accumulating attention scores per-head instead of cross-head, with no normalization), you may keep tokens that one head cares about but no other head ever queries. The cache shrinks but quality degrades. See the discussion in `practice/03_kv_cache_techniques.md` of the [GENOMA LABS handbook](https://github.com/genoma-labs/subquadratic-research) for details. - **If you evict the recent window**, generation stalls or produces topical drift, since the model loses its local context. The smoke test confirms all three boundary conditions are respected for at least one set of inputs. It does not catch every bug; before deploying on a production workload, you should run a perplexity sweep at the deploy budget on representative prompts (this requires a real model and GPU; not in this notebook). ## Memory model The KV cache size for the canonical DeepseekV3 layout (61 layers, 64 heads per layer, qk_dim=192, v_dim=128, FP16) is: ``` cache_bytes = layers * seq_len * heads * (qk_dim + v_dim) * 2 = 61 * seq_len * 64 * (192 + 128) * 2 = 2,498,560 * seq_len bytes per token ~= 40,960 KB per layer per 1K tokens ``` Plugging in: | seq_len | full cache | budget=4096 cache | savings | |---:|---:|---:|---:| | 8K | ~20 GB | ~10.2 GB | 2x | | 16K | ~41 GB | ~10.2 GB | 4x | | 32K | ~82 GB | ~10.2 GB | 8x | | 64K | ~164 GB | ~10.2 GB | 16x | | 128K | ~328 GB | ~10.2 GB | 32x | | 1M | ~2.5 TB | ~10.2 GB | 250x | The fixed-size budget is the whole point: cache memory does not grow with input length once the budget is hit. Eviction does add a small per-step overhead (the `argpartition` to find the bottom-k heavy hitters and the masked tensor copy), but in practice this is dominated by the model's per-step compute and is not a noticeable bottleneck. ## What to validate next on real hardware Before deploying, do at minimum: 1. **Perplexity sweep** at full cache vs your deploy budget on a sample of representative prompts. Plot perplexity as a function of budget; look for the elbow. 2. **Task-level evaluation** at the deploy budget. Perplexity is a coarse proxy. 3. **Worst-case benchmark** — run RULER NIAH at the deploy budget across a range of context lengths. NIAH is the canonical worst case for cache eviction; if your workload has any retrieval-of-specific-facts pattern, the NIAH score predicts your worst-case behavior under eviction. A reproducible RULER 128K benchmark on a public DeepSeek model with this eviction code will be added to this repository in a follow-up release. Until then, the smoke test verifies the logic; production validation is on you.