# Rolling-KV — overflow throughput evaluations (consolidated) **What this is.** The measured performance record for rolling-KV / POSITION_WINDOW (streaming the host KV tail over PCIe when the cache overflows VRAM). The numbers were historically scattered across three feature docs and one study; this file is the single evals-side index. **Design + code** live in the feature docs; this file is **results only**, with back-links. **Sources consolidated here:** - `docs/features/rolling_kv.md` — M7 design + M7-G gate + the bug-1843 fix + the #588 3090 cliff curve. - `docs/features/rolling_kv_step_prefetch.md` — the 3090 overflow-decode hyperbola (both model classes) + the step-prefetch design. - `docs/features/rolling_kv_compute_reserve.md` — the bs2 boot-OOM fix (bug-1342) that *unblocked* the RTX 6000 curve, + the prefill cliff. - `docs/evaluations/hal_kv_compression.md` §3a — the bs2 host-stream-exact spill cliff (route-(a), the reason compression wins over spill). --- ## 0. The one-line verdict Rolling-KV's re-stream design is **bandwidth-bound by physics**: every decoded token restreams the spilled tail over PCIe, so decode tps follows `t(token) ≈ resident_compute_ms + spilled_bytes / PCIe_BW`. After the bug-1843 compute-side fix, the curve sits within ~5 % of that model. **No copy-plumbing lever can beat the link itself** → the rolling-KV *perf* line is **CLOSED (NO-GO)** for deep overflow; window mode ships as the **graceful-overflow fallback** (linear, correctness-clean, no cliff pathology). The real capacity lever is **residency-side compression**, not spill — see [`hal_kv_compression.md`](hal_kv_compression.md) §4 and [`context.md`](context.md) §5b. --- ## 1. The mechanism (why it is bandwidth-bound) POSITION_WINDOW keeps a resident window of recent cells and streams the older tail from host RAM every decode step. At `n_q=1` an attention op's compute is a sub-ms VRAM read, but its tail-tile DMA is ~30× that, so the op-level 1-tile-ahead ping-pong is DMA-bound from MiB 1. Full-attention models stream a tail on **every** layer; iSWA models stream only their few global layers → an order-of-magnitude gap. Mechanism + the `cs_sync` cross-layer barrier: `rolling_kv_step_prefetch.md` §"Root cause". --- ## 2. 3090 overflow-decode cliff — the original hyperbola (2026-07-01) Op-level streaming, pre-bug-1843. PCIe ~6.4 GB/s, `+0.156 ms/MiB`. | model | attn | streaming layers | overflow decode | shape | |---|---|---|---|---| | Gemma-4-A4B-128e | iSWA | ~5 global | 71 → 19.5 tps @236 MiB tail | graceful hyperbola | | Qwen2.5-14B-1M | full | 48 | 26.9 → 0.86 tps @~1.2 GB tail | ~10× worse (48 streaming layers) | Qwen engages POSITION_WINDOW (GPU 60–84 % busy) — it is slow purely because full-attention restreams a tail on every layer. Source: `rolling_kv_step_prefetch.md` §"Problem (measured, 3090)". (#583/#584) --- ## 3. The compute-side fix — bug-1843 / patch 0098 (2026-07-05) The S0-cliff redirects (#586 staging ring / #587 barrier-drop) were **perf-inert** on the 3090 (bug-1838: the PCIe link is already saturated — no copy-side slack). The real waste was a per-tile `streaming_lse_kernel` **recompute** that every `head_dim ≤ 256` decode tile paid, because the Stage-3a LSE channel was gated on `head_dim > 256`. Patch `0098-rolling-kv-lse-decode` arms the LSE channel for D ≤ 256 at all three streaming-FA sites. **Gate (3090, DSO `df46dc36`, Qwen3-8B-Q8_0 ctx 40960 / vt 14000):** - spill decode **1.14 → 6.04 tps (5.3×, == NOOP ceiling)**, needle PRESENT - resident 42.5 tps unchanged; `OPENCOTI_LSE_NOOP` now speed-inert (recompute structurally off the path) - correctness: teacher-forced logit-equiv (ours-resident REF vs POSITION_WINDOW at 256/16384 cells resident) **real_frac 0.0, frac_full_agree 1.0, mean_tv 0.0068 → PASS** Source: `rolling_kv.md` §"2026-07-05 — bug-1843 shipped". (#623) --- ## 4. #588 perf gate — post-fix 3090 cliff curve → CLOSED NO-GO (2026-07-06) bug-1843-fixed DSO (`df46dc36`), 3090, Qwen3-8B-Q8_0, ctx 16384 (~13.1k-token prompt = 2304 MiB f16 KV), needle PRESENT in every cell. | vt (MiB) | window split | occupied tail | decode tps | ms/token | |---|---|---|---|---| | 22000/16000/14000/13000 | FULLY RESIDENT | 0 | 42.3–42.5 | 23.6 | | 12000 | 11008 / 16384 cells | ~2.1k cells ≈ 295 MiB | 15.5 | 64.6 | | 8000/4000/2000 | 256 / 16384 (floor) | ~12.9k cells ≈ 1815 MiB | 3.33 | 300 | The PCIe model fits within ~5 %: `t ≈ 23.6 ms + occupied_tail_bytes / 6.5 GB/s` (vt12000 predicted 67 ms vs 64.6 measured; floor predicted 303 vs 300). Spill decode is now purely bandwidth-bound — the structural floor for a re-stream design. **Verdict: NO-GO for further re-stream perf work.** The "within 10–15 % of resident" bar is only reachable for tails ≤ ~25 MiB (~170 cells) on a 6.5 GB/s 3090. No #586-class copy-plumbing beats the link. Source: `rolling_kv.md` §"2026-07-06 — #588 perf gate". --- ## 5. bs2 RTX 6000 — the serving-relevant host opencoti serves on the RTX 6000 (PCIe ~50 GB/s, ~8× the 3090's link — see memory `reference_pcie_bandwidth_hosts`). Two bs2 measurements exist plus one open cell. ### 5a. Host-stream-exact spill cliff (route-(a), simulated 24 GB card, 2026-07-01) From `hal_kv_compression.md` §3a — pre-bug-1843, f16 KV, 8 GiB budget: | cell | resident / host tail | decode tps | needle | |---|---|---|---| | f16 @32k | all / 0 | **67.3** | ✓ | | f16 @49k | 43k cells / 1152 MiB | 57.6 | ✓ | | f16 @64k | 43k / **4224 MiB** | **0.99** | ✓ | | f16 @98k | 43k / 10368 MiB | 0.71 | ✓ | | q4_0 @64k (resident) | all / 0 | 27.1 | ✓ | | q4_0 @128k (resident) | all / 0 | 16.2 | ✓ | The cliff is catastrophic past ~2 GiB tail (0.99 tps @64k). Same-context resident **compression is 27× faster** — this is *why* route-(a) spill is demoted to beyond-max-compression fallback and compression is the primary mechanism. ### 5b. Prefill cliff after the boot-OOM fix (bug-1342, 2026-07-01) `rolling_kv_compute_reserve.md` §"Gate result": with the two-pass compute-reserve fix, deep-overflow configs boot on the RTX 6000; **prefill falls 4000 → ~450 tok/s** as the host tail grows (expected cliff). This fix is what makes the §5c curve *runnable* on bs2 at all. ### 5c. bs2 real-spill characterization — qw35a3, 256k fine tail ladder, DCA-on vs DCA-off (2026-07-10, #631/#633) **This closes the §5c gap with the first fine bs2 spill curve, and it OVERTURNS the initial coarse read.** Vehicle: the serving-size **Qwen3.6-35B-A3B** (`qw35a3`, hybrid DeltaNet+attention MoE, D256 attn, fast A3B-active decode), 256k ctx, f16 KV, fixed binary `llamafile.dualctx633`. A fine `--vram-target` ladder (200 MiB tail steps) run **both DCA-on and DCA-off** shows: **there IS a graceful, bandwidth-bound overflow band** (the earlier "hard cliff / no graceful band" claim is **RETRACTED**); the catastrophic collapse seen in the first coarse sweep was a **DCA-specific bug (bug-2144)**, not a hardware limit and not a hybrid-recurrent property — the coarse sweep's smallest tail (1710 MiB) simply landed *past* the DCA-on knee, so it only ever sampled the buggy regime. **Two gate fixes landed first.** (1) **bug-2142 / #633** — `--vram-target` / `--kv-residency-mode window` were silently ignored for **every hybrid-recurrent model** (`LLM_ARCH_QWEN35/QWEN35MOE`, Nemotron-H, Falcon-H1): the `llama_memory_hybrid` / `llama_memory_hybrid_iswa` wrappers (`llama-model.cpp` create_memory ~L2041/L2022) built their internal attention `llama_kv_cache` with hardcoded residency defaults, so the sizer always logged `FULLY RESIDENT` regardless of the flag. Fix (`opencoti-hook: rolling-KV hybrid residency (bug-2142)`, host-only, 6 edits/5 files): thread `vram_target_mib / pcie_bw_gbps / kv_residency_mode / kv_compute_reserve_mib / kv_window_measure_pass` through both hybrid ctors into the attention cache, mirroring the iSWA (L2100) / plain-dense (L2140) paths that already forwarded them. Verified-engaged + regression-clean (shield below). (2) Note **bug-1839**: `--vram-target` is a **TOTAL device cap** (weights+KV+compute), so a spill target must be `< weights + resident-KV + compute`; the old "24 GB card = `--vram-target 9600`" recipe (14B) now floors resident-KV to 0. Empirical tail↔target map for qw35a3 @256k: `tail_MiB ≈ 36708 − vram_target_MiB` (fits every cell to ±5 MiB). **qw35a3 KV is tiny** (hybrid — only the few full-attention layers grow): f16 KV = **5120 MiB @256k, 10240 MiB @512k** (~20 KiB/tok), an order of magnitude under the 14B-1M full-attention class (§5a). On the 96 GB RTX 6000 it never *needs* to spill; the ladder *forces* spill via `--vram-target` to characterize a smaller card / multi-session KV pressure. #### DCA-OFF — the correct rolling-KV behavior (bandwidth-bound, graceful) The non-DCA `build_attn` overload routes spilling layers through the bug-1840 streaming fast-path (`build_attn_mha_position_window`), which FAs the resident window on-device and streams only the **tail** once. Result: a smooth, bandwidth-bound decline — the right shape. | host tail | resident cells | prefill tps | **decode tps** | achieved GiB/s¹ | needle | |---|---|---|---|---|---| | 0 (resident) | 262144 / 262144 | 4799 | **123.5** | — | ✓ | | ~0 (VT 36308) | resident | 4797 | 123.0 | — | ✓ | | **80 MiB** | 258048 | 326 | **105.4** | ~15 (fixed engage cost) | ✓ | | **280 MiB** | 247808 | 326 | **103.6** | ~29 | ✓ | | **680 MiB** | 227328 | 326 | **104.4** | ~71 | ✓ | | **1080 MiB** | 206848 | 325 | **57.7** | **~62** | ✓ | | **1880 MiB** | 165888 | 324 | **31.7** | **~60** | ✓ | | **2680 MiB** | 124928 | 315 | **21.4** | **~57** | ✓ | ¹ `tail_MiB / (1000/decode_tps) ms`, tail-streamed-once framing. **Knee at ~680–1080 MiB:** decode holds **within ~15 % of resident to a 680 MiB tail** (104 vs 123.5), then declines cleanly, saturating at **~57–62 GiB/s** — i.e. **PCIe-bandwidth-bound near the RTX 6000 link**, exactly the correct rolling-KV shape (§4's `t ≈ compute + tail_bytes/BW`, now confirmed on the 50 GB/s bus). Needle PRESENT at every depth. The −15 % step at first spill (80 MiB → 105 tps) is the fixed streaming-path engagement cost. **⚠ DCA-off PREFILL is broken though (bug-2145) — a second, mirror-image gap.** The prefill column above **cliffs from 4797 → ~325 tok/s the instant any tail forms (80 MiB) and then stays FLAT ~325 all the way to a 2.7 GB tail** (325/325/325/324/315 — **tail-INDEPENDENT**). Tail-independence is the smoking gun: a bandwidth-bound prefill would slow *with* the tail; a flat 325 means the fast prefill path is simply **switched off whenever the window is active** (a routing gap, not a copy cost). So DCA-off is graceful for **decode** but not **prefill**. #### DCA-ON — perfect to ~800 MiB, then a sharp cliff (bug-2144 — NOW FIXED, see §5d) > **⚠ HISTORICAL / pre-fix.** The cliff below is the *diagnosis* that motivated the fix. bug-2144 > is **fixed** by the C2 streaming `dca_fused` kernel (§5d): DCA-on spill now tracks/beats the > DCA-off reference at matched tails, no cliff. The pre-fix ladder is retained because it is the > evidence base for the mechanism; read it as "what was broken," not current behaviour. | host tail | resident cells | prefill tps | **decode tps** | needle | |---|---|---|---|---|---| | 0 (resident) | 262144 | 4392 | **112.5** | ✓ | | **205 MiB** | 251648 | 4077 | **111.7** | ✓ | | **405 MiB** | 241408 | 4080 | **112.5** | ✓ | | **605 MiB** | 231168 | 4077 | **113.2** | ✓ | | **805 MiB** | 220928 | 4073 | **112.8** | ✓ | | **1205 MiB** | 200448 | 3668 | **2.93** ⚠ | ✓ | | **1605 MiB** | 179968 | 3011 | **3.09** ⚠ | ✓ | | **2405 MiB** | 139008 | 2330 | **3.19** ⚠ | ✓ | | **3205 MiB** | 98048 | 2010 | **3.33** ⚠ | ✓ | DCA-on is **penalty-free (full 112–113 tps) out to an ~805 MiB tail** — actually *better* than DCA-off in that band (DCA-off already pays −15 % at 80 MiB) — then **cliffs 38× in a single 400 MiB step (805 → 1205 MiB: 112.8 → 2.93 tps)** and stays dead-flat ~3 tps past it. Past-cliff decode is 341 ms/token: only **3.5 GiB/s** on the tail-once framing (≪ link → not bandwidth-bound), but **~26 GiB/s** on the ~9 GiB **full-KV-round-trip** framing (window-D2H 3920 MiB + full-concat-H2D 5120 MiB) — the fingerprint of bug-2144. **Matched-tail contrast (the empirical lock on the code diagnosis): at ~1.1 GiB tail, DCA-off decodes 57.7 tps vs DCA-on 2.93 tps — ~20× faster.** Same model, same binary, same tail; the *only* difference is DCA on/off. This confirms the JOB-1 code diagnosis below. #### Mechanism — bug-2144: DCA bypasses the streaming window path (confirmed) The bug-1840 perf fast-path `build_attn_mha_position_window` (`llama-graph.cpp:2503`) is reached **only from the non-DCA `build_attn` overload** (`llama-graph.cpp:2999-3007`). **DCA's `build_attn_dca_core` (`dca.cpp:229`) fetches K/V with the raw whole-cache accessors `get_k`/`get_v` (`dca.cpp:251-252`)**, which for a spilling layer return the **cross-backend correctness fallback** `ggml_concat(win_v /*device*/, tail_v /*pinned host*/, 2)` (`llama-kv-cache.cpp:2585-2611`; its own comment names `build_attn_mha_position_window` as the perf path it is *not* taking). The scheduler materialises that concat on the compute backend → a **synchronous ~2× full-KV host↔device round-trip every decode step** (bug-1840's documented "12.6 GiB/step both ways → 0.99 tps at 4.2 GB tail", `llama-graph.cpp:2989-2990`); the async DMA tile overlap lives in `ggml_cuda_streaming_flash_attn` (`fattn.cu:1170`) which DCA never enters. This predicts the ~26 GiB/s effective full-KV rate, the flat-past-cliff plateau (round-trip volume ≈ window-D2H + full-KV-H2D, ~constant), and the slight rise 2.93 → 3.33 as the tail grows (smaller window ⇒ less window-D2H). **The sharp ~1 GiB step (805 → 1205 MiB) — CONFIRMED as the ~1-chunk knee.** The DCA chunk is 32768 cells ≈ **640 MiB**; the step sits between a **1.26-chunk tail** (805 MiB, fast) and a **1.88-chunk tail** (1205 MiB, slow) — the tail crossing into a **second DCA chunk** is the trigger, exactly as hypothesised. The C2 fix (§5d) reproduces this knee on the fixed binary as a graceful *step down* (37 → 15 tps at the same crossing) instead of a cliff, since the streaming kernel reads only the chunks it needs rather than round-tripping the full KV. Below ~1 chunk of tail DCA-on stays entirely device-resident and pays nothing (unchanged pre/post fix). #### Resident-compression anchors (no spill; DCA-on, everything on-device) | ctx | KV type | KV MiB | prefill tps | **decode tps** | needle | |---|---|---|---|---|---| | 256k | f16/f16 | 5120 | 4358 | **116.3** | ✓ | | 256k | q8_0K/q4_0V | ~2080 | 4241 | **75.2** | ✓ | | 256k | turbo2/turbo2 | **680** | 3664 | **46.4** | ✓ | | 512k | f16/f16 | 10240 | 2856 | **87.65** | ✓ | | 512k | q8_0K/q4_0V | ~4110 | 2741 | **48.8** | ✓ | | 512k | turbo2/turbo2 | **1360** | 2384 | **27.3** | ✓ | **Regression shield.** Non-hybrid 14B-1M-Q8_0 @65536, `--vram-target 9600`, fixed binary → `256 / 65536 cells resident, host tail = 12240 MiB` — **byte-identical to the pre-fix binary**, so the bug-2142 fix is isolated to the hybrid path. **Dense control (Qwen-27B-omni, full attention) — DEFERRED.** Only the resident calibration landed: `KV 16384 MiB @256k` (full-attention → ~10× the hybrid KV, as expected), DCA-off resident decode **36.9 tps**. The dense spill ladder did not run (GPU time). It is **not blocking** — the qw35a3 DCA-on-vs-off contrast already isolates bug-2144 as DCA-specific (not hybrid-specific) — but a clean single-GPU dense ladder is worth running later (after #632's tps work) to confirm the DCA-off graceful band and the DCA-on cliff reproduce on a dense arch. #### The full 4-quadrant picture — phase × DCA (two halves of one routing gap) Both anomalies are the **same** defect seen from two sides: the streaming-window fast-path is wired for only **two** of the four (phase × DCA) quadrants. Where it is wired, behaviour is graceful; where it is not, decode/prefill falls to a slow fallback. | | **DCA-OFF** | **DCA-ON** | |---|---|---| | **DECODE** | ✅ **graceful** — bandwidth-bound, 123→104@680MiB→21@2.7GB, ~57–62 GiB/s (streaming window path) | ✅ **graceful (bug-2144 FIXED, §5d)** — was 112→2.9 cliff; C2 streaming `dca_fused` now tracks/beats DCA-off at matched tails | | **PREFILL** | ❌ **cliff** — 4797 → **flat ~325**, tail-INDEPENDENT from 80 MiB to 2.7 GB (**bug-2145**: fast prefill path disabled under window mode) | ✅ **graceful** — 4076@605MiB → 3668@1.2GB → 2010@2.7GB (declines with tail) | The diagonal was the tell: **non-DCA-decode ✓ and DCA-prefill ✓** always took the fast path; **non-DCA-prefill (bug-2145)** and **DCA-decode (bug-2144)** were the two uncovered quadrants. **DCA-decode is now covered (bug-2144 FIXED, §5d)** — the C2 streaming-tail kernel routes DCA spill through the same async tile machinery. **bug-2145 remains open** (non-DCA prefill): its tail-**independence** (80 MiB costs the same 325 tok/s as 2.7 GB) proves it is a **path switch** — the fast prefill kernel is turned off the moment `window_cells < kv_size`, regardless of how little actually spilled — not a bandwidth cost. It is an additive routing/plumbing gap around the *same* `build_attn_mha_position_window` / `ggml_streaming_flash_attn` machinery, not a model or hardware limit. **Neither quadrant ever lost the needle** — correctness is clean throughout; these are purely throughput-path gaps. #### Verdict (revised) - **A graceful overflow band exists on bs2 — for DECODE.** With DCA off, qw35a3 spill **decode** is bandwidth-bound (~57–62 GiB/s, needle clean) and holds within ~15 % of resident to a ~680 MiB tail — the correct rolling-KV behavior on the 50 GB/s link. The §0/§4 "deep overflow is bandwidth-bound" verdict stands (for decode); the "no graceful band for hybrid" over-claim is withdrawn. **DCA-on decode is now graceful too (bug-2144 FIXED, §5d).** The one remaining non-graceful quadrant is **DCA-off *prefill* (bug-2145)** — a DCA-off spill config still cliffs on prefill; closing it needs the same streaming-window routing applied to the non-DCA prefill path. - **DCA-on had a ~1 GiB decode cliff (bug-2144) — FIXED (§5d, C2 streaming `dca_fused`).** Was perfect below ~800 MiB tail then a full-KV concat round-trip (~20–55× slower). The C2 streaming-tail kernel now makes DCA-on spill track/beat the DCA-off reference at matched tails (18× recovery at a 6.4 GiB tail). Gated `WS2_DCA_STREAM=1` pending ship green-light. - **Compression remains the primary capacity lever** (§0, `hal_kv_compression.md` §3a/§4): turbo2 keeps qw35a3 resident at 0.68/1.36 GiB KV (46/27 tps), and qw35a3's KV is small enough to stay resident on a 96 GB card regardless. Spill (DCA-off) is a valid graceful-overflow *fallback* for smaller cards / multi-session pressure; the bug-2144 fix (§5d) now extends that graceful fallback to DCA-on configs too. **Contention caveat.** These ran on GPU1 with #632's tps work on GPU0 concurrently. GPU1 is a separate device (isolated VRAM/compute), and the load-bearing conclusions — the ~20× DCA-off vs DCA-on gap, the DCA-off bandwidth saturation, the DCA-on cliff position — are far outside any cross-GPU noise and internally consistent (monotone ladders). Fine-grained tps (±1–2 tps between adjacent cells) may carry minor host/PCIe contention; correctness (needle) and the resident/tail splits are contention-immune. Method: fixed binary, DSO `2377635512` (bug-2141 guard, inert for these non-degenerate rows), no rebuild during the run. ### 5d. bug-2144 C2 fix — DCA-on spill made graceful (streaming `dca_fused`), Qwen2.5-14B-1M @128k (2026-07-10) **The fix.** bug-2144's DCA-on decode cliff (§5c: full-KV concat round-trip, ~55× collapse past ~1 chunk of tail) is eliminated by giving the fused DCA kernel a two-region window+tail streaming path — the same async tile ping-pong the non-DCA `build_attn_mha_position_window` uses. Host plumbing threads `k_tail`/`v_tail`/`n_win` into `ggml_flash_attn_ext_dca_fused` (tail on `src[7]`/`src[8]`, `op_params[5]=n_win`); gated behind `WS2_DCA_STREAM=1`, decode-only via `WS2_NQ_MAX` (n_q ≤ 16). Legacy concat path is byte-identical when off. Marker `opencoti-hook: DCA streaming-tail (bug-2144-C / Stage C)` (dca.cpp). **Decisive ladder — DCA-on (C2 stream) vs DCA-off (rolling-KV reference).** Both modes re-measured on the **reserve-fixed C2r binary** (bug-2146 below), so these are the clean, accurate numbers — no reserve artifact, no run contamination. Qwen2.5-14B-1M Q8_0, f16 KV, `-c 131072`, needle QUASAR-4417, bs2 RTX 6000 single-GPU, ~51 GB/s eff link. Both **prefill / decode** tps recorded; needle PRESENT in **every** cell of both ladders: | vram-target | DCA-on tail | **DCA-on** prefill / decode | DCA-off tail | **DCA-off** prefill / decode | |---|---|---|---|---| | 41888 | 0 (resident) | 2094 / **38.50** | 0 (resident) | 1547 / **40.13** | | 41000 | ~864 MiB | 2016 / 37.27 | ~864 MiB | 1342 / 35.82 | | 40200 | ~1680 | 2015 / 37.35 | ~1680 | 1335 / 35.92 | | 39400 | ~2496 | 2016 / 37.06 | ~2496 | 1338 / 35.91 | | 38600 | ~3264 | 2016 / 37.23 | ~3264 | 1340 / 35.92 | | 37800 | ~4080 | 2015 / 37.23 | ~4080 | 1344 / 35.56 | | 37000 | ~4896 | 2017 / 37.19 | ~4896 | 1335 / 34.75 | | 36200 | ~5808 | 1598 / **14.45** (knee) | ~5664 | 1172 / **10.78** (knee) | | 35400 | ~6720 | 1361 / 11.69 | ~6480 | 1028 / 9.29 | | 34600 | ~7680 | 1164 / 9.72 | ~7296 | 933 / 8.17 | | 33000 | ~9600 | 925 / 7.29 | ~8880 | 782 / 6.55 | | 31000 | ~12000 | 765 / 5.56 | ~10896 | 670 / 5.22 | **Verdict: the C2 streaming path makes DCA-on spill track *and beat* the DCA-off "normal" rolling-KV reference at every point** — both hold their full-speed plateau out to a **~4.9 GiB tail** (DCA-on ~37.2, DCA-off ~35.8), knee together at **~5.6 GiB** (tail crossing ~1 DCA chunk = 32768 cells), then the same graceful bandwidth-bound rolloff, **no cliff.** DCA-on decode is faster than DCA-off at every cell (plateau *and* deep tail) because the fused kernel reads only the chunks it needs, whereas DCA-off window mode streams the whole tail; DCA-on prefill is also higher on the plateau (~2015 vs ~1340). Pre-fix DCA-on cliffed to ~0.68 tps past ~1 GiB of tail → now 11.69 at a 6.7 GiB tail = **~17× recovery.** bug-2144 CLOSED on perf. (vt=41888 DCA-on decode 38.50 confirms full residency under the bug-2146 fix; the harness's tail-grep prints a stale iter-0 value there because the final "FULLY RESIDENT" boot line carries no "host tail =" token — decode is ground truth.) > **Caveat for the pending bug-2145 fix (non-DCA prefill cliff, #636):** that fix, if it lands, can > only move the **DCA-off prefill column** (when the window engages) — decode is untouched in both > modes and the DCA-on ladder is entirely unaffected. Worst case is a cheap one-column prefill > re-run of the dcaoff mode, never the whole matrix. **Reserve artifact — root cause of the pre-fix forced ≥3 GiB tail (now fixed; the ladder above is already the fixed run).** The bug-1342 two-pass compute-reserve sizer measures pass 1 with a **degenerate `wc=256` window** (llama-kv-cache.cpp `measure_pass`), which for the streaming graph means a maximal ~130816-cell host tail. Instrumented boot (`-lv 5`, vt=41888): ``` Pass 1 (measure, wc=256, tail=130816): CUDA0 compute buffer = 4088.81 MiB → reserve 4344 MiB Pass 2 (real, wc=114176, tail=16896): CUDA0 compute buffer = 926.01 MiB ``` The streaming decode graph's **real** device compute buffer is only **926 MiB**; the reserve holds back **4344 MiB → a 3.16 GiB over-reservation**, a pure measurement artifact of the wc=256 pass (confirmed vs a `WS2_NQ_MAX=0` control: 18216 vs 15054 MiB budget at identical free VRAM). This is why vt=41888 was forced to a 3168 MiB tail instead of staying fully resident. **FIX — bug-2146, IMPLEMENTED + validated (host-only, no DSO rebuild).** `llama_context` ctor (llama-context.cpp) now **iterates the two-pass reserve to a fixpoint** instead of trusting the single degenerate wc=256 measurement: each re-size grows the window → shrinks the tail → shrinks the measured buffer, so the reserve decreases monotonically and converges (to the plain resident buffer, tail 0, when the KV fits). The resident graph's buffer is a subset of the streaming graph's, so a re-size only ever reveals a smaller-or-equal buffer — it can never under-reserve. Capped at 3 iterations; stops as soon as the reserve stops shrinking (128 MiB margin). For the non-streaming case pass 2 measures what pass 1 did → converges in ONE step, **byte-identical to prior behaviour**. Validated on the same C2r binary (Qwen2.5-14B-1M Q8_0 @128k, bs2): | config | pre-fix | **post-fix (bug-2146)** | |---|---|---| | vt=41888 (KV fits) | 3168 MiB tail, 37.10 tps | **FULLY RESIDENT, tail 0, 38.40 tps** (converged iter 0→1→2) | | vt=35400 (genuine spill) | 9648 MiB tail, 7.24 tps | **6720 MiB tail, 11.72 tps (+62%)** (reserve 14966→17858 budget) | Needle PRESENT both cells. This recovers the ~3.16 GiB and, because the corrected (smaller) reserve grows the resident window at every vram-target, **every §5d/§5e streaming ladder cell above was measured with the over-large reserve and improves under the fix** (the tables are the conservative lower bound). Still-open cleanup: strip leftover `get_k_tail: ENTRY` / `ACTUAL tail` `LLAMA_LOG_WARN`s (llama-kv-cache.cpp:3252/3389/3408) before ship. ### 5e. Rolling-KV spill vs upstream CPU-layer offload — bytes vs compute (Qwen2.5-14B-1M @128k, 2026-07-10) The honest baseline for "what does everyone else do when VRAM runs out" is **not** rolling-KV — it is upstream llama.cpp's `-ngl < n_layer`, which pushes whole transformer layers (weights **+** their KV **+** their compute) to CPU. This section pits that against rolling-KV spill on the same model/host, to show the two overflow mechanisms are categorically different. **Two ways to move data off a too-small card:** - **Rolling-KV spill** — all compute stays on GPU; only KV *bytes* stream over PCIe (window on device, tail on host). Cost is **PCIe-bandwidth-bound**. - **`-ngl` layer offload** — whole layers' *compute* moves to CPU. Cost is **CPU-matmul-bound** (and the CPU also owns those layers' KV). This is what stock llama.cpp does when the model+KV won't fit. **`-ngl` offload ladder** (DCA off, full-KV-in-VRAM for GPU-resident layers, `-c 131072`, needle QUASAR-4417, bs2 RTX 6000; RAM column = weights 279 MiB + KV 512 MiB per CPU layer, + ~790 MiB output tensor which drops to CPU at any `-ngl ≤ 48`, + ~790 MiB token-embedding always on CPU): | `-ngl` | layers on CPU | **≈ RAM offloaded** | prefill tps | **decode tps** | needle | |---|---|---|---|---|---| | 99 | 0 (output on GPU) | ~0.8 GiB (embd) | — | **39.81** (all-VRAM baseline) | ✓ | | 48 | 0 (output→CPU) | ~1.6 GiB | 1439 | **13.67** | ✓ | | 46 | 2 | ~3.1 GiB | 1213 | **7.02** | ✓ | | 44 | 4 | ~4.6 GiB | 1097 | **4.51** | ✓ | | 42 | 6 | ~6.2 GiB | 1008 | **3.41** | ✓ | | 40 | 8 | ~7.7 GiB | 917 | **2.67** | ✓ | | 38 | 10 | ~9.3 GiB | 847 | **2.19** | ✓ | | 36 | 12 | ~10.8 GiB | 795 | **1.86** | ✓ | | 34 | 14 | ~12.4 GiB | 745 | **1.63** | ✓ | | 32 | 16 | ~13.9 GiB | 721 | **1.44** | ✓ | | 31 | 17 | ~14.7 GiB | 678 | **1.38** | ✓ | **Gotcha (load-bearing for reading this table):** `-ngl 48` is **not** all-in-VRAM. llamafile keeps the **output/LM-head projection** (152k-vocab × 5120) on CPU until `-ngl > n_layer` (=48), so at `-ngl 48` all 48 transformer layers are on GPU but the final logit projection runs on CPU every decode step — that alone is the 39.81 → 13.67 drop (3×). The true all-VRAM number is `-ngl 99` (**39.81 tps**, bounded by reading model 14 GiB + KV 24.5 GiB per token on the RTX 6000). **The three-way verdict — decode tps at matched GiB moved off the GPU:** | off-GPU tail | rolling-KV C2 (DCA-on, §5d) | DCA-off window (§5d) | **`-ngl` CPU offload** | |---|---|---|---| | ~3.3 GiB | **37.2** | 35.9 | **7.0** (≈3.1 GiB) | | ~4.9 GiB | **37.2** | 34.8 | **4.5** (≈4.6 GiB) | | ~5.7 GiB | **14.5** | 10.8 | **3.4** (≈6.2 GiB) | | ~6.6 GiB | **11.7** | 9.3 | **2.7** (≈7.7 GiB) | | ~9.5 GiB | **7.3** | 6.6 | **2.2** (≈9.3 GiB) | | ~11 GiB | **5.6** | 5.2 | **1.9** (≈11 GiB) | **Rolling-KV spill beats upstream CPU-layer offload by ~3–8× at every matched off-GPU volume.** Of the two rolling-KV variants, **DCA-on C2 is fastest at every point** (it reads only the DCA chunks it needs; DCA-off window streams the whole tail) — the streaming `dca_fused` path is not merely on-par with the non-DCA reference, it edges past it. Even at the *shallowest* possible offload (just the LM head, 1.6 GiB), CPU offload (13.67) is already below rolling-KV at a 3 GiB tail (37). The reason is structural, not tuning: rolling-KV relocates **bytes** (streamed over the ~50 GB/s link while the GPU keeps computing), whereas `-ngl` relocates **compute** (CPU matmul at a fraction of GPU throughput, and the CPU-owned KV must be read on the CPU too). This is the quantitative case for rolling-KV as the overflow fallback: it degrades gracefully on the PCIe axis instead of falling off the CPU-compute cliff. (Compression remains the *primary* capacity lever above both — §0; spill of either kind is the fallback when compression is exhausted.) Needle PRESENT in every cell of all three ladders. **Full merged three-way ladder (decode / prefill tps).** The three modes don't share an identical x-grid — the two rolling-KV modes are keyed by **KV host-tail**, `-ngl` by **whole layers (weights+KV+compute) pushed to CPU** — so rows are aligned on *approximate GiB moved off the GPU* and each cell shows the mode's own nearest measured point (`—` = no cell near that bin). Format is **decode (prefill)**: | ≈ GiB off GPU | DCA-on streaming (C2r) | DCA-off rolling-KV window (C2r) | CPU layer offload, resident KV (`-ngl`) | |---|---|---|---| | **0 (all resident)** | **38.5** (2094) | **40.1** (1547) | **39.8** (—) `-ngl 99` | | ~0.9 | 37.3 (2016) | 35.8 (1342) | — | | ~1.6 | — | — | 13.67 (1439) `-ngl 48`¹ | | ~1.7 | 37.4 (2015) | 35.9 (1335) | — | | ~2.5 | 37.1 (2016) | 35.9 (1338) | — | | ~3.1–3.3 | 37.2 (2016) | 35.9 (1340) | 7.02 (1213) `-ngl 46` | | ~4.1 | 37.2 (2015) | 35.6 (1344) | — | | ~4.6–4.9 | 37.2 (2017) | 34.8 (1335) | 4.51 (1097) `-ngl 44` | | ~5.7 | **14.5** (1598) ← knee | **10.8** (1172) ← knee | — | | ~6.2 | — | — | 3.41 (1008) `-ngl 42` | | ~6.5–6.7 | 11.7 (1361) | 9.3 (1028) | — | | ~7.3–7.7 | 9.7 (1164) | 8.2 (933) | 2.67 (917) `-ngl 40` | | ~8.9 | — | 6.6 (782) | — | | ~9.3–9.6 | 7.3 (925) | — | 2.19 (847) `-ngl 38` | | ~10.8–10.9 | — | 5.2 (670) | 1.86 (795) `-ngl 36` | | ~12–12.4 | 5.6 (765) | — | 1.63 (745) `-ngl 34` | | ~14.7 | — | — | 1.38 (678) `-ngl 31` | ¹ `-ngl 48` is *not* all-VRAM — llamafile keeps the LM-head on CPU until `-ngl > n_layer`, so the 152k-vocab logit projection runs on CPU and alone costs 39.8→13.7 (3×). True all-VRAM is `-ngl 99`. Full-residency ceiling ≈ 38–40 tps for all three; in the spill regime the ordering is unambiguous: **DCA-on ≥ DCA-off ≫ CPU-offload.** --- ## 6. What ships - **Window mode = graceful-overflow fallback** (shipping default): linear degradation ∝ spilled bytes, correctness clean (needle + logit-equiv real_frac 0), no cliff pathology. Not a compression substitute. - **Rolling-KV perf line: CLOSED.** No further re-stream/copy-plumbing investment — the link is the wall. - **Capacity levers are residency-side:** quant-KV decode (#582, made bandwidth-bound), DCA, PolyKV, low-bit-resident. See [`hal_kv_compression.md`](hal_kv_compression.md) §4 + [`context.md`](context.md) §5b. ## 7. Provenance Consolidated 2026-07-10 from `rolling_kv.md`, `rolling_kv_step_prefetch.md`, `rolling_kv_compute_reserve.md`, `hal_kv_compression.md` §3a. Tasks: #583/#584 (3090 cliffs), #585–#589 (S0–S3 + boot fix), #623 (bug-1843), #588 (perf gate), #626 (eval refresh in progress). Memory: `reference_pcie_bandwidth_hosts`, `project_qwen_retrieval_distributed_hal`.