ManniX-ITA's picture
Upload folder using huggingface_hub
5ef4cc5 verified
|
Raw
History Blame
32.5 kB

Rolling-KV β€” overflow throughput evaluations (consolidated)

What this is. The measured performance record for rolling-KV / POSITION_WINDOW (streaming the host KV tail over PCIe when the cache overflows VRAM). The numbers were historically scattered across three feature docs and one study; this file is the single evals-side index. Design + code live in the feature docs; this file is results only, with back-links.

Sources consolidated here:

  • docs/features/rolling_kv.md β€” M7 design + M7-G gate + the bug-1843 fix + the #588 3090 cliff curve.
  • docs/features/rolling_kv_step_prefetch.md β€” the 3090 overflow-decode hyperbola (both model classes) + the step-prefetch design.
  • docs/features/rolling_kv_compute_reserve.md β€” the bs2 boot-OOM fix (bug-1342) that unblocked the RTX 6000 curve, + the prefill cliff.
  • docs/evaluations/hal_kv_compression.md Β§3a β€” the bs2 host-stream-exact spill cliff (route-(a), the reason compression wins over spill).

0. The one-line verdict

Rolling-KV's re-stream design is bandwidth-bound by physics: every decoded token restreams the spilled tail over PCIe, so decode tps follows t(token) β‰ˆ resident_compute_ms + spilled_bytes / PCIe_BW. After the bug-1843 compute-side fix, the curve sits within ~5 % of that model. No copy-plumbing lever can beat the link itself β†’ the rolling-KV perf line is CLOSED (NO-GO) for deep overflow; window mode ships as the graceful-overflow fallback (linear, correctness-clean, no cliff pathology). The real capacity lever is residency-side compression, not spill β€” see hal_kv_compression.md Β§4 and context.md Β§5b.


1. The mechanism (why it is bandwidth-bound)

POSITION_WINDOW keeps a resident window of recent cells and streams the older tail from host RAM every decode step. At n_q=1 an attention op's compute is a sub-ms VRAM read, but its tail-tile DMA is ~30Γ— that, so the op-level 1-tile-ahead ping-pong is DMA-bound from MiB 1. Full-attention models stream a tail on every layer; iSWA models stream only their few global layers β†’ an order-of-magnitude gap. Mechanism + the cs_sync cross-layer barrier: rolling_kv_step_prefetch.md Β§"Root cause".


2. 3090 overflow-decode cliff β€” the original hyperbola (2026-07-01)

Op-level streaming, pre-bug-1843. PCIe ~6.4 GB/s, +0.156 ms/MiB.

model attn streaming layers overflow decode shape
Gemma-4-A4B-128e iSWA ~5 global 71 β†’ 19.5 tps @236 MiB tail graceful hyperbola
Qwen2.5-14B-1M full 48 26.9 β†’ 0.86 tps @~1.2 GB tail ~10Γ— worse (48 streaming layers)

Qwen engages POSITION_WINDOW (GPU 60–84 % busy) β€” it is slow purely because full-attention restreams a tail on every layer. Source: rolling_kv_step_prefetch.md Β§"Problem (measured, 3090)". (#583/#584)


3. The compute-side fix β€” bug-1843 / patch 0098 (2026-07-05)

The S0-cliff redirects (#586 staging ring / #587 barrier-drop) were perf-inert on the 3090 (bug-1838: the PCIe link is already saturated β€” no copy-side slack). The real waste was a per-tile streaming_lse_kernel recompute that every head_dim ≀ 256 decode tile paid, because the Stage-3a LSE channel was gated on head_dim > 256. Patch 0098-rolling-kv-lse-decode arms the LSE channel for D ≀ 256 at all three streaming-FA sites.

Gate (3090, DSO df46dc36, Qwen3-8B-Q8_0 ctx 40960 / vt 14000):

  • spill decode 1.14 β†’ 6.04 tps (5.3Γ—, == NOOP ceiling), needle PRESENT
  • resident 42.5 tps unchanged; OPENCOTI_LSE_NOOP now speed-inert (recompute structurally off the path)
  • correctness: teacher-forced logit-equiv (ours-resident REF vs POSITION_WINDOW at 256/16384 cells resident) real_frac 0.0, frac_full_agree 1.0, mean_tv 0.0068 β†’ PASS

Source: rolling_kv.md Β§"2026-07-05 β€” bug-1843 shipped". (#623)


4. #588 perf gate β€” post-fix 3090 cliff curve β†’ CLOSED NO-GO (2026-07-06)

bug-1843-fixed DSO (df46dc36), 3090, Qwen3-8B-Q8_0, ctx 16384 (~13.1k-token prompt = 2304 MiB f16 KV), needle PRESENT in every cell.

vt (MiB) window split occupied tail decode tps ms/token
22000/16000/14000/13000 FULLY RESIDENT 0 42.3–42.5 23.6
12000 11008 / 16384 cells ~2.1k cells β‰ˆ 295 MiB 15.5 64.6
8000/4000/2000 256 / 16384 (floor) ~12.9k cells β‰ˆ 1815 MiB 3.33 300

The PCIe model fits within ~5 %: t β‰ˆ 23.6 ms + occupied_tail_bytes / 6.5 GB/s (vt12000 predicted 67 ms vs 64.6 measured; floor predicted 303 vs 300). Spill decode is now purely bandwidth-bound β€” the structural floor for a re-stream design.

Verdict: NO-GO for further re-stream perf work. The "within 10–15 % of resident" bar is only reachable for tails ≀ 25 MiB (170 cells) on a 6.5 GB/s 3090. No #586-class copy-plumbing beats the link. Source: rolling_kv.md Β§"2026-07-06 β€” #588 perf gate".


5. bs2 RTX 6000 β€” the serving-relevant host

opencoti serves on the RTX 6000 (PCIe ~50 GB/s, ~8Γ— the 3090's link β€” see memory reference_pcie_bandwidth_hosts). Two bs2 measurements exist plus one open cell.

5a. Host-stream-exact spill cliff (route-(a), simulated 24 GB card, 2026-07-01)

From hal_kv_compression.md Β§3a β€” pre-bug-1843, f16 KV, 8 GiB budget:

cell resident / host tail decode tps needle
f16 @32k all / 0 67.3 βœ“
f16 @49k 43k cells / 1152 MiB 57.6 βœ“
f16 @64k 43k / 4224 MiB 0.99 βœ“
f16 @98k 43k / 10368 MiB 0.71 βœ“
q4_0 @64k (resident) all / 0 27.1 βœ“
q4_0 @128k (resident) all / 0 16.2 βœ“

The cliff is catastrophic past ~2 GiB tail (0.99 tps @64k). Same-context resident compression is 27Γ— faster β€” this is why route-(a) spill is demoted to beyond-max-compression fallback and compression is the primary mechanism.

5b. Prefill cliff after the boot-OOM fix (bug-1342, 2026-07-01)

rolling_kv_compute_reserve.md Β§"Gate result": with the two-pass compute-reserve fix, deep-overflow configs boot on the RTX 6000; prefill falls 4000 β†’ ~450 tok/s as the host tail grows (expected cliff). This fix is what makes the Β§5c curve runnable on bs2 at all.

5c. bs2 real-spill characterization β€” qw35a3, 256k fine tail ladder, DCA-on vs DCA-off (2026-07-10, #631/#633)

This closes the Β§5c gap with the first fine bs2 spill curve, and it OVERTURNS the initial coarse read. Vehicle: the serving-size Qwen3.6-35B-A3B (qw35a3, hybrid DeltaNet+attention MoE, D256 attn, fast A3B-active decode), 256k ctx, f16 KV, fixed binary llamafile.dualctx633. A fine --vram-target ladder (200 MiB tail steps) run both DCA-on and DCA-off shows: there IS a graceful, bandwidth-bound overflow band (the earlier "hard cliff / no graceful band" claim is RETRACTED); the catastrophic collapse seen in the first coarse sweep was a DCA-specific bug (bug-2144), not a hardware limit and not a hybrid-recurrent property β€” the coarse sweep's smallest tail (1710 MiB) simply landed past the DCA-on knee, so it only ever sampled the buggy regime.

Two gate fixes landed first. (1) bug-2142 / #633 β€” --vram-target / --kv-residency-mode window were silently ignored for every hybrid-recurrent model (LLM_ARCH_QWEN35/QWEN35MOE, Nemotron-H, Falcon-H1): the llama_memory_hybrid / llama_memory_hybrid_iswa wrappers (llama-model.cpp create_memory ~L2041/L2022) built their internal attention llama_kv_cache with hardcoded residency defaults, so the sizer always logged FULLY RESIDENT regardless of the flag. Fix (opencoti-hook: rolling-KV hybrid residency (bug-2142), host-only, 6 edits/5 files): thread vram_target_mib / pcie_bw_gbps / kv_residency_mode / kv_compute_reserve_mib / kv_window_measure_pass through both hybrid ctors into the attention cache, mirroring the iSWA (L2100) / plain-dense (L2140) paths that already forwarded them. Verified-engaged + regression-clean (shield below). (2) Note bug-1839: --vram-target is a TOTAL device cap (weights+KV+compute), so a spill target must be < weights + resident-KV + compute; the old "24 GB card = --vram-target 9600" recipe (14B) now floors resident-KV to 0. Empirical tail↔target map for qw35a3 @256k: tail_MiB β‰ˆ 36708 βˆ’ vram_target_MiB (fits every cell to Β±5 MiB).

qw35a3 KV is tiny (hybrid β€” only the few full-attention layers grow): f16 KV = 5120 MiB @256k, 10240 MiB @512k (~20 KiB/tok), an order of magnitude under the 14B-1M full-attention class (Β§5a). On the 96 GB RTX 6000 it never needs to spill; the ladder forces spill via --vram-target to characterize a smaller card / multi-session KV pressure.

DCA-OFF β€” the correct rolling-KV behavior (bandwidth-bound, graceful)

The non-DCA build_attn overload routes spilling layers through the bug-1840 streaming fast-path (build_attn_mha_position_window), which FAs the resident window on-device and streams only the tail once. Result: a smooth, bandwidth-bound decline β€” the right shape.

host tail resident cells prefill tps decode tps achieved GiB/sΒΉ needle
0 (resident) 262144 / 262144 4799 123.5 β€” βœ“
~0 (VT 36308) resident 4797 123.0 β€” βœ“
80 MiB 258048 326 105.4 ~15 (fixed engage cost) βœ“
280 MiB 247808 326 103.6 ~29 βœ“
680 MiB 227328 326 104.4 ~71 βœ“
1080 MiB 206848 325 57.7 ~62 βœ“
1880 MiB 165888 324 31.7 ~60 βœ“
2680 MiB 124928 315 21.4 ~57 βœ“

ΒΉ tail_MiB / (1000/decode_tps) ms, tail-streamed-once framing. Knee at ~680–1080 MiB: decode holds within ~15 % of resident to a 680 MiB tail (104 vs 123.5), then declines cleanly, saturating at ~57–62 GiB/s β€” i.e. PCIe-bandwidth-bound near the RTX 6000 link, exactly the correct rolling-KV shape (Β§4's t β‰ˆ compute + tail_bytes/BW, now confirmed on the 50 GB/s bus). Needle PRESENT at every depth. The βˆ’15 % step at first spill (80 MiB β†’ 105 tps) is the fixed streaming-path engagement cost.

⚠ DCA-off PREFILL is broken though (bug-2145) β€” a second, mirror-image gap. The prefill column above cliffs from 4797 β†’ ~325 tok/s the instant any tail forms (80 MiB) and then stays FLAT ~325 all the way to a 2.7 GB tail (325/325/325/324/315 β€” tail-INDEPENDENT). Tail-independence is the smoking gun: a bandwidth-bound prefill would slow with the tail; a flat 325 means the fast prefill path is simply switched off whenever the window is active (a routing gap, not a copy cost). So DCA-off is graceful for decode but not prefill.

DCA-ON β€” perfect to ~800 MiB, then a sharp cliff (bug-2144 β€” NOW FIXED, see Β§5d)

⚠ HISTORICAL / pre-fix. The cliff below is the diagnosis that motivated the fix. bug-2144 is fixed by the C2 streaming dca_fused kernel (§5d): DCA-on spill now tracks/beats the DCA-off reference at matched tails, no cliff. The pre-fix ladder is retained because it is the evidence base for the mechanism; read it as "what was broken," not current behaviour.

| host tail | resident cells | prefill tps | decode tps | needle | |---|---|---|---|---|---| | 0 (resident) | 262144 | 4392 | 112.5 | βœ“ | | 205 MiB | 251648 | 4077 | 111.7 | βœ“ | | 405 MiB | 241408 | 4080 | 112.5 | βœ“ | | 605 MiB | 231168 | 4077 | 113.2 | βœ“ | | 805 MiB | 220928 | 4073 | 112.8 | βœ“ | | 1205 MiB | 200448 | 3668 | 2.93 ⚠ | βœ“ | | 1605 MiB | 179968 | 3011 | 3.09 ⚠ | βœ“ | | 2405 MiB | 139008 | 2330 | 3.19 ⚠ | βœ“ | | 3205 MiB | 98048 | 2010 | 3.33 ⚠ | βœ“ |

DCA-on is penalty-free (full 112–113 tps) out to an ~805 MiB tail β€” actually better than DCA-off in that band (DCA-off already pays βˆ’15 % at 80 MiB) β€” then cliffs 38Γ— in a single 400 MiB step (805 β†’ 1205 MiB: 112.8 β†’ 2.93 tps) and stays dead-flat 3 tps past it. Past-cliff decode is 341 ms/token: only 3.5 GiB/s on the tail-once framing (β‰ͺ link β†’ not bandwidth-bound), but **26 GiB/s** on the ~9 GiB full-KV-round-trip framing (window-D2H 3920 MiB + full-concat-H2D 5120 MiB) β€” the fingerprint of bug-2144.

Matched-tail contrast (the empirical lock on the code diagnosis): at ~1.1 GiB tail, DCA-off decodes 57.7 tps vs DCA-on 2.93 tps β€” ~20Γ— faster. Same model, same binary, same tail; the only difference is DCA on/off. This confirms the JOB-1 code diagnosis below.

Mechanism β€” bug-2144: DCA bypasses the streaming window path (confirmed)

The bug-1840 perf fast-path build_attn_mha_position_window (llama-graph.cpp:2503) is reached only from the non-DCA build_attn overload (llama-graph.cpp:2999-3007). DCA's build_attn_dca_core (dca.cpp:229) fetches K/V with the raw whole-cache accessors get_k/get_v (dca.cpp:251-252), which for a spilling layer return the cross-backend correctness fallback ggml_concat(win_v /*device*/, tail_v /*pinned host*/, 2) (llama-kv-cache.cpp:2585-2611; its own comment names build_attn_mha_position_window as the perf path it is not taking). The scheduler materialises that concat on the compute backend β†’ a synchronous ~2Γ— full-KV host↔device round-trip every decode step (bug-1840's documented "12.6 GiB/step both ways β†’ 0.99 tps at 4.2 GB tail", llama-graph.cpp:2989-2990); the async DMA tile overlap lives in ggml_cuda_streaming_flash_attn (fattn.cu:1170) which DCA never enters. This predicts the ~26 GiB/s effective full-KV rate, the flat-past-cliff plateau (round-trip volume β‰ˆ window-D2H + full-KV-H2D, ~constant), and the slight rise 2.93 β†’ 3.33 as the tail grows (smaller window β‡’ less window-D2H).

The sharp ~1 GiB step (805 β†’ 1205 MiB) β€” CONFIRMED as the ~1-chunk knee. The DCA chunk is 32768 cells β‰ˆ 640 MiB; the step sits between a 1.26-chunk tail (805 MiB, fast) and a 1.88-chunk tail (1205 MiB, slow) β€” the tail crossing into a second DCA chunk is the trigger, exactly as hypothesised. The C2 fix (Β§5d) reproduces this knee on the fixed binary as a graceful step down (37 β†’ 15 tps at the same crossing) instead of a cliff, since the streaming kernel reads only the chunks it needs rather than round-tripping the full KV. Below ~1 chunk of tail DCA-on stays entirely device-resident and pays nothing (unchanged pre/post fix).

Resident-compression anchors (no spill; DCA-on, everything on-device)

ctx KV type KV MiB prefill tps decode tps needle
256k f16/f16 5120 4358 116.3 βœ“
256k q8_0K/q4_0V ~2080 4241 75.2 βœ“
256k turbo2/turbo2 680 3664 46.4 βœ“
512k f16/f16 10240 2856 87.65 βœ“
512k q8_0K/q4_0V ~4110 2741 48.8 βœ“
512k turbo2/turbo2 1360 2384 27.3 βœ“

Regression shield. Non-hybrid 14B-1M-Q8_0 @65536, --vram-target 9600, fixed binary β†’ 256 / 65536 cells resident, host tail = 12240 MiB β€” byte-identical to the pre-fix binary, so the bug-2142 fix is isolated to the hybrid path.

Dense control (Qwen-27B-omni, full attention) β€” DEFERRED. Only the resident calibration landed: KV 16384 MiB @256k (full-attention β†’ ~10Γ— the hybrid KV, as expected), DCA-off resident decode 36.9 tps. The dense spill ladder did not run (GPU time). It is not blocking β€” the qw35a3 DCA-on-vs-off contrast already isolates bug-2144 as DCA-specific (not hybrid-specific) β€” but a clean single-GPU dense ladder is worth running later (after #632's tps work) to confirm the DCA-off graceful band and the DCA-on cliff reproduce on a dense arch.

The full 4-quadrant picture β€” phase Γ— DCA (two halves of one routing gap)

Both anomalies are the same defect seen from two sides: the streaming-window fast-path is wired for only two of the four (phase Γ— DCA) quadrants. Where it is wired, behaviour is graceful; where it is not, decode/prefill falls to a slow fallback.

DCA-OFF DCA-ON
DECODE βœ… graceful β€” bandwidth-bound, 123β†’104@680MiBβ†’21@2.7GB, ~57–62 GiB/s (streaming window path) βœ… graceful (bug-2144 FIXED, Β§5d) β€” was 112β†’2.9 cliff; C2 streaming dca_fused now tracks/beats DCA-off at matched tails
PREFILL ❌ cliff β€” 4797 β†’ flat ~325, tail-INDEPENDENT from 80 MiB to 2.7 GB (bug-2145: fast prefill path disabled under window mode) βœ… graceful β€” 4076@605MiB β†’ 3668@1.2GB β†’ 2010@2.7GB (declines with tail)

The diagonal was the tell: non-DCA-decode βœ“ and DCA-prefill βœ“ always took the fast path; non-DCA-prefill (bug-2145) and DCA-decode (bug-2144) were the two uncovered quadrants. DCA-decode is now covered (bug-2144 FIXED, Β§5d) β€” the C2 streaming-tail kernel routes DCA spill through the same async tile machinery. bug-2145 remains open (non-DCA prefill): its tail-independence (80 MiB costs the same 325 tok/s as 2.7 GB) proves it is a path switch β€” the fast prefill kernel is turned off the moment window_cells < kv_size, regardless of how little actually spilled β€” not a bandwidth cost. It is an additive routing/plumbing gap around the same build_attn_mha_position_window / ggml_streaming_flash_attn machinery, not a model or hardware limit. Neither quadrant ever lost the needle β€” correctness is clean throughout; these are purely throughput-path gaps.

Verdict (revised)

  • A graceful overflow band exists on bs2 β€” for DECODE. With DCA off, qw35a3 spill decode is bandwidth-bound (~57–62 GiB/s, needle clean) and holds within ~15 % of resident to a ~680 MiB tail β€” the correct rolling-KV behavior on the 50 GB/s link. The Β§0/Β§4 "deep overflow is bandwidth-bound" verdict stands (for decode); the "no graceful band for hybrid" over-claim is withdrawn. DCA-on decode is now graceful too (bug-2144 FIXED, Β§5d). The one remaining non-graceful quadrant is DCA-off prefill (bug-2145) β€” a DCA-off spill config still cliffs on prefill; closing it needs the same streaming-window routing applied to the non-DCA prefill path.
  • DCA-on had a ~1 GiB decode cliff (bug-2144) β€” FIXED (Β§5d, C2 streaming dca_fused). Was perfect below 800 MiB tail then a full-KV concat round-trip (20–55Γ— slower). The C2 streaming-tail kernel now makes DCA-on spill track/beat the DCA-off reference at matched tails (18Γ— recovery at a 6.4 GiB tail). Gated WS2_DCA_STREAM=1 pending ship green-light.
  • Compression remains the primary capacity lever (Β§0, hal_kv_compression.md Β§3a/Β§4): turbo2 keeps qw35a3 resident at 0.68/1.36 GiB KV (46/27 tps), and qw35a3's KV is small enough to stay resident on a 96 GB card regardless. Spill (DCA-off) is a valid graceful-overflow fallback for smaller cards / multi-session pressure; the bug-2144 fix (Β§5d) now extends that graceful fallback to DCA-on configs too.

Contention caveat. These ran on GPU1 with #632's tps work on GPU0 concurrently. GPU1 is a separate device (isolated VRAM/compute), and the load-bearing conclusions β€” the ~20Γ— DCA-off vs DCA-on gap, the DCA-off bandwidth saturation, the DCA-on cliff position β€” are far outside any cross-GPU noise and internally consistent (monotone ladders). Fine-grained tps (Β±1–2 tps between adjacent cells) may carry minor host/PCIe contention; correctness (needle) and the resident/tail splits are contention-immune. Method: fixed binary, DSO 2377635512 (bug-2141 guard, inert for these non-degenerate rows), no rebuild during the run.

5d. bug-2144 C2 fix β€” DCA-on spill made graceful (streaming dca_fused), Qwen2.5-14B-1M @128k (2026-07-10)

The fix. bug-2144's DCA-on decode cliff (Β§5c: full-KV concat round-trip, ~55Γ— collapse past ~1 chunk of tail) is eliminated by giving the fused DCA kernel a two-region window+tail streaming path β€” the same async tile ping-pong the non-DCA build_attn_mha_position_window uses. Host plumbing threads k_tail/v_tail/n_win into ggml_flash_attn_ext_dca_fused (tail on src[7]/src[8], op_params[5]=n_win); gated behind WS2_DCA_STREAM=1, decode-only via WS2_NQ_MAX (n_q ≀ 16). Legacy concat path is byte-identical when off. Marker opencoti-hook: DCA streaming-tail (bug-2144-C / Stage C) (dca.cpp).

Decisive ladder β€” DCA-on (C2 stream) vs DCA-off (rolling-KV reference). Both modes re-measured on the reserve-fixed C2r binary (bug-2146 below), so these are the clean, accurate numbers β€” no reserve artifact, no run contamination. Qwen2.5-14B-1M Q8_0, f16 KV, -c 131072, needle QUASAR-4417, bs2 RTX 6000 single-GPU, ~51 GB/s eff link. Both prefill / decode tps recorded; needle PRESENT in every cell of both ladders:

vram-target DCA-on tail DCA-on prefill / decode DCA-off tail DCA-off prefill / decode
41888 0 (resident) 2094 / 38.50 0 (resident) 1547 / 40.13
41000 ~864 MiB 2016 / 37.27 ~864 MiB 1342 / 35.82
40200 ~1680 2015 / 37.35 ~1680 1335 / 35.92
39400 ~2496 2016 / 37.06 ~2496 1338 / 35.91
38600 ~3264 2016 / 37.23 ~3264 1340 / 35.92
37800 ~4080 2015 / 37.23 ~4080 1344 / 35.56
37000 ~4896 2017 / 37.19 ~4896 1335 / 34.75
36200 ~5808 1598 / 14.45 (knee) ~5664 1172 / 10.78 (knee)
35400 ~6720 1361 / 11.69 ~6480 1028 / 9.29
34600 ~7680 1164 / 9.72 ~7296 933 / 8.17
33000 ~9600 925 / 7.29 ~8880 782 / 6.55
31000 ~12000 765 / 5.56 ~10896 670 / 5.22

Verdict: the C2 streaming path makes DCA-on spill track and beat the DCA-off "normal" rolling-KV reference at every point β€” both hold their full-speed plateau out to a ~4.9 GiB tail (DCA-on 37.2, DCA-off ~35.8), knee together at **5.6 GiB** (tail crossing 1 DCA chunk = 32768 cells), then the same graceful bandwidth-bound rolloff, no cliff. DCA-on decode is faster than DCA-off at every cell (plateau and deep tail) because the fused kernel reads only the chunks it needs, whereas DCA-off window mode streams the whole tail; DCA-on prefill is also higher on the plateau (2015 vs 1340). Pre-fix DCA-on cliffed to ~0.68 tps past ~1 GiB of tail β†’ now 11.69 at a 6.7 GiB tail = **17Γ— recovery.** bug-2144 CLOSED on perf. (vt=41888 DCA-on decode 38.50 confirms full residency under the bug-2146 fix; the harness's tail-grep prints a stale iter-0 value there because the final "FULLY RESIDENT" boot line carries no "host tail =" token β€” decode is ground truth.)

Caveat for the pending bug-2145 fix (non-DCA prefill cliff, #636): that fix, if it lands, can only move the DCA-off prefill column (when the window engages) β€” decode is untouched in both modes and the DCA-on ladder is entirely unaffected. Worst case is a cheap one-column prefill re-run of the dcaoff mode, never the whole matrix.

Reserve artifact β€” root cause of the pre-fix forced β‰₯3 GiB tail (now fixed; the ladder above is already the fixed run). The bug-1342 two-pass compute-reserve sizer measures pass 1 with a degenerate wc=256 window (llama-kv-cache.cpp measure_pass), which for the streaming graph means a maximal ~130816-cell host tail. Instrumented boot (-lv 5, vt=41888):

Pass 1 (measure, wc=256, tail=130816):  CUDA0 compute buffer = 4088.81 MiB β†’ reserve 4344 MiB
Pass 2 (real,    wc=114176, tail=16896): CUDA0 compute buffer =  926.01 MiB

The streaming decode graph's real device compute buffer is only 926 MiB; the reserve holds back 4344 MiB β†’ a 3.16 GiB over-reservation, a pure measurement artifact of the wc=256 pass (confirmed vs a WS2_NQ_MAX=0 control: 18216 vs 15054 MiB budget at identical free VRAM). This is why vt=41888 was forced to a 3168 MiB tail instead of staying fully resident.

FIX β€” bug-2146, IMPLEMENTED + validated (host-only, no DSO rebuild). llama_context ctor (llama-context.cpp) now iterates the two-pass reserve to a fixpoint instead of trusting the single degenerate wc=256 measurement: each re-size grows the window β†’ shrinks the tail β†’ shrinks the measured buffer, so the reserve decreases monotonically and converges (to the plain resident buffer, tail 0, when the KV fits). The resident graph's buffer is a subset of the streaming graph's, so a re-size only ever reveals a smaller-or-equal buffer β€” it can never under-reserve. Capped at 3 iterations; stops as soon as the reserve stops shrinking (128 MiB margin). For the non-streaming case pass 2 measures what pass 1 did β†’ converges in ONE step, byte-identical to prior behaviour. Validated on the same C2r binary (Qwen2.5-14B-1M Q8_0 @128k, bs2):

config pre-fix post-fix (bug-2146)
vt=41888 (KV fits) 3168 MiB tail, 37.10 tps FULLY RESIDENT, tail 0, 38.40 tps (converged iter 0β†’1β†’2)
vt=35400 (genuine spill) 9648 MiB tail, 7.24 tps 6720 MiB tail, 11.72 tps (+62%) (reserve 14966β†’17858 budget)

Needle PRESENT both cells. This recovers the ~3.16 GiB and, because the corrected (smaller) reserve grows the resident window at every vram-target, every Β§5d/Β§5e streaming ladder cell above was measured with the over-large reserve and improves under the fix (the tables are the conservative lower bound). Still-open cleanup: strip leftover get_k_tail: ENTRY / ACTUAL tail LLAMA_LOG_WARNs (llama-kv-cache.cpp:3252/3389/3408) before ship.

5e. Rolling-KV spill vs upstream CPU-layer offload β€” bytes vs compute (Qwen2.5-14B-1M @128k, 2026-07-10)

The honest baseline for "what does everyone else do when VRAM runs out" is not rolling-KV β€” it is upstream llama.cpp's -ngl < n_layer, which pushes whole transformer layers (weights + their KV + their compute) to CPU. This section pits that against rolling-KV spill on the same model/host, to show the two overflow mechanisms are categorically different.

Two ways to move data off a too-small card:

  • Rolling-KV spill β€” all compute stays on GPU; only KV bytes stream over PCIe (window on device, tail on host). Cost is PCIe-bandwidth-bound.
  • -ngl layer offload β€” whole layers' compute moves to CPU. Cost is CPU-matmul-bound (and the CPU also owns those layers' KV). This is what stock llama.cpp does when the model+KV won't fit.

-ngl offload ladder (DCA off, full-KV-in-VRAM for GPU-resident layers, -c 131072, needle QUASAR-4417, bs2 RTX 6000; RAM column = weights 279 MiB + KV 512 MiB per CPU layer, + ~790 MiB output tensor which drops to CPU at any -ngl ≀ 48, + ~790 MiB token-embedding always on CPU):

-ngl layers on CPU β‰ˆ RAM offloaded prefill tps decode tps needle
99 0 (output on GPU) ~0.8 GiB (embd) β€” 39.81 (all-VRAM baseline) βœ“
48 0 (outputβ†’CPU) ~1.6 GiB 1439 13.67 βœ“
46 2 ~3.1 GiB 1213 7.02 βœ“
44 4 ~4.6 GiB 1097 4.51 βœ“
42 6 ~6.2 GiB 1008 3.41 βœ“
40 8 ~7.7 GiB 917 2.67 βœ“
38 10 ~9.3 GiB 847 2.19 βœ“
36 12 ~10.8 GiB 795 1.86 βœ“
34 14 ~12.4 GiB 745 1.63 βœ“
32 16 ~13.9 GiB 721 1.44 βœ“
31 17 ~14.7 GiB 678 1.38 βœ“

Gotcha (load-bearing for reading this table): -ngl 48 is not all-in-VRAM. llamafile keeps the output/LM-head projection (152k-vocab Γ— 5120) on CPU until -ngl > n_layer (=48), so at -ngl 48 all 48 transformer layers are on GPU but the final logit projection runs on CPU every decode step β€” that alone is the 39.81 β†’ 13.67 drop (3Γ—). The true all-VRAM number is -ngl 99 (39.81 tps, bounded by reading model 14 GiB + KV 24.5 GiB per token on the RTX 6000).

The three-way verdict β€” decode tps at matched GiB moved off the GPU:

off-GPU tail rolling-KV C2 (DCA-on, Β§5d) DCA-off window (Β§5d) -ngl CPU offload
~3.3 GiB 37.2 35.9 7.0 (β‰ˆ3.1 GiB)
~4.9 GiB 37.2 34.8 4.5 (β‰ˆ4.6 GiB)
~5.7 GiB 14.5 10.8 3.4 (β‰ˆ6.2 GiB)
~6.6 GiB 11.7 9.3 2.7 (β‰ˆ7.7 GiB)
~9.5 GiB 7.3 6.6 2.2 (β‰ˆ9.3 GiB)
~11 GiB 5.6 5.2 1.9 (β‰ˆ11 GiB)

Rolling-KV spill beats upstream CPU-layer offload by ~3–8Γ— at every matched off-GPU volume. Of the two rolling-KV variants, DCA-on C2 is fastest at every point (it reads only the DCA chunks it needs; DCA-off window streams the whole tail) β€” the streaming dca_fused path is not merely on-par with the non-DCA reference, it edges past it. Even at the shallowest possible offload (just the LM head, 1.6 GiB), CPU offload (13.67) is already below rolling-KV at a 3 GiB tail (37). The reason is structural, not tuning: rolling-KV relocates bytes (streamed over the ~50 GB/s link while the GPU keeps computing), whereas -ngl relocates compute (CPU matmul at a fraction of GPU throughput, and the CPU-owned KV must be read on the CPU too). This is the quantitative case for rolling-KV as the overflow fallback: it degrades gracefully on the PCIe axis instead of falling off the CPU-compute cliff. (Compression remains the primary capacity lever above both β€” Β§0; spill of either kind is the fallback when compression is exhausted.) Needle PRESENT in every cell of all three ladders.

Full merged three-way ladder (decode / prefill tps). The three modes don't share an identical x-grid β€” the two rolling-KV modes are keyed by KV host-tail, -ngl by whole layers (weights+KV+compute) pushed to CPU β€” so rows are aligned on approximate GiB moved off the GPU and each cell shows the mode's own nearest measured point (β€” = no cell near that bin). Format is decode (prefill):

β‰ˆ GiB off GPU DCA-on streaming (C2r) DCA-off rolling-KV window (C2r) CPU layer offload, resident KV (-ngl)
0 (all resident) 38.5 (2094) 40.1 (1547) 39.8 (β€”) -ngl 99
~0.9 37.3 (2016) 35.8 (1342) β€”
~1.6 β€” β€” 13.67 (1439) -ngl 48ΒΉ
~1.7 37.4 (2015) 35.9 (1335) β€”
~2.5 37.1 (2016) 35.9 (1338) β€”
~3.1–3.3 37.2 (2016) 35.9 (1340) 7.02 (1213) -ngl 46
~4.1 37.2 (2015) 35.6 (1344) β€”
~4.6–4.9 37.2 (2017) 34.8 (1335) 4.51 (1097) -ngl 44
~5.7 14.5 (1598) ← knee 10.8 (1172) ← knee β€”
~6.2 β€” β€” 3.41 (1008) -ngl 42
~6.5–6.7 11.7 (1361) 9.3 (1028) β€”
~7.3–7.7 9.7 (1164) 8.2 (933) 2.67 (917) -ngl 40
~8.9 β€” 6.6 (782) β€”
~9.3–9.6 7.3 (925) β€” 2.19 (847) -ngl 38
~10.8–10.9 β€” 5.2 (670) 1.86 (795) -ngl 36
~12–12.4 5.6 (765) β€” 1.63 (745) -ngl 34
~14.7 β€” β€” 1.38 (678) -ngl 31

ΒΉ -ngl 48 is not all-VRAM β€” llamafile keeps the LM-head on CPU until -ngl > n_layer, so the 152k-vocab logit projection runs on CPU and alone costs 39.8β†’13.7 (3Γ—). True all-VRAM is -ngl 99. Full-residency ceiling β‰ˆ 38–40 tps for all three; in the spill regime the ordering is unambiguous: DCA-on β‰₯ DCA-off ≫ CPU-offload.


6. What ships

  • Window mode = graceful-overflow fallback (shipping default): linear degradation ∝ spilled bytes, correctness clean (needle + logit-equiv real_frac 0), no cliff pathology. Not a compression substitute.
  • Rolling-KV perf line: CLOSED. No further re-stream/copy-plumbing investment β€” the link is the wall.
  • Capacity levers are residency-side: quant-KV decode (#582, made bandwidth-bound), DCA, PolyKV, low-bit-resident. See hal_kv_compression.md Β§4 + context.md Β§5b.

7. Provenance

Consolidated 2026-07-10 from rolling_kv.md, rolling_kv_step_prefetch.md, rolling_kv_compute_reserve.md, hal_kv_compression.md Β§3a. Tasks: #583/#584 (3090 cliffs), #585–#589 (S0–S3 + boot fix), #623 (bug-1843), #588 (perf gate), #626 (eval refresh in progress). Memory: reference_pcie_bandwidth_hosts, project_qwen_retrieval_distributed_hal.