File size: 29,891 Bytes
b69d9d8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 | From: opencoti
Subject: [PATCH 2/N] lazy slot-context allocation β Phase 2 (grow on demand + partial zero-fill)
Phase 2 of the F4 M3 lazy slot-context rollout. Builds on Phase 1
(deferred zero-fill) by adding a soft allocation cap that
find_slot honors, a grow-on-demand path that bumps the cap when
real prompts need more cells, and a partial-range zero-fill so
RSS commits scale with actual workload β not with the configured
n_ctx_slot Γ slot_count product.
The end-state for the production embedder
(qwen3-0.6b, --parallel 2 --ctx-size 16384, CPU mode):
- With --slot-initial-ctx 0 (the default) Phase 2 is a no-op
layered on Phase 1: cap == n_ctx_seq from construction; full
buffer cleared on first batch (matches Phase 1 RSS profile).
- With --slot-initial-ctx 4096 (the recommended embedder
setting) the warm RSS scales with whatever cells real prompts
have touched. A run that only ever embeds 8β56-token requests
commits ~4096 Γ per-cell bytes, not the full 16384.
The patch threads `slot_initial_ctx` through:
- public API: llama_context_params (include/llama.h)
- internal: llama_cparams (src/llama-cparams.h)
- ctor chain: llama_context::ctor (src/llama-context.cpp)
llama_kv_cache::ctor (src/llama-kv-cache.{h,cpp})
llama_kv_cache_iswa::ctor (src/llama-kv-cache-iswa.{h,cpp})
llama_model::create_memory call sites (src/llama-model.cpp)
- hybrid mem: src/llama-memory-hybrid.cpp + -iswa.cpp pass 0
(kept on Phase-1 behavior β recurrent/hybrid models
are not the F4 embedder target)
- CLI: common_params.slot_initial_ctx + --slot-initial-ctx
(common/common.{h,cpp}, common/arg.cpp)
Mechanism (inside llama_kv_cache):
- kv_size_max: hard cap (== constructor's kv_size).
- kv_size_alloc: soft cap; cells [0, kv_size_alloc) visible to
find_slot. Starts at min(slot_initial_ctx,
kv_size_max) β Phase 1 default of 0 collapses
to kv_size_max so find_slot sees everything.
- n_cells_cleared: half-open range of cells already zero-filled.
Phase 1's `bool cleared` is replaced; the cells
beyond n_cells_cleared have malloc'd address
space but uncommitted pages.
- ensure_cleared(up_to):
Partial memset of cells [n_cells_cleared, up_to) on each
layer's K and V tensor across every stream, using
ggml_backend_tensor_memset (host-or-device dispatch). The
common inference path passes up_to=0 which expands to the
current soft cap; state_read passes up_to=kv_size_max for
full coverage on the deserialize path.
- ensure_capacity(cells_needed):
Geometric grow (next-pow2, n_pad-aligned) of kv_size_alloc
up to kv_size_max. Monotonic β never shrinks (Phase 3 adds
shrink-on-idle).
- prepare(ubatches):
When find_slot fails inside the current soft cap, retries
after ensure_capacity(used_max + ubatch.n_tokens) +
ensure_cleared(). Bounded by kv_size_max; a genuine
overflow at the hard cap propagates up unchanged.
Bench: docs/decisions/0001-lazy-slot-context.md Β§6.2 Phase 2 row.
Milestone: F4 M3 β see docs/features/memory_embedder.md
Design: docs/decisions/0001-lazy-slot-context.md
Upstreaming: deferred; propose to llamafile + llama.cpp
maintainers after F4 M3's full phase-set ships and bench
numbers are public.
---
diff --git a/llama.cpp/common/arg.cpp b/llama.cpp/common/arg.cpp
--- a/llama.cpp/common/arg.cpp
+++ b/llama.cpp/common/arg.cpp
@@ -1361,6 +1361,16 @@ common_params_context common_params_parser_init(common_params & params, llama_ex
params.kv_unified = value;
}
).set_env("LLAMA_ARG_KV_UNIFIED").set_examples({LLAMA_EXAMPLE_SERVER, LLAMA_EXAMPLE_PERPLEXITY, LLAMA_EXAMPLE_BATCHED, LLAMA_EXAMPLE_BENCH, LLAMA_EXAMPLE_PARALLEL}));
+ add_opt(common_arg(
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ {"--slot-initial-ctx"}, "N",
+ string_format("initial KV-cell allocation per sequence (default: %d, 0 = full n_ctx_seq, "
+ "positive values cap initial allocation and enable grow-on-demand)",
+ params.slot_initial_ctx),
+ [](common_params & params, int value) {
+ params.slot_initial_ctx = value;
+ }
+ ).set_env("LLAMA_ARG_SLOT_INITIAL_CTX").set_examples({LLAMA_EXAMPLE_SERVER, LLAMA_EXAMPLE_EMBEDDING, LLAMA_EXAMPLE_PARALLEL}));
add_opt(common_arg(
{"--cache-idle-slots"},
{"--no-cache-idle-slots"},
diff --git a/llama.cpp/common/common.cpp b/llama.cpp/common/common.cpp
--- a/llama.cpp/common/common.cpp
+++ b/llama.cpp/common/common.cpp
@@ -1628,6 +1628,8 @@ struct llama_context_params common_context_params_to_llama(const common_params &
cparams.n_ctx = params.n_ctx;
cparams.n_seq_max = params.n_parallel;
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ cparams.slot_initial_ctx = params.slot_initial_ctx < 0 ? 0u : (uint32_t) params.slot_initial_ctx;
cparams.n_rs_seq = params.speculative.need_n_rs_seq();
cparams.n_batch = params.n_batch;
cparams.n_ubatch = params.n_ubatch;
diff --git a/llama.cpp/common/common.h b/llama.cpp/common/common.h
--- a/llama.cpp/common/common.h
+++ b/llama.cpp/common/common.h
@@ -425,6 +425,11 @@ struct ggml_opt_optimizer_params common_opt_lr_pars(void * userdata);
struct common_params {
int32_t n_predict = -1; // max. number of new tokens to predict, -1 == no limit
int32_t n_ctx = 0; // context size, 0 == context the model was trained with
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Initial KV-cell allocation per sequence. 0 = full n_ctx_seq allocation
+ // (Phase 1 behavior β only zero-fill is deferred). Positive values cap
+ // the initial allocation and enable grow-on-demand up to n_ctx_seq.
+ int32_t slot_initial_ctx = 0;
int32_t n_batch = 2048; // logical batch size for prompt processing (must be >=32 to use BLAS)
int32_t n_ubatch = 512; // physical batch size for prompt processing (must be >=32 to use BLAS)
int32_t n_keep = 0; // number of tokens to keep from initial prompt
diff --git a/llama.cpp/include/llama.h b/llama.cpp/include/llama.h
--- a/llama.cpp/include/llama.h
+++ b/llama.cpp/include/llama.h
@@ -338,6 +338,14 @@ extern "C" {
uint32_t n_batch; // logical maximum batch size that can be submitted to llama_decode
uint32_t n_ubatch; // physical maximum batch size
uint32_t n_seq_max; // max number of sequences (i.e. distinct states for recurrent models)
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Initial KV-cell allocation per sequence. 0 = behave like Phase 1
+ // (allocate the full n_ctx_seq up front, only zero-fill is deferred).
+ // Positive values cap the initial allocation at min(slot_initial_ctx,
+ // n_ctx_seq); subsequent prompts that exceed this cap trigger a grow
+ // up to next_pow2 (still capped at n_ctx_seq). Reduces warm-state RSS
+ // when typical prompts are much shorter than n_ctx_seq.
+ uint32_t slot_initial_ctx;
uint32_t n_rs_seq; // number of recurrent-state snapshots per seq for rollback (0 = no rollback) [EXPERIMENTAL]
int32_t n_threads; // number of threads to use for generation
int32_t n_threads_batch; // number of threads to use for batch processing
diff --git a/llama.cpp/src/llama-context.cpp b/llama.cpp/src/llama-context.cpp
--- a/llama.cpp/src/llama-context.cpp
+++ b/llama.cpp/src/llama-context.cpp
@@ -58,6 +58,9 @@ llama_context::llama_context(
cparams.n_rs_seq = 0;
}
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ cparams.slot_initial_ctx = params.slot_initial_ctx;
+
cparams.n_threads = params.n_threads;
cparams.n_threads_batch = params.n_threads_batch;
cparams.yarn_ext_factor = params.yarn_ext_factor >= 0.0f ? params.yarn_ext_factor : hparams.yarn_ext_factor;
@@ -3346,6 +3349,7 @@ llama_context_params llama_context_default_params() {
/*.n_batch =*/ 2048,
/*.n_ubatch =*/ 512,
/*.n_seq_max =*/ 1,
+ /*.slot_initial_ctx =*/ 0,
/*.n_rs_seq =*/ 0,
/*.n_threads =*/ GGML_DEFAULT_N_THREADS, // TODO: better default
/*.n_threads_batch =*/ GGML_DEFAULT_N_THREADS,
diff --git a/llama.cpp/src/llama-cparams.h b/llama.cpp/src/llama-cparams.h
--- a/llama.cpp/src/llama-cparams.h
+++ b/llama.cpp/src/llama-cparams.h
@@ -12,6 +12,10 @@ struct llama_cparams {
uint32_t n_batch;
uint32_t n_ubatch;
uint32_t n_seq_max;
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // 0 = full n_ctx_seq allocation. Positive value caps the initial KV-cell
+ // allocation per sequence and enables grow-on-demand up to n_ctx_seq.
+ uint32_t slot_initial_ctx;
uint32_t n_rs_seq; // number of recurrent-state snapshots per seq for rollback
int32_t n_threads; // number of threads to use for generation
int32_t n_threads_batch; // number of threads to use for batch processing
diff --git a/llama.cpp/src/llama-kv-cache-iswa.cpp b/llama.cpp/src/llama-kv-cache-iswa.cpp
--- a/llama.cpp/src/llama-kv-cache-iswa.cpp
+++ b/llama.cpp/src/llama-kv-cache-iswa.cpp
@@ -20,6 +20,7 @@ llama_kv_cache_iswa::llama_kv_cache_iswa(
bool swa_full,
bool unified,
uint32_t kv_size,
+ uint32_t kv_size_initial,
uint32_t n_seq_max,
uint32_t n_ubatch,
uint32_t n_pad,
@@ -61,14 +62,14 @@ llama_kv_cache_iswa::llama_kv_cache_iswa(
kv_base = std::make_unique<llama_kv_cache>(
model, type_k, type_v,
- v_trans, offload, unified, size_base, n_seq_max, n_pad,
+ v_trans, offload, unified, size_base, kv_size_initial, n_seq_max, n_pad,
0, LLAMA_SWA_TYPE_NONE, filter_base, reuse);
LLAMA_LOG_INFO("%s: creating SWA KV cache, size = %u cells\n", __func__, size_swa);
kv_swa = std::make_unique<llama_kv_cache>(
model, type_k, type_v,
- v_trans, offload, unified, size_swa, n_seq_max, n_pad,
+ v_trans, offload, unified, size_swa, kv_size_initial, n_seq_max, n_pad,
hparams.n_swa, hparams.swa_type, filter_swa, reuse);
}
diff --git a/llama.cpp/src/llama-kv-cache-iswa.h b/llama.cpp/src/llama-kv-cache-iswa.h
--- a/llama.cpp/src/llama-kv-cache-iswa.h
+++ b/llama.cpp/src/llama-kv-cache-iswa.h
@@ -22,6 +22,9 @@ public:
bool swa_full,
bool unified,
uint32_t kv_size,
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Forwarded to both kv_base and kv_swa. 0 = full kv_size.
+ uint32_t kv_size_initial,
uint32_t n_seq_max,
uint32_t n_ubatch,
uint32_t n_pad,
diff --git a/llama.cpp/src/llama-kv-cache.cpp b/llama.cpp/src/llama-kv-cache.cpp
--- a/llama.cpp/src/llama-kv-cache.cpp
+++ b/llama.cpp/src/llama-kv-cache.cpp
@@ -85,6 +85,7 @@ llama_kv_cache::llama_kv_cache(
bool offload,
bool unified,
uint32_t kv_size,
+ uint32_t kv_size_initial,
uint32_t n_seq_max,
uint32_t n_pad,
uint32_t n_swa,
@@ -96,6 +97,24 @@ llama_kv_cache::llama_kv_cache(
GGML_ASSERT(kv_size % n_pad == 0);
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Initialize the hard cap and the soft cap. kv_size_initial==0 (the
+ // legacy / Phase-1 default) means "no soft cap" so the whole kv_size
+ // is visible to find_slot. A positive value is clamped to [n_pad,
+ // kv_size] and rounded up to the next n_pad multiple.
+ kv_size_max = kv_size;
+ if (kv_size_initial == 0) {
+ kv_size_alloc = kv_size_max;
+ } else {
+ uint32_t cap = std::min(kv_size_initial, kv_size_max);
+ cap = std::max(cap, n_pad);
+ cap = ((cap + n_pad - 1) / n_pad) * n_pad;
+ cap = std::min(cap, kv_size_max);
+ kv_size_alloc = cap;
+ LLAMA_LOG_INFO("%s: slot_initial_ctx = %u cells (cap = %u)\n",
+ __func__, kv_size_alloc, kv_size_max);
+ }
+
const uint32_t n_layer_kv = hparams.n_layer_kv();
// define a comparator for the buft -> ctx map to ensure that the order is well-defined:
@@ -287,7 +306,9 @@ llama_kv_cache::llama_kv_cache(
}
if (model.hparams.no_alloc) {
- cleared = true;
+ // opencoti F4 M3 Phase 1+2 β dummy buffers are size 0; treat all
+ // cells as "already cleared" so ensure_cleared is a no-op.
+ n_cells_cleared = kv_size_max;
} else {
LLAMA_LOG_INFO("%s: KV cache zero-fill deferred until first batch\n", __func__);
}
@@ -348,20 +369,81 @@ llama_kv_cache::llama_kv_cache(
debug = LLAMA_KV_CACHE_DEBUG ? atoi(LLAMA_KV_CACHE_DEBUG) : 0;
}
-// opencoti F4 M3 Phase 1 β see docs/decisions/0001-lazy-slot-context.md
-void llama_kv_cache::ensure_cleared() {
- if (cleared) {
+// opencoti F4 M3 Phase 1+2 β see docs/decisions/0001-lazy-slot-context.md
+void llama_kv_cache::ensure_cleared(uint32_t up_to_cells) {
+ // up_to_cells == 0 means "clear up to the current soft cap". Callers
+ // that need full coverage (state_read) pass kv_size_max.
+ const uint32_t target = up_to_cells == 0
+ ? kv_size_alloc
+ : std::min(up_to_cells, kv_size_max);
+
+ if (n_cells_cleared >= target) {
return;
}
- for (auto & [_, buf] : ctxs_bufs) {
- LLAMA_LOG_INFO("%s: %10s KV buffer zero-fill = %8.2f MiB (deferred)\n", __func__,
- ggml_backend_buffer_name(buf.get()),
- ggml_backend_buffer_get_size(buf.get())/1024.0/1024.0);
- ggml_backend_buffer_clear(buf.get(), 0);
+ const uint32_t a = n_cells_cleared;
+ const uint32_t b = target;
+
+ // Per-layer, per-stream partial memset of the K and V tensors.
+ // ggml_backend_tensor_memset handles the host/device dispatch β for the
+ // CPU backend (the only one the embedder uses today) this is a direct
+ // memset on the malloc'd buffer; for offloaded buffers it's a backend
+ // op. Either way only bytes for cells [a, b) get touched, so on Linux
+ // only those pages get committed to RSS.
+ for (const auto & layer : layers) {
+ for (uint32_t st = 0; st < n_stream; ++st) {
+ if (layer.k) {
+ const size_t off = (size_t) st * (size_t) layer.k->nb[2]
+ + (size_t) a * (size_t) layer.k->nb[1];
+ const size_t sz = (size_t) (b - a) * (size_t) layer.k->nb[1];
+ ggml_backend_tensor_memset(layer.k, 0, off, sz);
+ }
+ if (layer.v) {
+ const size_t off = (size_t) st * (size_t) layer.v->nb[2]
+ + (size_t) a * (size_t) layer.v->nb[1];
+ const size_t sz = (size_t) (b - a) * (size_t) layer.v->nb[1];
+ ggml_backend_tensor_memset(layer.v, 0, off, sz);
+ }
+ }
+ }
+
+ LLAMA_LOG_INFO("%s: cleared KV cells [%u, %u) of %u (%s)\n",
+ __func__, a, b, kv_size_max,
+ up_to_cells == 0 ? "soft cap" : "full");
+
+ n_cells_cleared = b;
+}
+
+// opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+bool llama_kv_cache::ensure_capacity(uint32_t cells_needed) {
+ if (cells_needed <= kv_size_alloc) {
+ return true;
+ }
+
+ if (kv_size_alloc >= kv_size_max) {
+ return false;
+ }
+
+ // Grow geometrically (next power-of-two) and align up to n_pad. The
+ // pow2 ramp avoids re-growing every few tokens for long prompts; n_pad
+ // alignment preserves the constructor's GGML_ASSERT(kv_size % n_pad == 0)
+ // invariant for find_slot.
+ uint32_t new_cap = std::max<uint32_t>(kv_size_alloc, n_pad);
+ while (new_cap < cells_needed && new_cap < kv_size_max) {
+ new_cap *= 2;
}
+ new_cap = ((new_cap + n_pad - 1) / n_pad) * n_pad;
+ new_cap = std::min(new_cap, kv_size_max);
- cleared = true;
+ if (new_cap == kv_size_alloc) {
+ return cells_needed <= kv_size_alloc;
+ }
+
+ LLAMA_LOG_INFO("%s: grow soft cap %u -> %u (need %u, max %u)\n",
+ __func__, kv_size_alloc, new_cap, cells_needed, kv_size_max);
+
+ kv_size_alloc = new_cap;
+ return cells_needed <= kv_size_alloc;
}
void llama_kv_cache::clear(bool data) {
@@ -371,14 +453,16 @@ void llama_kv_cache::clear(bool data) {
}
if (data) {
- // opencoti F4 M3 Phase 1: if the deferred zero-fill has not happened
- // yet there is nothing to clear β v_cells.reset() above already
- // satisfies the logical "data cleared" state. Don't commit RSS just
- // to zero pages that have never been written.
- if (cleared) {
- for (auto & [_, buf] : ctxs_bufs) {
- ggml_backend_buffer_clear(buf.get(), 0);
- }
+ // opencoti F4 M3 Phase 1+2: if no zero-fill has happened yet there
+ // is nothing to clear β v_cells.reset() above already satisfies the
+ // logical "data cleared" state. Don't commit RSS just to zero pages
+ // that have never been written. When a partial clear is already on
+ // the books, re-issue it against the current soft cap so any cells
+ // touched between the first ensure_cleared and now are re-zeroed.
+ if (n_cells_cleared > 0) {
+ const uint32_t b = n_cells_cleared;
+ n_cells_cleared = 0;
+ ensure_cleared(b);
}
}
}
@@ -743,10 +827,28 @@ llama_kv_cache::slot_info_vec_t llama_kv_cache::prepare(const std::vector<llama_
for (const auto & ubatch : ubatches) {
// only find a suitable slot for the ubatch. don't modify the cells yet
- const auto sinfo_new = find_slot(ubatch, false);
+ auto sinfo_new = find_slot(ubatch, false);
if (sinfo_new.empty()) {
- success = false;
- break;
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Phase 2 keeps the soft alloc cap (kv_size_alloc) below kv_size_max
+ // until a real ubatch needs more cells than the cap exposes. When
+ // find_slot fails, try to grow the cap to cover what we already
+ // know is used plus the ubatch's token count, then retry. The grow
+ // is bounded by kv_size_max; if we are already at the hard cap the
+ // failure is genuine and we propagate it up.
+ uint32_t used_max = 0;
+ for (uint32_t st = 0; st < n_stream; ++st) {
+ used_max = std::max(used_max, v_cells[st].used_max_p1());
+ }
+ const uint32_t want = used_max + ubatch.n_tokens;
+ if (ensure_capacity(want)) {
+ ensure_cleared();
+ sinfo_new = find_slot(ubatch, false);
+ }
+ if (sinfo_new.empty()) {
+ success = false;
+ break;
+ }
}
// remember the position that we found
@@ -981,14 +1083,20 @@ llama_kv_cache::slot_info llama_kv_cache::find_slot(const llama_ubatch & ubatch,
uint32_t head_cur = v_heads[seq_to_stream[seq_id]];
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // The Phase-1 search swept the full hard cap (cells.size() == kv_size_max).
+ // Phase 2 caps it at the current soft alloc; prepare() grows the cap and
+ // retries when the soft cap is too small for the requested ubatch.
+ const uint32_t cap = kv_size_alloc;
+
// if we have enough unused cells before the current head ->
// better to start searching from the beginning of the cache, hoping to fill it
if (head_cur > cells.get_used() + 2*n_tokens) {
head_cur = 0;
}
- if (n_tokens > cells.size()) {
- LLAMA_LOG_ERROR("%s: n_tokens = %d > size = %u\n", __func__, n_tokens, cells.size());
+ if (n_tokens > cap) {
+ // not necessarily fatal β prepare() may grow the cap and retry
return { };
}
@@ -999,8 +1107,8 @@ llama_kv_cache::slot_info llama_kv_cache::find_slot(const llama_ubatch & ubatch,
const uint32_t n_test = cont ? n_tokens : 1;
while (true) {
- if (head_cur + n_test > cells.size()) {
- n_tested += cells.size() - head_cur;
+ if (head_cur + n_test > cap) {
+ n_tested += cap - head_cur;
head_cur = 0;
continue;
}
@@ -1058,7 +1166,7 @@ llama_kv_cache::slot_info llama_kv_cache::find_slot(const llama_ubatch & ubatch,
res.idxs[s].clear();
}
- if (n_tested >= cells.size()) {
+ if (n_tested >= cap) {
//LLAMA_LOG_ERROR("%s: failed to find a slot for %d tokens\n", __func__, n_tokens);
return { };
}
@@ -1961,10 +2069,13 @@ void llama_kv_cache::state_read(llama_io_read_i & io, llama_seq_id seq_id, llama
GGML_ASSERT(seq_id == -1 || (seq_id >= 0 && (size_t) seq_id < seq_to_stream.size()));
- // opencoti F4 M3 Phase 1 β see docs/decisions/0001-lazy-slot-context.md
- // state_read writes the saved KV blob into our tensors, which requires
- // backing storage. Force the deferred allocation before we read.
- ensure_cleared();
+ // opencoti F4 M3 Phase 1+2 β see docs/decisions/0001-lazy-slot-context.md
+ // state_read writes the saved KV blob into our tensors at arbitrary
+ // offsets, so we cannot rely on the soft cap. Force a full-range
+ // zero-fill before we read β this commits all KV pages but only on
+ // the explicit deserialize path.
+ ensure_capacity(kv_size_max);
+ ensure_cleared(kv_size_max);
uint32_t n_stream_cur;
io.read(&n_stream_cur, sizeof(n_stream_cur));
diff --git a/llama.cpp/src/llama-kv-cache.h b/llama.cpp/src/llama-kv-cache.h
--- a/llama.cpp/src/llama-kv-cache.h
+++ b/llama.cpp/src/llama-kv-cache.h
@@ -101,6 +101,11 @@ public:
bool offload,
bool unified,
uint32_t kv_size,
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Initial allocation cap. 0 means the same as kv_size
+ // (Phase 1 behavior). Otherwise clamped to [n_pad, kv_size]
+ // and rounded up to the next n_pad multiple.
+ uint32_t kv_size_initial,
uint32_t n_seq_max,
uint32_t n_pad,
uint32_t n_swa,
@@ -152,15 +157,28 @@ public:
bool get_has_shift() const;
- // opencoti F4 M3 Phase 1 β see docs/decisions/0001-lazy-slot-context.md
- // Idempotent: zeroes the KV-cache backend buffers on first call.
- // The buffers are malloc'd eagerly at construction (so the scheduler
- // can resolve tensor->buffer at reserve time) but NOT cleared β on
- // Linux, malloc reserves address space without committing pages, so
- // RSS stays low until something writes. This method does the memset
- // that commits the pages. Called from inference entry points
- // (init_batch, prepare, update, state_read).
- void ensure_cleared();
+ // opencoti F4 M3 Phase 1+2 β see docs/decisions/0001-lazy-slot-context.md
+ // Idempotent. Zeroes cells [n_cells_cleared, up_to_cells) of every KV
+ // layer's K and V tensors across every stream. The buffers are malloc'd
+ // eagerly at construction (so the scheduler can resolve tensor->buffer
+ // at reserve time) but their pages are not committed to RSS until they
+ // are first written. Partial memset (via ggml_backend_tensor_memset)
+ // commits only the requested range β that is the Phase 2 RSS win.
+ //
+ // up_to_cells == 0 means "use the current soft cap kv_size_alloc"
+ // (the normal inference path). Pass kv_size_max for paths that need
+ // full coverage (state_read).
+ void ensure_cleared(uint32_t up_to_cells = 0);
+
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Grow the soft allocation cap. Returns true iff the cap covers
+ // cells_needed after the grow. Bound by kv_size_max. Idempotent at
+ // the cap.
+ bool ensure_capacity(uint32_t cells_needed);
+
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // The current soft cap (cells visible to find_slot).
+ uint32_t get_kv_size_alloc() const { return kv_size_alloc; }
ggml_type type_k() const;
ggml_type type_v() const;
@@ -266,12 +284,22 @@ private:
// ggml contexts for the KV cache along with the allocated backend buffers:
std::vector<std::pair<ggml_context_ptr, ggml_backend_buffer_ptr>> ctxs_bufs;
- // opencoti F4 M3 Phase 1 β see docs/decisions/0001-lazy-slot-context.md
- // True once the backend buffers in ctxs_bufs have been zero-filled.
- // The constructor sets this true for no_alloc models (whose dummy
- // buffers are size 0 β clearing is a no-op) and false otherwise. The
- // first ensure_cleared() call zeroes the buffers and flips this true.
- bool cleared = false;
+ // opencoti F4 M3 Phase 1+2 β see docs/decisions/0001-lazy-slot-context.md
+ // Cells in the half-open range [0, n_cells_cleared) are guaranteed to
+ // be zeroed across every layer's K and V tensors and every stream.
+ // ensure_cleared(up_to) grows this monotonically up to kv_size_alloc
+ // (or kv_size_max for the explicit-full call). The constructor sets
+ // this to kv_size_max for no_alloc models (whose dummy buffers are
+ // size 0 β clearing is a no-op) and 0 otherwise.
+ uint32_t n_cells_cleared = 0;
+
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // kv_size_max: hard cap on cells (set from constructor's kv_size).
+ // kv_size_alloc: current soft cap. find_slot only scans cells [0, kv_size_alloc).
+ // Grows via ensure_capacity, never shrinks (Phase 2; Phase 3
+ // adds shrink-on-idle).
+ uint32_t kv_size_max = 0;
+ uint32_t kv_size_alloc = 0;
// the current index from where we start searching for a free slot in the ring buffer of KV cells (see find_slot())
// note: this is not part of the KV state and it's only used to speed-up the find_slot() method
diff --git a/llama.cpp/src/llama-memory-hybrid-iswa.cpp b/llama.cpp/src/llama-memory-hybrid-iswa.cpp
--- a/llama.cpp/src/llama-memory-hybrid-iswa.cpp
+++ b/llama.cpp/src/llama-memory-hybrid-iswa.cpp
@@ -40,6 +40,10 @@ llama_memory_hybrid_iswa::llama_memory_hybrid_iswa(
swa_full,
unified,
kv_size,
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Hybrid-iswa does not yet thread slot_initial_ctx; pass 0 for
+ // Phase 1 backward-compat (full eager allocation).
+ 0,
n_seq_max,
n_ubatch,
n_pad,
diff --git a/llama.cpp/src/llama-memory-hybrid.cpp b/llama.cpp/src/llama-memory-hybrid.cpp
--- a/llama.cpp/src/llama-memory-hybrid.cpp
+++ b/llama.cpp/src/llama-memory-hybrid.cpp
@@ -39,6 +39,10 @@ llama_memory_hybrid::llama_memory_hybrid(
offload,
unified,
kv_size,
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ // Hybrid memory does not yet thread slot_initial_ctx; pass 0 for
+ // Phase 1 backward-compat (full eager allocation).
+ 0,
n_seq_max,
n_pad,
n_swa,
diff --git a/llama.cpp/src/llama-model.cpp b/llama.cpp/src/llama-model.cpp
--- a/llama.cpp/src/llama-model.cpp
+++ b/llama.cpp/src/llama-model.cpp
@@ -2073,6 +2073,8 @@ llama_memory_i * llama_model::create_memory(const llama_memory_params & params,
params.swa_full,
cparams.kv_unified,
cparams.n_ctx_seq,
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ cparams.slot_initial_ctx,
cparams.n_seq_max,
cparams.n_ubatch,
1,
@@ -2089,6 +2091,8 @@ llama_memory_i * llama_model::create_memory(const llama_memory_params & params,
cparams.offload_kqv,
cparams.kv_unified,
cparams.n_ctx_seq,
+ // opencoti F4 M3 Phase 2 β see docs/decisions/0001-lazy-slot-context.md
+ cparams.slot_initial_ctx,
cparams.n_seq_max,
1,
hparams.n_swa,
|