File size: 31,607 Bytes
d9f0f95 756ec9e d9f0f95 4a740a4 d9f0f95 756ec9e d9f0f95 756ec9e 6e4de51 756ec9e 6e4de51 756ec9e d9f0f95 3b7efe6 d9f0f95 3b7efe6 d9f0f95 3b7efe6 d9f0f95 2b8ba64 4a740a4 2b8ba64 d9f0f95 3b7efe6 d9f0f95 e6bd945 d9f0f95 b2f2b76 d9f0f95 e6bd945 d9f0f95 3b7efe6 d9f0f95 e6bd945 4a740a4 e6bd945 d9f0f95 2b8ba64 d9f0f95 d051184 d9f0f95 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 | # opencoti-llamafile β usage guide
How this engine diverges from upstream
[Mozilla-Ocho llamafile](https://github.com/Mozilla-Ocho/llamafile),
what the added features are, how each is gated, its knobs and
defaults, its limitations, and which features are meant to be used
together.
Audience: anyone running the packaged
`opencoti-llamafile-<ver>-<tag>-<arch>.llamafile` artifact as a local
inference server. Deep-dive design docs live in
[`docs/features/`](features/), measured evaluations in
[`docs/evaluations/`](evaluations/).
---
## Supported / target model families β read this first
opencoti-llamafile loads **any GGUF that upstream llama.cpp
loads** β that part is inherited unchanged. But the opencoti feature
set (KV tiers, rolling-KV, DCA, MTP, sparse attention, RYS) is
developed, tuned, and correctness-gated on **two model families**, in
a deliberate primary/secondary split:
### Gemma-4 β PRIMARY target
| Model | Kind | Notes |
|---|---|---|
| Gemma-4 26B-A4B-128e ("A4B") | MoE | the flagship serving target; MTP-validated with its native gemma4-assistant drafter |
| Gemma-4 12B / 31B | dense | full feature validation incl. RYS + DCA |
| Gemma-4 E2B / E4B | elastic (E-series) | shared-KV elastic layers supported; MTP drafters available |
Gemma-4 is what the engine is *for*: its unusual head dims (256 and
512), the iSWA sliding/global dual KV cache, and the per-size
[gemma4-assistant MTP drafters](https://huggingface.co/ManniX-ITA)
all have dedicated kernels and graph paths here that upstream lacks
or handles slowly. `--spec-type draft-assistant`, D256/D512 FA-VEC +
scalar-MMA decode, iSWA-aware rolling-KV/SharedKVPool/DCA wiring are
all Gemma-4-first features.
### Qwen β SECONDARY target / verification family
| Model | Kind | Notes |
|---|---|---|
| Qwen3.5 / Qwen3.6 (e.g. 35B-A3B) | dense / MoE / hybrid (gated-delta-net) | NextN self-spec MTP (`--spec-type draft-mtp`, no external drafter needed) |
| Qwen2.5-14B-1M | dense, native-1M | the long-context/DCA validation vehicle |
Qwen is the standard-architecture (head_dim 128) counterweight:
every feature that ships is verified on it too, and it carries one
feature Gemma doesn't β **NextN self-speculation** (the model's own
MTP head drafts; fused multi-step, at/above upstream parity).
### Everything else
Other architectures run with upstream behavior and safe fallbacks,
but opencoti features are **unvalidated** there, and some are
arch-gated: MTP needs NextN tensors (Qwen-style) or a
gemma4-assistant drafter; RYS `--repeat-layers` supports the
qwen2/qwen3(+MoE)/qwen3.5/qwen3next/gemma-4 forward loops; DCA is
validated on Gemma-4 and Qwen2.5-1M. Quality gates (KLD,
RULER-niah) were run on the two families above β re-gate before
trusting aggressive KV tiers on anything else.
---
## 1. Relationship to upstream llamafile
opencoti-llamafile is **upstream llamafile 0.10.3 plus an additive
patch series** (`patches/` in the HF repo,
`vendors/patches/llamafile/` in the git repo β ~81 patches, numbered
`0006`β`0135`). Three properties are contractual:
1. **Off means off.** Every opencoti feature is opt-in behind a flag,
env var, or per-request JSON field. With no opencoti flags set, the
engine's compute path is **byte-identical to upstream** β this is a
regression gate on every patch, not an aspiration.
2. **Lossless by proof, not vibes.** Features that touch the forward
pass are gated by logit-equivalence / KLD / RULER-retrieval
against vanilla, never by "the output looks fine". Speculative
decode is verified-lossless (the output *is* the target model's).
3. **Single file, zero dependencies.** The artifact is a Cosmopolitan
APE: one file runs on Linux/macOS/Windows/BSD, x86_64 and aarch64.
In the full x86_64 artifact the CUDA backend (`ggml-cuda.so`) is
embedded and self-extracts to `~/.llamafile/v/<ver>/` on first GPU
run; the `-win` variant ships without it (see Β§1.1). TCQ codebooks
and quantization tables are compiled in. No installer, no downloads.
What upstream gives you is unchanged: the server API
(`/completion`, `/v1/chat/completions`, `/props`, `/slots`, β¦), GGUF
loading, sampling, chat templates. opencoti adds serving-efficiency
machinery on top, aimed at **multi-session agentic serving on a fixed
VRAM budget**: more concurrent sessions per card, longer usable
context, faster decode.
```bash
chmod +x opencoti-llamafile-0.10.3-c3-x86_64.llamafile
sh ./opencoti-llamafile-0.10.3-c3-x86_64.llamafile --server --port 8080 \
-m model.gguf -ngl 99 --flash-attn on
# --version β opencoti-0.10.3-c3 ; without --server you get the chat CLI
```
> **Note (Linux):** launch via `sh ./file.llamafile` if your kernel
> lacks binfmt_misc APE registration.
### 1.1 Artifact variants β which file to download
| Artifact | Size | Runs on | GPU story |
|---|---:|---|---|
| `β¦-x86_64.llamafile` | ~4.8 GB | Linux/macOS/BSD, x86_64 + aarch64 | CUDA DSO **embedded** (x86_64-linux, sm_75/80/86/89/90/120f); self-extracts on first `-ngl` run |
| `β¦-win-x86_64.llamafile.exe` | ~55 MB | **Windows** + Linux/macOS/BSD, x86_64 + aarch64 | none embedded β CPU works everywhere out of the box; GPU via side-load (below) |
| `β¦-aarch64.llamafile` | ~1.9 GB | same platforms | CUDA sbsa DSO **embedded** (aarch64-linux, sm_110f DGX Spark GB10 / Jetson Thor + sm_121a); GPU out of the box on aarch64-linux CUDA hosts |
The host binary inside every variant is byte-for-byte the same APE
(same patches, same features); they differ only in the embedded GPU
backend. The `-win` variant exists because **Windows refuses to run
executables larger than 4 GB**, so the full artifact β 4.8 GB with
the CUDA DSO inside β cannot ship for Windows; the bare APE renamed
to `.exe` (Windows needs the extension) can.
**GPU with the `-win`/bare variant:**
- **Windows:** install the NVIDIA CUDA Toolkit + MSVC and run with
`-ngl 99`; llamafile compiles a native `ggml-cuda.dll` on the fly
into `%USERPROFILE%\.llamafile\` on first use (upstream llamafile
mechanism, unchanged).
- **Linux x86_64:** download the published standalone DSO
(`dso/<ver>-<tag>/ggml-cuda-x86_64.so` in the HF repo β same bytes
as the one embedded in the full artifact, sha in
`releases/<tag>/SHA256SUMS`) and place it where the engine looks
before attempting extraction:
```bash
mkdir -p ~/.llamafile/v/0.10.3/
cp ggml-cuda-x86_64.so ~/.llamafile/v/0.10.3/ggml-cuda.so
```
- **Linux aarch64 (DGX Spark / GB10 class):** same mechanism with the
sbsa DSO β `dso/<ver>-<tag>/ggml-cuda-sbsa-aarch64.so` β
`~/.llamafile/v/0.10.3/ggml-cuda.so`. Or just download the
`β¦-aarch64.llamafile` artifact, which has it embedded.
- **No CUDA available:** nothing to do β it falls back to CPU
inference automatically.
---
## 2. Feature map β what exists and how it's gated
| Feature | Default | Turn on with | Class |
|---|---|---|---|
| Session-keyed KV reuse | off (per request) | `session_id` JSON field | latency |
| ReST-KV retention eviction | **off** | `--rest-kv-eviction` | quality-under-overflow |
| KV quantization (scalar) | f16 | `-ctk` / `-ctv` | capacity |
| TurboQuant / TCQ KV tiers | off | `-ctk`/`-ctv turbo*` | capacity |
| Auto KV-tier policy | **off** | `OPENCOTI_KV_AUTO_TIER=1` | capacity (policy) |
| PolyKV pool (SharedKVPool) | off (per request) | `shared_pool_slot` JSON field | multi-agent capacity |
| Rolling-KV window / spill | **auto** (engages only under pressure) | `--vram-target`, `--kv-residency-mode` | capacity |
| Mixed-KV spilled tail | off | `-ctkt` / `-ctvt` | capacity |
| DCA long-context | **off** | `--dca on` | context extension |
| Sparse attention (block-selector) | **off** | `--sparse-attn on` | long-ctx decode speed |
| Sparse-V | auto on iSWA+quant-V, else off | `TURBO_SPARSE_V_TAU` | decode speed |
| MTP speculative decode | **off** | `--spec-type` + drafter | decode speed |
| RYS layer duplication | **off** | `--repeat-layers` | quality |
| RYS probe | off | `--rys-probe` | tooling |
| Lazy slot context | off | `--slot-initial-ctx`, `--slot-shrink-idle-ms` | embedder memory |
| Introspection API | **always on** | `GET /props`, `GET /slots` | observability |
Every boot flag also has an env twin
(`OPENCOTI_LLAMAFILE_<SNAKE_CASE>` for adapter-typed fields,
`LLAMA_ARG_*` for llama.cpp-registered ones).
---
## 3. KV capacity stack β PolyKV
**PolyKV** is the umbrella name for this whole stack: the compressed
shared KV pool. Concretely it is the KV quantization tiers of Β§3.1
plus the multi-agent SharedKVPool of Β§3.4, stacking with the
auto-tier policy (Β§3.2) and the rolling-KV window (Β§3.3). If you
arrived here looking for "PolyKV" from an announcement: Β§3.4 is the
shared-prefix pool itself; Β§3.1 is what the pooled cells are made of.
These four features share one goal β **fit more context / more
sessions in fixed VRAM** β and are designed to stack. Recommended
order of adoption: scalar quant β auto-tier β rolling-KV β
turbo tiers β SharedKVPool.
### 3.1 KV quantization: scalar types + TurboQuant/TCQ tiers (PolyKV M6)
The KV cache type is set per-tensor-half: `-ctk <type>` (keys) and
`-ctv <type>` (values), independently β **asymmetric pairs are
first-class** (e.g. `-ctk q8_0 -ctv q4_0`).
Supported types: `f16`, `bf16`, `q8_0`, `q6_0`, `q5_1`, `q5_0`,
`q4_0` (scalar), `turbo2`, `turbo3`, `turbo4`, `turbo8`
(TurboQuant, MSE-optimal with Walsh-Hadamard rotation + InnerQ),
`turbo2_tcq`, `turbo3_tcq` (trellis-coded, Viterbi-encoded).
**Which to pick (measured):**
- **8-bit / 4-bit: use `q8_0` / `q4_0`.** The native scalar types
dominate turbo8/turbo4 at equal width β turbo earns nothing there.
- **`-ctk q8_0 -ctv q4_0`** is the workhorse asymmetric pair: keys
keep 8-bit fidelity (attention logits are K-sensitive), values
take the compression.
- **3 bits and below is TurboQuant territory:** `turbo3`
Pareto-beats q4_0 (90 vs 129 MiB KV at equal quality, teacher-forced
TV 0.0067 vs 0.0094); `turbo2` is the smallest logit-equivalent KV
that exists (~2 bit) β the 256k-context play. TCQ variants trade
encode cost for a further fidelity step at the same width.
- All shipped tiers pass logit-equivalence gates; decode runs the
quantized data **in-register** in the flash-attention kernel (no
f16 materialization) for turbo2/3/4 and TCQ.
Limitations: turbo8 uses a materialize fallback (not fused);
at Gemma-4's head_dim 512 only turbo2/turbo3 have fused D=512
instances; prefill on very long prompts uses a hybrid path
automatically. Quality validation on Gemma franken-merges must use
retrieval (niah), not perplexity.
### 3.2 Auto KV-tier (`OPENCOTI_KV_AUTO_TIER=1`)
Boot policy: pick the **least-compressing scalar pair that keeps the
whole KV resident** in the VRAM budget; if even that spills, the T\*
model decides between "small f16 spill" and "quantize one tier down"
by predicted tokens/s drop.
Knobs (env): `OPENCOTI_KV_AUTO_TIER=1` (master),
`OPENCOTI_KV_TSTAR_DROP` (target drop, default 20%),
`OPENCOTI_KV_TSTAR_MAX_SPILL_MIB` (default 800),
`OPENCOTI_KV_AUTO_TIER_TAIL=1` (also auto-pick a q4_0 spilled tail).
Explicit `-ctv` disables auto entirely; explicit `-ctk` holds K and
walks only V. Dense full-attention models only (iSWA models keep f16).
Read back what it decided: `GET /props β .opencoti.kv.effective` β
the *configured vs effective* split exists exactly because auto-tier
may override you.
### 3.3 Rolling-KV window (residency / spill)
"KV doesn't have to fit." Each layer keeps a device-resident window
of recent positions; the tail `[0 β¦ window_start)` lives in pinned
host RAM and is streamed through the attention kernel per-tile,
merged exactly via online-softmax (LSE). When everything fits, every
layer is `GPU_RESIDENT` and the path is byte-identical to vanilla β
the feature only *engages* under memory pressure.
Flags: `--vram-target <MiB>` (budget cap; `0` = all free VRAM minus
reserve), `--kv-residency-mode {auto,head,window}` (default `auto`;
leave it), `-ctkt` / `-ctvt` (distinct, more-compressed types for the
spilled tail β "mixed-KV": f16 recent window β q4_0 tail).
Performance model (RTX 3090, PCIe ~6.5 GB/s): spill decode sits at
the PCIe floor, `t(token) β fixed + tail_bytes / link_bw` β linear,
no cliff. On a fast-link host (RTX 6000, ~50 GB/s) window-mode spill
is genuinely usable; on consumer PCIe it's a last resort β prefer
quantizing (that's what auto-tier automates).
Limitations: while a window is spilled, context-shift and
prompt-cache-reuse are guarded off (requests bounded at `n_ctx`);
the compute-buffer reserve for long contexts is measured
automatically at boot (two-pass reserve β no knob).
### 3.4 PolyKV SharedKVPool (multi-agent shared prefix)
N agents attending **one physical copy** of a common prefix (system
prompt + tool defs). Per-request JSON, no CLI flag:
```jsonc
{ "shared_pool_slot": 0, "shared_prefix_n_tokens": 4096, β¦ }
```
Server must run `--kv-unified --no-cache-idle-slots` (the latter is
mandatory β the default idle-slot save/clear would evict the pooled
prefix).
**What the pool speeds up β measured** (Gemma-4-26B-A4B Q4_K_M,
RTX 3090, q8/q8 unified KV, Pβ5k-token shared prefix, greedy
fixed-length decode, 8 concurrent sessions unless noted):
| axis | naive (N private copies) | shared pool | gain |
|---|---|---|---|
| KV cells (N=8) | ~8Β·P | P + suffixes | **6.9Γ** (~306 vs ~9 agents on a fixed buffer) |
| prefill, 8 sessions joining | 22.8 s | 4.6 s | **~5Γ** (prefix enters KV once per pool) |
| steady-state batched decode (N=8) | ~190 tok/s | 217 tok/s | **+14%** (8 queries read one physical prefix β L2 reuse, smaller cell span) |
| multi-turn re-query (N=8) | 99 tok/s | 225 tok/s | **2.3Γ** (see note) |
| iso-speed capacity | 8 sessions @ 24.0 tok/s each | β₯12 sessions @ β₯26.9 tok/s each | **β₯1.5Γ** sessions (crossover not reached at N=12; aggregate 315 tok/s) |
The multi-turn row is iSWA-specific and easy to miss: a private slot
that has decoded past its prompt cannot partially rewind (upstream
SWA-checkpoint semantics, llama.cpp PR #13194), so re-querying it
pays a checkpoint restore or a full re-prefill every turn. The pool
slot never decodes, so its prefix never slides β every re-attach is
free. Note the capacity row is about per-session speed, not just
aggregate: 12 pooled sessions each decode faster than 8 private ones.
**Cross-architecture results** (RTX 6000 96GB, Pβ5073, GEN=256,
A/B/A naive/shared/naive): the pool is validated on all three
attention architectures, and the memory axis is
architecture-independent (~6.8β6.9Γ at N=8 β it counts cells, not
attention math).
| axis | Qwen2.5-14B-1M Q8_0 (pure full attention) | Qwen3.6-27B-Omnimerge-v4 Q4_K_M (hybrid GDN + NextN MTP n=3) |
|---|---|---|
| KV cells (N=8) | **6.78Γ** (~406 vs ~9 agents on a fixed 8192-cell buffer) | **6.89Γ** (~304 vs ~9 agents) |
| batched decode N=8 | 358 β 403 tok/s (**+12.5%**) | 106 β 123 tok/s (**+15.6%**) |
| batched decode N=24 | 412 β 742 tok/s (**+80%**; 17.2 β 30.9 tok/s per session) | 90 β 125 tok/s (**+39.7%**; 3.7 β 5.2 per session) |
| shared-only sweep N=32/48/64 | 813 / 871 / 865 tok/s (plateau ~870 near N=48) | 126 / 124 / 124 tok/s (saturates by Nβ24β32) |
The shared-vs-naive decode gain **grows with N** on both. On
hybrid/recurrent models (delta-net, mamba) the *absolute* aggregate
saturates much earlier than on pure attention β the recurrent layers
batch worse β so there the pool buys **concurrency capacity and
memory**, not aggregate throughput past Nβ24.
**How to prefill the pool β use the common-prefix token array, not
the document text.** Tokenizers merge across the document/suffix
boundary (on the Qwen tokenizer the last prefix token fuses with the
suffix start), so `tok(DOC)` can be one token longer than the common
prefix the agents actually share β and a pool that is even one token
longer than `shared_prefix_n_tokens` cannot be shared exactly. The
correct client sequence:
```jsonc
// 1. tokenize the FULL agent prompts and compute
// P = min over agents of commonPrefixLen(tok(DOC), tok(DOC+suffix_i))
// 2. prefill the pool slot with the token array itself (llama.cpp
// /completion accepts token arrays) β pool state == P on ANY tokenizer:
{ "prompt": [/* tok(DOC+suffix_1)[:P] */], "id_slot": 0,
"cache_prompt": true, "n_predict": 1 }
// 3. agents attach with prompts STRICTLY longer than P:
{ "prompt": "<DOC + private suffix>", "id_slot": 1,
"shared_pool_slot": 0, "shared_prefix_n_tokens": P, β¦ }
```
On attention-only models a text prefill happens to work (the ranged
cell copy tolerates the extra token); on hybrid/recurrent targets it
silently disables every share β see the gotcha below. The token-array
prefill is correct everywhere.
**Hybrid/recurrent gotcha (GDN / mamba / `qwen35moe`-class models).**
A recurrent cache has one rolling state per sequence, not per-position
cells, so a pool share is only possible as an **exact full-state**
share. The server enforces this: the share engages only when
`shared_prefix_n_tokens == pool state length` *and* the request prompt
is strictly longer than the shared prefix; anything else logs
`poly-kv-pool: hybrid/recurrent target needs exact full-state share β¦
skipping share, full reprocess (bug-2203)` and falls back to a full
(correct, slower) reprocess. If you see zero speedup on a hybrid
model β or mass HTTP 500s at high N because N unshared full prompt
copies overflow the unified KV β grep the server log for that WARN:
it almost always means the pool was prefilled with text instead of
the token array. (Older builds crashed outright here β
`failed to remove sequence N with p0=β¦` β fixed by patch `0135`.)
Sizing note: the pooled prefix pins P cells in **both** iSWA caches
(global + SWA) for the pool's lifetime. Budget `-c` for pool prefix
+ N session windows + generation headroom, or long-running sessions
can exhaust slot allocation mid-round.
Composes with KV quantization (the pool holds quantized cells) and
with session KV-reuse. The pool is read-only for consumers; each
agent's divergent suffix is private.
**Tiering is pinned per-pool, never per-session.** The K/V tiers β
including the mixed-KV recent-window β compressed-tail pair β are
properties of the boot-allocated cache tensors, chosen once at boot
(by you or by auto-tier) before any session exists. The window/tail
boundary is a per-layer residency budget over the *physical cell
axis*, so a prefix cell is in the resident window or evicted (and
quantized exactly once, on eviction) for **all** sequences
simultaneously. Sharing itself is not copy-on-write: a sharer joins
the prefix by adding its sequence bit to the existing cells, and a
diverging session just appends private suffix cells β there is no
per-session copy that could be re-quantized, and no way for two
sessions to see the same prefix at different tiers. The flip side:
you cannot give one session a higher-precision read of a shared
prefix than another; that would require forking the prefix into a
private copy, which is exactly the O(N) memory cost the pool exists
to avoid.
---
## 4. Long context
### 4.1 DCA β Dual Chunk Attention (training-free context extension)
Splits attention into intra-chunk / successive / inter-chunk position
regimes and merges them exactly by LSE, so a model trained at
`n_ctx_train` serves multiples of it **without retraining**.
Flags: `--dca on` (default **off**), `--dca-chunk-size N` (default
derives from the model's training context; explicit 8192 is the
validated recipe), `--dca-yarn-factor F` (default 1.0; measured
neutral for retrieval β leave it). Serve beyond the GGUF's declared
context with
`--override-kv <arch>.context_length=int:1048576`.
Validated recipe (Gemma-4-A4B, n_ctx_train 256k):
```bash
--dca on --dca-chunk-size 8192 -fa on --parallel 1 \
--override-kv gemma4.context_length=int:1048576
```
Measured retrieval (RULER-VT, n=50): **256k 0.964 Β· 512k 0.996 Β·
768k 0.984 Β· 1M 0.916** β a gentle β7 pp at 4Γ native, no cliff.
Counter-proof on Qwen3-8B (native 41k): plain attention collapses at
128k (PPL 19.2) while DCA holds PPL 7.3.
Works on Gemma-4 (its 5 global layers; SWA layers untouched) and
Qwen2.5/3/3.5 (all layers). Composes with quantized KV (scalar pairs
all pass; q8-DCA decode costs ~2Γ vs f16-DCA), sparse attention, and
rolling-KV.
Limitations: DCA caches K un-rope'd β **launch-time toggle only** (a
server booted DCA-on can't switch off per request); expect
approximation, not identity, past one chunk. On models that are already *native* long-context
(e.g. Qwen2.5-1M), DCA can only approximate down β don't use it
there.
### 4.2 Sparse attention (Quest block-selector) + sparse-V
Two independent decode-bandwidth levers:
- **Block-selector** (`--sparse-attn on`): per-block min/max key
bounds give an upper bound on each block's attention mass; decode
visits only the top-K blocks (+ sinks + recent). Flags:
`--sparse-attn-block-size` (128), `--sparse-attn-topk` (default 0 =
visit **all** blocks, i.e. no skipping; pass `auto` for adaptive
max(64, n_blocks/4), or an explicit block count),
`--sparse-attn-recent`, `--sparse-attn-sink` (1),
`--sparse-attn-refresh` (8 β re-select every N decode steps),
`--sparse-attn-mode` (0). Default **off**.
- **Sparse-V**: skips V-dequant for negligible-weight positions
inside visited blocks. **Self-configuring**: on iSWA models with
quantized V it auto-sets Ο=0.05; elsewhere it stays off. Manual
override: `TURBO_SPARSE_V_TAU=<float>`.
When to use: **long context on quantized KV.** The win grows with
context (selectivity 0.91@16k β 0.99@40k and climbing) and lives on
quantized KV: q8_0 β sparse at 50% coverage measured **1.34Γ decode
at niah 100**. Both levers stack (1.31Γ combined measured).
When *not* to use: short contexts or f16 KV on mid-size models β the
decode isn't KV-bandwidth-bound there and the selector overhead can
make it *slower* than dense. Ο values don't transfer across models;
retune if you override manually.
---
## 5. Decode speed β MTP speculative decoding
Lossless speculative decode; the emitted text is the target model's
own (verified). Two flavours, chosen by `--spec-type`:
### 5.1 `--spec-type draft-assistant` (external drafter β Gemma-4)
A small `gemma4-assistant` drafter GGUF rides the target's
embeddings:
```bash
--spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \
-ngld 99 --spec-draft-n-max 2
```
`--mtp-head` (alias `-md`) names the drafter; **`-ngld 99` matters**
(a CPU-resident draft head erases the win). Drafters for
A4B/12B/27B/E2B/E4B are published per-size. Setting `mtpHead` in the
TS adapter auto-derives the rest.
### 5.2 `--spec-type draft-mtp` (NextN self-spec β Qwen)
Qwen 3.5/3.6 GGUFs that embed a NextN/MTP head self-speculate β no
second file:
```bash
--spec-type draft-mtp --spec-draft-n-max 3
```
Runs per-slot under `--parallel` (multi-session capable).
**Measured (RTX 3090 + upstream-parity campaign):** A4B assistant
decode beats upstream llama.cpp b9859 at every depth (+6.6/+12.9/+7.6%
at n_max 1/2/3); combined with `turbo3_tcq` KV it reaches **~89 tok/s
vs 52.9 plain (+69%)**. Qwen-35B NextN sits at parity with upstream.
Recommended depth: `--spec-draft-n-max 2β3` (A4B), `3` (Qwen NextN).
Notes/limits: acceptance dips a few pp at depth β₯2 (chained-draft
numerics β expected); with `--parallel > 1`, assistant-MTP requires
`--kv-unified` (auto-forced with a boot warning). Composes with
turbo/TCQ KV tiers (its biggest lever), DCA, and quantized KV. Watch
live acceptance per slot via `GET /slots` (Β§7).
---
## 6. Quality β RYS layer duplication
`--repeat-layers` re-runs a contiguous block of **middle** layers,
weight-shared: zero extra parameter VRAM, no new GGUF, quant-agnostic.
You pay in KV cache and tokens/s proportional to the extra effective
layers; you buy quality-per-token.
```bash
--repeat-layers 33,34 # +1 layer (RYS-S)
--repeat-layers 26-34 # +8 layers ([26,34) half-open, RYS-XL)
--repeat-layers 8-12;20-24 # disjoint blocks
```
Rules that matter:
- **Middle layers only.** Duplicating first/last layers reliably
produces incoherent output on merge-fragile models β this is a
model property, not an engine bug; the engine prints a boot
advisory when a plan touches the boundary band.
- Absent flag = identity = byte-identical to stock.
- Composes with the full stack: quantized KV, DCA, rolling-KV
window/spill, sparse-attn (the residency/DCA/sparse sizing paths
are effective-plan-aware), and MTP β where the draft context
deliberately runs the un-duplicated base stack while the target
keeps RYS (still lossless: the target verifies every drafted
token). Wired across all text archs (dense, MoE, Gemma-4 iSWA
dual-cache, Qwen 3.5/3.6 recurrent-hybrid); unsupported archs
fail loudly at load rather than silently ignoring the plan.
**Finding a good plan:** `--rys-probe` enumerates safe-band blocks,
scores each by ΞPPL + a task-probe battery, and prints two
ready-to-paste templates (most-efficient and max-gain):
```bash
sh ./opencoti-llamafile β¦ --rys-probe -m model.gguf -f corpus.txt \
--rys-probe-widths auto --rys-probe-topk 10
```
Treat its output as a shortlist to verify with your own eval, not a
verdict.
---
## 7. Instrumentation β monitor & control API
Three planes (full reference: `docs/features/introspection.md`):
### Boot knobs
Everything in Β§Β§3β6 is a boot flag: set at launch, echoed back at
runtime. By design, tier/residency/DCA/retention **cannot** change
per request (KV layout would differ).
### Per-request control (JSON body fields)
| Field | Default | Effect |
|---|---|---|
| `session_id` | `""` | Sessionβslot affinity: the same session returns to the slot holding its KV (prevents cross-session eviction at `--parallel > 1`). Pair with `cache_prompt: true`. |
| `shared_pool_slot` | `-1` | Attach this request to SharedKVPool slot N (read-only prefix share). |
| `shared_prefix_n_tokens` | `0` | Length of the shared prefix. |
### Runtime introspection
**`GET /props` β `"opencoti"` object** β boot-state echo plus the
*effective* KV state read back from the live cache:
```jsonc
"opencoti": {
"kv": { "cache_type_k": "q8_0", "cache_type_v": "q4_0",
"auto_tier": false,
"effective": { "type_k": "q8_0", "type_v": "q4_0",
"n_cells": 524288, "n_cells_resident": 524288,
"n_layers_spilling": 0, "fully_resident": true,
"is_iswa": true } },
"residency": { "kv_residency_mode": 0, "vram_target_mib": 0 },
"dca": { "enabled": true, "chunk_size": 0, "yarn_factor": 1.0 },
"sparse_attn":{ "enabled": false, "block_size": 128, "topk": 0 },
"speculative":{ "types": ["none","draft-assistant"], "n_max": 3 },
"kv_reuse": { "n_parallel": 4, "kv_unified": true, "cache_ram_mib": 8192 },
"rest_kv": { "eviction": false, "recent": 256, "layer": -1 },
"repeat_layers": null
}
```
`kv.effective` is the only authoritative record of the auto-tier
decision β `configured != effective` is expected when auto-tier
engaged. `fully_resident` / `n_layers_spilling` tell you whether
rolling-KV is streaming.
**`GET /slots` β per-slot `"opencoti"` object** (requires `--slots`):
lifetime `draft_n_total` / `draft_n_accepted` / `draft_acceptance`
per slot, plus the slot's current `session_id` and pool binding.
Operational tell: **sustained draft_acceptance β³ 0.95 at turn end
usually means the model is looping/ruminating** (healthy agentic
decode sits ~0.4β0.9) β pollable, no log-scraping.
**Per-completion `timings`**: `cache_n` (prefix-reuse hits),
`draft_n` / `draft_n_accepted` for that response.
**Quick recipes**
```bash
curl -s :8080/props | jq .opencoti # what is this server running?
curl -s :8080/props | jq .opencoti.kv.effective # did auto-tier/spill engage?
curl -s :8080/slots | jq '.[] | {id, acc: .opencoti.draft_acceptance}'
```
For embedders/tools linking the C API:
`llama_memory_opencoti_kv_info()` (in `llama.h`) returns the same
effective-KV struct.
### Still log-only
SharedKVPool share/reject events, retention-eviction discards,
rolling-KV tactic selection detail, and the auto-tier WARN line
currently appear only in the server log.
---
## 8. Composition matrix
| | quant-KV | auto-tier | rolling-KV | PolyKV pool | DCA | sparse-attn | MTP | RYS |
|---|---|---|---|---|---|---|---|---|
| **quant-KV** | β | K-only honors | β
(tiles dequant-on-lift) | β
| β
| β
(the win case) | β
(turbo+MTP is the top decode combo) | β
|
| **auto-tier** | | β | β
(it *manages* spill) | β
| β
(probes in DCA state) | β
| β
| β
(sizing is eff-plan-aware) |
| **rolling-KV** | | | β | β
| β
| β
| β
| β
(validated: window spill Γ RYS on hybrid) |
| **PolyKV pool (SharedKVPool)** | | | | β | β
| β
| β
| β
(validated: 2-agent share gate Γ `--repeat-layers` on A4B; hybrid-GDN omnimerge Γ NextN MTP full gate, patch `0135`) |
| **DCA** | | | | | β | β
| β
(dual-ctx) | β
(effβsrc mapped) |
| **sparse-attn** | | | | | | β | β
| β
|
| **MTP** | | | | | | | β | β
(draft runs base stack; target keeps RYS) |
Two known guards worth restating: assistant-MTP + `--parallel > 1`
forces `--kv-unified`; SharedKVPool requires
`--kv-unified --no-cache-idle-slots`.
**Reference "agentic serving" launch** (Gemma-4-A4B on a 24 GB card β
quantized KV + MTP + introspection):
```bash
sh ./opencoti-llamafile-0.10.3-c3-x86_64.llamafile --server --port 8080 \
-m gemma4-A4B-Q4_K_M.gguf -ngl 99 --flash-attn on \
-c 262144 --parallel 4 --kv-unified \
-ctk q8_0 -ctv q4_0 \
--spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \
-ngld 99 --spec-draft-n-max 2 \
--slots
```
---
## 9. Internal / superseded machinery (so you don't chase ghosts)
Present in the patch series but **not** user-facing knobs anymore:
- **HeadInfer head-split** (`--headinfer-gpu-heads-frac`): retired as
a manual knob; it survives as one tactic inside rolling-KV's auto
ladder (`auto` is the only value you should pass, and the adapter
does it for you).
- **NEO GPU/CPU FA pipelining** (`--neo-pipeline`): structurally
shipped, default off; no measurable win on single-GPU consumer
hardware. Leave off.
- **Fused-MoE up-gate** (`--fused-moe-up-gate`): niche (+2.4% decode
on OLMoE-class MoE; Gemma-4 already fuses). Default off.
- **Fused-NextN draft graph** (`OPENCOTI_MTP_FUSED_NEXTN=1`): built and
shipped (patch 0093) but default off for a measured reason β the fused
graph isn't shape-invariant, so it rebuilds every cycle and decodes
*slower* than the default autoregressive draft loop (which, post-0128,
is at upstream parity or better). Leave off.
- **ScoutAttention, LMCache**: design-only / deferred β the flags don't
exist.
---
## 10. Verifying an artifact
```bash
# 1. artifact hash matches MANIFEST.json / SHA256SUMS
sha256sum opencoti-llamafile-*.llamafile
# 2. embedded CUDA DSO matches the manifest's backend sha β no execution needed
unzip -p opencoti-llamafile-*.llamafile ggml-cuda.so | sha256sum
# 3. version string
sh ./opencoti-llamafile-*.llamafile --version
```
`MANIFEST.json` also records the git commit and the exact patch list
the artifact was built from; the patch series in `patches/`
reproduces the tree from upstream llamafile 0.10.3.
|