# opencoti-llamafile — usage guide How this engine diverges from upstream [Mozilla-Ocho llamafile](https://github.com/Mozilla-Ocho/llamafile), what the added features are, how each is gated, its knobs and defaults, its limitations, and which features are meant to be used together. Audience: anyone running the packaged `opencoti-llamafile---.llamafile` artifact as a local inference server. Deep-dive design docs live in [`docs/features/`](features/), measured evaluations in [`docs/evaluations/`](evaluations/). --- ## Supported / target model families — read this first opencoti-llamafile loads **any GGUF that upstream llama.cpp loads** — that part is inherited unchanged. But the opencoti feature set (KV tiers, rolling-KV, DCA, MTP, sparse attention, RYS) is developed, tuned, and correctness-gated on **two model families**, in a deliberate primary/secondary split: ### Gemma-4 — PRIMARY target | Model | Kind | Notes | |---|---|---| | Gemma-4 26B-A4B-128e ("A4B") | MoE | the flagship serving target; MTP-validated with its native gemma4-assistant drafter | | Gemma-4 12B / 31B | dense | full feature validation incl. RYS + DCA | | Gemma-4 E2B / E4B | elastic (E-series) | shared-KV elastic layers supported; MTP drafters available | Gemma-4 is what the engine is *for*: its unusual head dims (256 and 512), the iSWA sliding/global dual KV cache, and the per-size [gemma4-assistant MTP drafters](https://huggingface.co/ManniX-ITA) all have dedicated kernels and graph paths here that upstream lacks or handles slowly. `--spec-type draft-assistant`, D256/D512 FA-VEC + scalar-MMA decode, iSWA-aware rolling-KV/SharedKVPool/DCA wiring are all Gemma-4-first features. ### Qwen — SECONDARY target / verification family | Model | Kind | Notes | |---|---|---| | Qwen3.5 / Qwen3.6 (e.g. 35B-A3B) | dense / MoE / hybrid (gated-delta-net) | NextN self-spec MTP (`--spec-type draft-mtp`, no external drafter needed) | | Qwen2.5-14B-1M | dense, native-1M | the long-context/DCA validation vehicle | Qwen is the standard-architecture (head_dim 128) counterweight: every feature that ships is verified on it too, and it carries one feature Gemma doesn't — **NextN self-speculation** (the model's own MTP head drafts; fused multi-step, at/above upstream parity). ### Everything else Other architectures run with upstream behavior and safe fallbacks, but opencoti features are **unvalidated** there, and some are arch-gated: MTP needs NextN tensors (Qwen-style) or a gemma4-assistant drafter; RYS `--repeat-layers` supports the qwen2/qwen3(+MoE)/qwen3.5/qwen3next/gemma-4 forward loops; DCA is validated on Gemma-4 and Qwen2.5-1M. Quality gates (KLD, RULER-niah) were run on the two families above — re-gate before trusting aggressive KV tiers on anything else. --- ## 1. Relationship to upstream llamafile opencoti-llamafile is **upstream llamafile 0.10.3 plus an additive patch series** (`patches/` in the HF repo, `vendors/patches/llamafile/` in the git repo — ~81 patches, numbered `0006`–`0135`). Three properties are contractual: 1. **Off means off.** Every opencoti feature is opt-in behind a flag, env var, or per-request JSON field. With no opencoti flags set, the engine's compute path is **byte-identical to upstream** — this is a regression gate on every patch, not an aspiration. 2. **Lossless by proof, not vibes.** Features that touch the forward pass are gated by logit-equivalence / KLD / RULER-retrieval against vanilla, never by "the output looks fine". Speculative decode is verified-lossless (the output *is* the target model's). 3. **Single file, zero dependencies.** The artifact is a Cosmopolitan APE: one file runs on Linux/macOS/Windows/BSD, x86_64 and aarch64. In the full x86_64 artifact the CUDA backend (`ggml-cuda.so`) is embedded and self-extracts to `~/.llamafile/v//` on first GPU run; the `-win` variant ships without it (see §1.1). TCQ codebooks and quantization tables are compiled in. No installer, no downloads. What upstream gives you is unchanged: the server API (`/completion`, `/v1/chat/completions`, `/props`, `/slots`, …), GGUF loading, sampling, chat templates. opencoti adds serving-efficiency machinery on top, aimed at **multi-session agentic serving on a fixed VRAM budget**: more concurrent sessions per card, longer usable context, faster decode. ```bash chmod +x opencoti-llamafile-0.10.3-c3-x86_64.llamafile sh ./opencoti-llamafile-0.10.3-c3-x86_64.llamafile --server --port 8080 \ -m model.gguf -ngl 99 --flash-attn on # --version → opencoti-0.10.3-c3 ; without --server you get the chat CLI ``` > **Note (Linux):** launch via `sh ./file.llamafile` if your kernel > lacks binfmt_misc APE registration. ### 1.1 Artifact variants — which file to download | Artifact | Size | Runs on | GPU story | |---|---:|---|---| | `…-x86_64.llamafile` | ~4.8 GB | Linux/macOS/BSD, x86_64 + aarch64 | CUDA DSO **embedded** (x86_64-linux, sm_75/80/86/89/90/120f); self-extracts on first `-ngl` run | | `…-win-x86_64.llamafile.exe` | ~55 MB | **Windows** + Linux/macOS/BSD, x86_64 + aarch64 | none embedded — CPU works everywhere out of the box; GPU via side-load (below) | | `…-aarch64.llamafile` | ~1.9 GB | same platforms | CUDA sbsa DSO **embedded** (aarch64-linux, sm_110f DGX Spark GB10 / Jetson Thor + sm_121a); GPU out of the box on aarch64-linux CUDA hosts | The host binary inside every variant is byte-for-byte the same APE (same patches, same features); they differ only in the embedded GPU backend. The `-win` variant exists because **Windows refuses to run executables larger than 4 GB**, so the full artifact — 4.8 GB with the CUDA DSO inside — cannot ship for Windows; the bare APE renamed to `.exe` (Windows needs the extension) can. **GPU with the `-win`/bare variant:** - **Windows:** install the NVIDIA CUDA Toolkit + MSVC and run with `-ngl 99`; llamafile compiles a native `ggml-cuda.dll` on the fly into `%USERPROFILE%\.llamafile\` on first use (upstream llamafile mechanism, unchanged). - **Linux x86_64:** download the published standalone DSO (`dso/-/ggml-cuda-x86_64.so` in the HF repo — same bytes as the one embedded in the full artifact, sha in `releases//SHA256SUMS`) and place it where the engine looks before attempting extraction: ```bash mkdir -p ~/.llamafile/v/0.10.3/ cp ggml-cuda-x86_64.so ~/.llamafile/v/0.10.3/ggml-cuda.so ``` - **Linux aarch64 (DGX Spark / GB10 class):** same mechanism with the sbsa DSO — `dso/-/ggml-cuda-sbsa-aarch64.so` → `~/.llamafile/v/0.10.3/ggml-cuda.so`. Or just download the `…-aarch64.llamafile` artifact, which has it embedded. - **No CUDA available:** nothing to do — it falls back to CPU inference automatically. --- ## 2. Feature map — what exists and how it's gated | Feature | Default | Turn on with | Class | |---|---|---|---| | Session-keyed KV reuse | off (per request) | `session_id` JSON field | latency | | ReST-KV retention eviction | **off** | `--rest-kv-eviction` | quality-under-overflow | | KV quantization (scalar) | f16 | `-ctk` / `-ctv` | capacity | | TurboQuant / TCQ KV tiers | off | `-ctk`/`-ctv turbo*` | capacity | | Auto KV-tier policy | **off** | `OPENCOTI_KV_AUTO_TIER=1` | capacity (policy) | | PolyKV pool (SharedKVPool) | off (per request) | `shared_pool_slot` JSON field | multi-agent capacity | | Rolling-KV window / spill | **auto** (engages only under pressure) | `--vram-target`, `--kv-residency-mode` | capacity | | Mixed-KV spilled tail | off | `-ctkt` / `-ctvt` | capacity | | DCA long-context | **off** | `--dca on` | context extension | | Sparse attention (block-selector) | **off** | `--sparse-attn on` | long-ctx decode speed | | Sparse-V | auto on iSWA+quant-V, else off | `TURBO_SPARSE_V_TAU` | decode speed | | MTP speculative decode | **off** | `--spec-type` + drafter | decode speed | | RYS layer duplication | **off** | `--repeat-layers` | quality | | RYS probe | off | `--rys-probe` | tooling | | Lazy slot context | off | `--slot-initial-ctx`, `--slot-shrink-idle-ms` | embedder memory | | Introspection API | **always on** | `GET /props`, `GET /slots` | observability | Every boot flag also has an env twin (`OPENCOTI_LLAMAFILE_` for adapter-typed fields, `LLAMA_ARG_*` for llama.cpp-registered ones). --- ## 3. KV capacity stack — PolyKV **PolyKV** is the umbrella name for this whole stack: the compressed shared KV pool. Concretely it is the KV quantization tiers of §3.1 plus the multi-agent SharedKVPool of §3.4, stacking with the auto-tier policy (§3.2) and the rolling-KV window (§3.3). If you arrived here looking for "PolyKV" from an announcement: §3.4 is the shared-prefix pool itself; §3.1 is what the pooled cells are made of. These four features share one goal — **fit more context / more sessions in fixed VRAM** — and are designed to stack. Recommended order of adoption: scalar quant → auto-tier → rolling-KV → turbo tiers → SharedKVPool. ### 3.1 KV quantization: scalar types + TurboQuant/TCQ tiers (PolyKV M6) The KV cache type is set per-tensor-half: `-ctk ` (keys) and `-ctv ` (values), independently — **asymmetric pairs are first-class** (e.g. `-ctk q8_0 -ctv q4_0`). Supported types: `f16`, `bf16`, `q8_0`, `q6_0`, `q5_1`, `q5_0`, `q4_0` (scalar), `turbo2`, `turbo3`, `turbo4`, `turbo8` (TurboQuant, MSE-optimal with Walsh-Hadamard rotation + InnerQ), `turbo2_tcq`, `turbo3_tcq` (trellis-coded, Viterbi-encoded). **Which to pick (measured):** - **8-bit / 4-bit: use `q8_0` / `q4_0`.** The native scalar types dominate turbo8/turbo4 at equal width — turbo earns nothing there. - **`-ctk q8_0 -ctv q4_0`** is the workhorse asymmetric pair: keys keep 8-bit fidelity (attention logits are K-sensitive), values take the compression. - **3 bits and below is TurboQuant territory:** `turbo3` Pareto-beats q4_0 (90 vs 129 MiB KV at equal quality, teacher-forced TV 0.0067 vs 0.0094); `turbo2` is the smallest logit-equivalent KV that exists (~2 bit) — the 256k-context play. TCQ variants trade encode cost for a further fidelity step at the same width. - All shipped tiers pass logit-equivalence gates; decode runs the quantized data **in-register** in the flash-attention kernel (no f16 materialization) for turbo2/3/4 and TCQ. Limitations: turbo8 uses a materialize fallback (not fused); at Gemma-4's head_dim 512 only turbo2/turbo3 have fused D=512 instances; prefill on very long prompts uses a hybrid path automatically. Quality validation on Gemma franken-merges must use retrieval (niah), not perplexity. ### 3.2 Auto KV-tier (`OPENCOTI_KV_AUTO_TIER=1`) Boot policy: pick the **least-compressing scalar pair that keeps the whole KV resident** in the VRAM budget; if even that spills, the T\* model decides between "small f16 spill" and "quantize one tier down" by predicted tokens/s drop. Knobs (env): `OPENCOTI_KV_AUTO_TIER=1` (master), `OPENCOTI_KV_TSTAR_DROP` (target drop, default 20%), `OPENCOTI_KV_TSTAR_MAX_SPILL_MIB` (default 800), `OPENCOTI_KV_AUTO_TIER_TAIL=1` (also auto-pick a q4_0 spilled tail). Explicit `-ctv` disables auto entirely; explicit `-ctk` holds K and walks only V. Dense full-attention models only (iSWA models keep f16). Read back what it decided: `GET /props → .opencoti.kv.effective` — the *configured vs effective* split exists exactly because auto-tier may override you. ### 3.3 Rolling-KV window (residency / spill) "KV doesn't have to fit." Each layer keeps a device-resident window of recent positions; the tail `[0 … window_start)` lives in pinned host RAM and is streamed through the attention kernel per-tile, merged exactly via online-softmax (LSE). When everything fits, every layer is `GPU_RESIDENT` and the path is byte-identical to vanilla — the feature only *engages* under memory pressure. Flags: `--vram-target ` (budget cap; `0` = all free VRAM minus reserve), `--kv-residency-mode {auto,head,window}` (default `auto`; leave it), `-ctkt` / `-ctvt` (distinct, more-compressed types for the spilled tail — "mixed-KV": f16 recent window ⊕ q4_0 tail). Performance model (RTX 3090, PCIe ~6.5 GB/s): spill decode sits at the PCIe floor, `t(token) ≈ fixed + tail_bytes / link_bw` — linear, no cliff. On a fast-link host (RTX 6000, ~50 GB/s) window-mode spill is genuinely usable; on consumer PCIe it's a last resort — prefer quantizing (that's what auto-tier automates). Limitations: while a window is spilled, context-shift and prompt-cache-reuse are guarded off (requests bounded at `n_ctx`); the compute-buffer reserve for long contexts is measured automatically at boot (two-pass reserve — no knob). ### 3.4 PolyKV SharedKVPool (multi-agent shared prefix) N agents attending **one physical copy** of a common prefix (system prompt + tool defs). Per-request JSON, no CLI flag: ```jsonc { "shared_pool_slot": 0, "shared_prefix_n_tokens": 4096, … } ``` Server must run `--kv-unified --no-cache-idle-slots` (the latter is mandatory — the default idle-slot save/clear would evict the pooled prefix). **What the pool speeds up — measured** (Gemma-4-26B-A4B Q4_K_M, RTX 3090, q8/q8 unified KV, P≈5k-token shared prefix, greedy fixed-length decode, 8 concurrent sessions unless noted): | axis | naive (N private copies) | shared pool | gain | |---|---|---|---| | KV cells (N=8) | ~8·P | P + suffixes | **6.9×** (~306 vs ~9 agents on a fixed buffer) | | prefill, 8 sessions joining | 22.8 s | 4.6 s | **~5×** (prefix enters KV once per pool) | | steady-state batched decode (N=8) | ~190 tok/s | 217 tok/s | **+14%** (8 queries read one physical prefix — L2 reuse, smaller cell span) | | multi-turn re-query (N=8) | 99 tok/s | 225 tok/s | **2.3×** (see note) | | iso-speed capacity | 8 sessions @ 24.0 tok/s each | ≥12 sessions @ ≥26.9 tok/s each | **≥1.5×** sessions (crossover not reached at N=12; aggregate 315 tok/s) | The multi-turn row is iSWA-specific and easy to miss: a private slot that has decoded past its prompt cannot partially rewind (upstream SWA-checkpoint semantics, llama.cpp PR #13194), so re-querying it pays a checkpoint restore or a full re-prefill every turn. The pool slot never decodes, so its prefix never slides — every re-attach is free. Note the capacity row is about per-session speed, not just aggregate: 12 pooled sessions each decode faster than 8 private ones. **Cross-architecture results** (RTX 6000 96GB, P≈5073, GEN=256, A/B/A naive/shared/naive): the pool is validated on all three attention architectures, and the memory axis is architecture-independent (~6.8–6.9× at N=8 — it counts cells, not attention math). | axis | Qwen2.5-14B-1M Q8_0 (pure full attention) | Qwen3.6-27B-Omnimerge-v4 Q4_K_M (hybrid GDN + NextN MTP n=3) | |---|---|---| | KV cells (N=8) | **6.78×** (~406 vs ~9 agents on a fixed 8192-cell buffer) | **6.89×** (~304 vs ~9 agents) | | batched decode N=8 | 358 → 403 tok/s (**+12.5%**) | 106 → 123 tok/s (**+15.6%**) | | batched decode N=24 | 412 → 742 tok/s (**+80%**; 17.2 → 30.9 tok/s per session) | 90 → 125 tok/s (**+39.7%**; 3.7 → 5.2 per session) | | shared-only sweep N=32/48/64 | 813 / 871 / 865 tok/s (plateau ~870 near N=48) | 126 / 124 / 124 tok/s (saturates by N≈24–32) | The shared-vs-naive decode gain **grows with N** on both. On hybrid/recurrent models (delta-net, mamba) the *absolute* aggregate saturates much earlier than on pure attention — the recurrent layers batch worse — so there the pool buys **concurrency capacity and memory**, not aggregate throughput past N≈24. **How to prefill the pool — use the common-prefix token array, not the document text.** Tokenizers merge across the document/suffix boundary (on the Qwen tokenizer the last prefix token fuses with the suffix start), so `tok(DOC)` can be one token longer than the common prefix the agents actually share — and a pool that is even one token longer than `shared_prefix_n_tokens` cannot be shared exactly. The correct client sequence: ```jsonc // 1. tokenize the FULL agent prompts and compute // P = min over agents of commonPrefixLen(tok(DOC), tok(DOC+suffix_i)) // 2. prefill the pool slot with the token array itself (llama.cpp // /completion accepts token arrays) — pool state == P on ANY tokenizer: { "prompt": [/* tok(DOC+suffix_1)[:P] */], "id_slot": 0, "cache_prompt": true, "n_predict": 1 } // 3. agents attach with prompts STRICTLY longer than P: { "prompt": "", "id_slot": 1, "shared_pool_slot": 0, "shared_prefix_n_tokens": P, … } ``` On attention-only models a text prefill happens to work (the ranged cell copy tolerates the extra token); on hybrid/recurrent targets it silently disables every share — see the gotcha below. The token-array prefill is correct everywhere. **Hybrid/recurrent gotcha (GDN / mamba / `qwen35moe`-class models).** A recurrent cache has one rolling state per sequence, not per-position cells, so a pool share is only possible as an **exact full-state** share. The server enforces this: the share engages only when `shared_prefix_n_tokens == pool state length` *and* the request prompt is strictly longer than the shared prefix; anything else logs `poly-kv-pool: hybrid/recurrent target needs exact full-state share … skipping share, full reprocess (bug-2203)` and falls back to a full (correct, slower) reprocess. If you see zero speedup on a hybrid model — or mass HTTP 500s at high N because N unshared full prompt copies overflow the unified KV — grep the server log for that WARN: it almost always means the pool was prefilled with text instead of the token array. (Older builds crashed outright here — `failed to remove sequence N with p0=…` — fixed by patch `0135`.) Sizing note: the pooled prefix pins P cells in **both** iSWA caches (global + SWA) for the pool's lifetime. Budget `-c` for pool prefix + N session windows + generation headroom, or long-running sessions can exhaust slot allocation mid-round. Composes with KV quantization (the pool holds quantized cells) and with session KV-reuse. The pool is read-only for consumers; each agent's divergent suffix is private. **Tiering is pinned per-pool, never per-session.** The K/V tiers — including the mixed-KV recent-window ⊕ compressed-tail pair — are properties of the boot-allocated cache tensors, chosen once at boot (by you or by auto-tier) before any session exists. The window/tail boundary is a per-layer residency budget over the *physical cell axis*, so a prefix cell is in the resident window or evicted (and quantized exactly once, on eviction) for **all** sequences simultaneously. Sharing itself is not copy-on-write: a sharer joins the prefix by adding its sequence bit to the existing cells, and a diverging session just appends private suffix cells — there is no per-session copy that could be re-quantized, and no way for two sessions to see the same prefix at different tiers. The flip side: you cannot give one session a higher-precision read of a shared prefix than another; that would require forking the prefix into a private copy, which is exactly the O(N) memory cost the pool exists to avoid. --- ## 4. Long context ### 4.1 DCA — Dual Chunk Attention (training-free context extension) Splits attention into intra-chunk / successive / inter-chunk position regimes and merges them exactly by LSE, so a model trained at `n_ctx_train` serves multiples of it **without retraining**. Flags: `--dca on` (default **off**), `--dca-chunk-size N` (default derives from the model's training context; explicit 8192 is the validated recipe), `--dca-yarn-factor F` (default 1.0; measured neutral for retrieval — leave it). Serve beyond the GGUF's declared context with `--override-kv .context_length=int:1048576`. Validated recipe (Gemma-4-A4B, n_ctx_train 256k): ```bash --dca on --dca-chunk-size 8192 -fa on --parallel 1 \ --override-kv gemma4.context_length=int:1048576 ``` Measured retrieval (RULER-VT, n=50): **256k 0.964 · 512k 0.996 · 768k 0.984 · 1M 0.916** — a gentle −7 pp at 4× native, no cliff. Counter-proof on Qwen3-8B (native 41k): plain attention collapses at 128k (PPL 19.2) while DCA holds PPL 7.3. Works on Gemma-4 (its 5 global layers; SWA layers untouched) and Qwen2.5/3/3.5 (all layers). Composes with quantized KV (scalar pairs all pass; q8-DCA decode costs ~2× vs f16-DCA), sparse attention, and rolling-KV. Limitations: DCA caches K un-rope'd → **launch-time toggle only** (a server booted DCA-on can't switch off per request); expect approximation, not identity, past one chunk. On models that are already *native* long-context (e.g. Qwen2.5-1M), DCA can only approximate down — don't use it there. ### 4.2 Sparse attention (Quest block-selector) + sparse-V Two independent decode-bandwidth levers: - **Block-selector** (`--sparse-attn on`): per-block min/max key bounds give an upper bound on each block's attention mass; decode visits only the top-K blocks (+ sinks + recent). Flags: `--sparse-attn-block-size` (128), `--sparse-attn-topk` (default 0 = visit **all** blocks, i.e. no skipping; pass `auto` for adaptive max(64, n_blocks/4), or an explicit block count), `--sparse-attn-recent`, `--sparse-attn-sink` (1), `--sparse-attn-refresh` (8 — re-select every N decode steps), `--sparse-attn-mode` (0). Default **off**. - **Sparse-V**: skips V-dequant for negligible-weight positions inside visited blocks. **Self-configuring**: on iSWA models with quantized V it auto-sets τ=0.05; elsewhere it stays off. Manual override: `TURBO_SPARSE_V_TAU=`. When to use: **long context on quantized KV.** The win grows with context (selectivity 0.91@16k → 0.99@40k and climbing) and lives on quantized KV: q8_0 ⊕ sparse at 50% coverage measured **1.34× decode at niah 100**. Both levers stack (1.31× combined measured). When *not* to use: short contexts or f16 KV on mid-size models — the decode isn't KV-bandwidth-bound there and the selector overhead can make it *slower* than dense. τ values don't transfer across models; retune if you override manually. --- ## 5. Decode speed — MTP speculative decoding Lossless speculative decode; the emitted text is the target model's own (verified). Two flavours, chosen by `--spec-type`: ### 5.1 `--spec-type draft-assistant` (external drafter — Gemma-4) A small `gemma4-assistant` drafter GGUF rides the target's embeddings: ```bash --spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \ -ngld 99 --spec-draft-n-max 2 ``` `--mtp-head` (alias `-md`) names the drafter; **`-ngld 99` matters** (a CPU-resident draft head erases the win). Drafters for A4B/12B/27B/E2B/E4B are published per-size. Setting `mtpHead` in the TS adapter auto-derives the rest. ### 5.2 `--spec-type draft-mtp` (NextN self-spec — Qwen) Qwen 3.5/3.6 GGUFs that embed a NextN/MTP head self-speculate — no second file: ```bash --spec-type draft-mtp --spec-draft-n-max 3 ``` Runs per-slot under `--parallel` (multi-session capable). **Measured (RTX 3090 + upstream-parity campaign):** A4B assistant decode beats upstream llama.cpp b9859 at every depth (+6.6/+12.9/+7.6% at n_max 1/2/3); combined with `turbo3_tcq` KV it reaches **~89 tok/s vs 52.9 plain (+69%)**. Qwen-35B NextN sits at parity with upstream. Recommended depth: `--spec-draft-n-max 2–3` (A4B), `3` (Qwen NextN). Notes/limits: acceptance dips a few pp at depth ≥2 (chained-draft numerics — expected); with `--parallel > 1`, assistant-MTP requires `--kv-unified` (auto-forced with a boot warning). Composes with turbo/TCQ KV tiers (its biggest lever), DCA, and quantized KV. Watch live acceptance per slot via `GET /slots` (§7). --- ## 6. Quality — RYS layer duplication `--repeat-layers` re-runs a contiguous block of **middle** layers, weight-shared: zero extra parameter VRAM, no new GGUF, quant-agnostic. You pay in KV cache and tokens/s proportional to the extra effective layers; you buy quality-per-token. ```bash --repeat-layers 33,34 # +1 layer (RYS-S) --repeat-layers 26-34 # +8 layers ([26,34) half-open, RYS-XL) --repeat-layers 8-12;20-24 # disjoint blocks ``` Rules that matter: - **Middle layers only.** Duplicating first/last layers reliably produces incoherent output on merge-fragile models — this is a model property, not an engine bug; the engine prints a boot advisory when a plan touches the boundary band. - Absent flag = identity = byte-identical to stock. - Composes with the full stack: quantized KV, DCA, rolling-KV window/spill, sparse-attn (the residency/DCA/sparse sizing paths are effective-plan-aware), and MTP — where the draft context deliberately runs the un-duplicated base stack while the target keeps RYS (still lossless: the target verifies every drafted token). Wired across all text archs (dense, MoE, Gemma-4 iSWA dual-cache, Qwen 3.5/3.6 recurrent-hybrid); unsupported archs fail loudly at load rather than silently ignoring the plan. **Finding a good plan:** `--rys-probe` enumerates safe-band blocks, scores each by ΔPPL + a task-probe battery, and prints two ready-to-paste templates (most-efficient and max-gain): ```bash sh ./opencoti-llamafile … --rys-probe -m model.gguf -f corpus.txt \ --rys-probe-widths auto --rys-probe-topk 10 ``` Treat its output as a shortlist to verify with your own eval, not a verdict. --- ## 7. Instrumentation — monitor & control API Three planes (full reference: `docs/features/introspection.md`): ### Boot knobs Everything in §§3–6 is a boot flag: set at launch, echoed back at runtime. By design, tier/residency/DCA/retention **cannot** change per request (KV layout would differ). ### Per-request control (JSON body fields) | Field | Default | Effect | |---|---|---| | `session_id` | `""` | Session→slot affinity: the same session returns to the slot holding its KV (prevents cross-session eviction at `--parallel > 1`). Pair with `cache_prompt: true`. | | `shared_pool_slot` | `-1` | Attach this request to SharedKVPool slot N (read-only prefix share). | | `shared_prefix_n_tokens` | `0` | Length of the shared prefix. | ### Runtime introspection **`GET /props` → `"opencoti"` object** — boot-state echo plus the *effective* KV state read back from the live cache: ```jsonc "opencoti": { "kv": { "cache_type_k": "q8_0", "cache_type_v": "q4_0", "auto_tier": false, "effective": { "type_k": "q8_0", "type_v": "q4_0", "n_cells": 524288, "n_cells_resident": 524288, "n_layers_spilling": 0, "fully_resident": true, "is_iswa": true } }, "residency": { "kv_residency_mode": 0, "vram_target_mib": 0 }, "dca": { "enabled": true, "chunk_size": 0, "yarn_factor": 1.0 }, "sparse_attn":{ "enabled": false, "block_size": 128, "topk": 0 }, "speculative":{ "types": ["none","draft-assistant"], "n_max": 3 }, "kv_reuse": { "n_parallel": 4, "kv_unified": true, "cache_ram_mib": 8192 }, "rest_kv": { "eviction": false, "recent": 256, "layer": -1 }, "repeat_layers": null } ``` `kv.effective` is the only authoritative record of the auto-tier decision — `configured != effective` is expected when auto-tier engaged. `fully_resident` / `n_layers_spilling` tell you whether rolling-KV is streaming. **`GET /slots` → per-slot `"opencoti"` object** (requires `--slots`): lifetime `draft_n_total` / `draft_n_accepted` / `draft_acceptance` per slot, plus the slot's current `session_id` and pool binding. Operational tell: **sustained draft_acceptance ≳ 0.95 at turn end usually means the model is looping/ruminating** (healthy agentic decode sits ~0.4–0.9) — pollable, no log-scraping. **Per-completion `timings`**: `cache_n` (prefix-reuse hits), `draft_n` / `draft_n_accepted` for that response. **Quick recipes** ```bash curl -s :8080/props | jq .opencoti # what is this server running? curl -s :8080/props | jq .opencoti.kv.effective # did auto-tier/spill engage? curl -s :8080/slots | jq '.[] | {id, acc: .opencoti.draft_acceptance}' ``` For embedders/tools linking the C API: `llama_memory_opencoti_kv_info()` (in `llama.h`) returns the same effective-KV struct. ### Still log-only SharedKVPool share/reject events, retention-eviction discards, rolling-KV tactic selection detail, and the auto-tier WARN line currently appear only in the server log. --- ## 8. Composition matrix | | quant-KV | auto-tier | rolling-KV | PolyKV pool | DCA | sparse-attn | MTP | RYS | |---|---|---|---|---|---|---|---|---| | **quant-KV** | — | K-only honors | ✅ (tiles dequant-on-lift) | ✅ | ✅ | ✅ (the win case) | ✅ (turbo+MTP is the top decode combo) | ✅ | | **auto-tier** | | — | ✅ (it *manages* spill) | ✅ | ✅ (probes in DCA state) | ✅ | ✅ | ✅ (sizing is eff-plan-aware) | | **rolling-KV** | | | — | ✅ | ✅ | ✅ | ✅ | ✅ (validated: window spill × RYS on hybrid) | | **PolyKV pool (SharedKVPool)** | | | | — | ✅ | ✅ | ✅ | ✅ (validated: 2-agent share gate × `--repeat-layers` on A4B; hybrid-GDN omnimerge × NextN MTP full gate, patch `0135`) | | **DCA** | | | | | — | ✅ | ✅ (dual-ctx) | ✅ (eff→src mapped) | | **sparse-attn** | | | | | | — | ✅ | ✅ | | **MTP** | | | | | | | — | ✅ (draft runs base stack; target keeps RYS) | Two known guards worth restating: assistant-MTP + `--parallel > 1` forces `--kv-unified`; SharedKVPool requires `--kv-unified --no-cache-idle-slots`. **Reference "agentic serving" launch** (Gemma-4-A4B on a 24 GB card — quantized KV + MTP + introspection): ```bash sh ./opencoti-llamafile-0.10.3-c3-x86_64.llamafile --server --port 8080 \ -m gemma4-A4B-Q4_K_M.gguf -ngl 99 --flash-attn on \ -c 262144 --parallel 4 --kv-unified \ -ctk q8_0 -ctv q4_0 \ --spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \ -ngld 99 --spec-draft-n-max 2 \ --slots ``` --- ## 9. Internal / superseded machinery (so you don't chase ghosts) Present in the patch series but **not** user-facing knobs anymore: - **HeadInfer head-split** (`--headinfer-gpu-heads-frac`): retired as a manual knob; it survives as one tactic inside rolling-KV's auto ladder (`auto` is the only value you should pass, and the adapter does it for you). - **NEO GPU/CPU FA pipelining** (`--neo-pipeline`): structurally shipped, default off; no measurable win on single-GPU consumer hardware. Leave off. - **Fused-MoE up-gate** (`--fused-moe-up-gate`): niche (+2.4% decode on OLMoE-class MoE; Gemma-4 already fuses). Default off. - **Fused-NextN draft graph** (`OPENCOTI_MTP_FUSED_NEXTN=1`): built and shipped (patch 0093) but default off for a measured reason — the fused graph isn't shape-invariant, so it rebuilds every cycle and decodes *slower* than the default autoregressive draft loop (which, post-0128, is at upstream parity or better). Leave off. - **ScoutAttention, LMCache**: design-only / deferred — the flags don't exist. --- ## 10. Verifying an artifact ```bash # 1. artifact hash matches MANIFEST.json / SHA256SUMS sha256sum opencoti-llamafile-*.llamafile # 2. embedded CUDA DSO matches the manifest's backend sha — no execution needed unzip -p opencoti-llamafile-*.llamafile ggml-cuda.so | sha256sum # 3. version string sh ./opencoti-llamafile-*.llamafile --version ``` `MANIFEST.json` also records the git commit and the exact patch list the artifact was built from; the patch series in `patches/` reproduces the tree from upstream llamafile 0.10.3.