# F4 M6 — Production llamafile build pipeline Owns: producing the `.llamafile` artifacts the embedder package consumes at runtime, from source under `vendors/sources/llamafile/` plus the opencoti patch series at `vendors/patches/llamafile/`. The build orchestrator is `packages/opencoti-llamafile/script/build-pipeline.ts`. Helpers live at `packages/opencoti-llamafile/src/build-pipeline.ts`. ## Why this exists Through F4 M0-M4 the embedder ran a prebuilt `mann1x/claude-hooks v1.4.0 qwen3-embedding-0.6b-16k.llamafile` asset pinned at `vendors/pin/embedder.txt`. That binary embeds `ggml-cuda.so` and a Qwen3 GGUF, and predates every opencoti patch (lazy slot context, shrink-on-idle, et al.). For F4 production we need: - **All backend DSOs built from source** (CUDA / ROCm / Vulkan), not a runtime side-load nobody controls. - **The opencoti patch series actually compiled in** — the prebuilt rejects `--slot-initial-ctx` and `--slot-shrink-idle-ms` because those flags don't exist in upstream 0.10.1. - **No cross-repo dependency** on `mann1x/claude-hooks` for an opencoti artifact. The result is one or two artifacts under `dist/llamafile/`: - `opencoti-llamafile--.llamafile` — thin binary + every staged DSO. The "GPU runtime", no model bundled. - `opencoti-embedder--.llamafile` — the above + an embedded GGUF + `.args`. The embedder package's launcher target. Each artifact has a sibling `MANIFEST.json` recording SHA256 of the artifact, the embedded patches, and per-DSO digests. ## Toolchains per backend | Backend | Build script | Requirement | Probe env | Default path | |---------|-------------------------------------------------|--------------------------|------------------------------|------------------------| | CUDA | `vendors/sources/llamafile/llamafile/cuda.sh` | NVIDIA CUDA 12.x or 13.x | `OPENCOTI_CUDA_PATH` | `/usr/local/cuda-12.6` | | ROCm | `vendors/sources/llamafile/llamafile/rocm.sh` | AMD ROCm 5+ (hipcc) | `OPENCOTI_ROCM_PATH` | `/opt/rocm` | | Vulkan | `vendors/sources/llamafile/llamafile/vulkan.sh` | Vulkan SDK 1.3+ (glslc) | `OPENCOTI_VULKAN_SDK` | system `glslc` on PATH | Each backend probe runs read-only — missing toolchain produces a WARN and the backend is skipped, not a hard build failure. Only the `package` subcommand hard-fails if zero backends produced a DSO (and `--cpu-only` wasn't passed explicitly). ### CUDA version per GPU generation Mozilla-Ocho's `cuda.sh` enables different `-gencode` arch lists depending on the detected CUDA toolkit version: | Target GPU | Compute cap | Minimum CUDA | Notes | |------------------------------------|---------------|--------------|--------------------------------------------------------------------------| | Turing (RTX 20xx, T4) | sm_75 | 12.x | Baseline arch — included in every build | | Ampere (A100, RTX 30xx) | sm_80 / sm_86 | 12.x | Baseline arch | | Ada Lovelace (RTX 40xx) | sm_89 | 12.x | Baseline arch | | Hopper (H100) | sm_90 | 12.x | Baseline arch | | **Blackwell consumer (RTX 50xx)** | **sm_120** | **13.x** | Requires CUDA 13.x branch — adds `sm_120f` + `--compress-mode=size` (4x DSO size reduction) | | Blackwell Jetson Thor (aarch64) | sm_110 | 13.x | aarch64 only | | Blackwell DGX Spark GB10 (aarch64) | sm_121 | 13.x | aarch64 only | **For RTX 50xx targets, set `OPENCOTI_CUDA_PATH=/usr/local/cuda-13.x` when building** — the 12.x branch will emit a DSO with no sm_120 kernels and the driver cannot JIT-fallback (the `code=sm_X` form embeds SASS only, not PTX). Symptom on RTX 50xx with a 12.x-built DSO: `ggml-cuda.cu:97: CUDA error` during first kernel launch. ## End-to-end pipeline ```bash # 1. Reset vendored sources, init nested submodules, apply # Mozilla-Ocho's own patches, then apply opencoti's series. bun run build:llamafile:apply # 2. Build the thin patched llamafile binary (no DSOs embedded yet). bun run build:llamafile # = apply + make # Output: vendors/sources/llamafile/o/llamafile/llamafile # 3. Build every available backend DSO. Soft-fails per backend on # missing toolchain; hard-fails only if zero built. bun run build:llamafile:all-backends # Output (per backend that built): # vendors/dist/llamafile///ggml-.so # ~/.llamafile/v//ggml-.so (runtime side-load mirror) # 4. Assemble production artifacts. bun run build:llamafile:package # Always emits: # dist/llamafile/opencoti-llamafile---.llamafile # dist/llamafile/opencoti-llamafile---.llamafile.MANIFEST.json # 5. Or, for the composite-embedder asset (binary + DSOs + Qwen3 GGUF + .args): bun run build:llamafile:package -- --with-model ~/.opencoti/qwen3-embedding-0.6b.gguf # Additionally emits: # dist/llamafile/opencoti-embedder--.llamafile # dist/llamafile/opencoti-embedder--.llamafile.MANIFEST.json ``` Iteration on a single backend uses its dedicated subcommand (`build:llamafile:cuda`, `:rocm`, `:vulkan`). The `:make` subcommand re-runs cosmocc-make alone — keeps in-tree edits in place while you're polishing a patch you haven't yet captured under `vendors/patches/llamafile/`. ## `package` subcommand details ``` build:llamafile:package [--with-model PATH] [--with-args PATH] [--arch ID] [--cpu-only] [--win-variant] ``` - `--with-model PATH` — GGUF to embed. Filename is preserved; the embedded `.args` points at `/zip/`. - `--with-args PATH` — explicit `.args` to embed instead of the generated one. The generated default matches `buildEmbedderArgs()` in `packages/opencoti-embedder/src/manager.ts` (--server, --embedding, --pooling last, --ctx-size 16384, --parallel 2, --slot-initial-ctx 4096, --slot-shrink-idle-ms 30000). - `--arch ID` — architecture string in the output filename. Detected from `os.arch()` by default (`x86_64`/`aarch64`). - `--cpu-only` — skip embedding any DSO. Produces a binary that works on any host but always runs on CPU. Useful for CI smoke on machines with no GPU. - `--win-variant` — package the bare thin APE (no embedded DSO) as `…-win-x86_64.llamafile.exe` and hard-fail if it exceeds 4 GiB: Windows refuses executables >4 GB, so this is the only variant that can run there (CPU out of the box; GPU via side-load DSO on Linux or llamafile's on-the-fly nvcc DLL compile on Windows). MANIFEST records `arch: "win-x86_64"`, `backends: []`. Embedding uses the vendored zipalign at `vendors/sources/llamafile/o/third_party/zipalign/zipalign`. For DSOs the `-j0` flag strips directory prefixes so cosmo_dlopen resolves the entry as the bare filename it expects. ## Embedder package: local-asset modes (dev escape hatch) The embedder downloader at `packages/opencoti-embedder/src/download.ts` understands two local-asset modes added in M6 to avoid a GitHub Release round trip during dev cycles: ### `file:` pin Change the first field of `vendors/pin/embedder.txt` from `/` to `file:`. Path is anything `path.resolve()` can interpret (absolute or relative to caller cwd). The SHA256 is still verified. ``` file:./dist/llamafile/opencoti-embedder-0.10.1-x86_64.llamafile v0 opencoti-embedder.llamafile ``` ### `OPENCOTI_EMBEDDER_LOCAL` env Bypasses the pin's path entirely while still verifying SHA. Set `OPENCOTI_EMBEDDER_LOCAL=$(pwd)/dist/llamafile/opencoti-embedder-…llamafile` and run the embedder smoke as usual — the downloader copies bytes from disk instead of fetching from GitHub. Both modes still write to `destDir/` so the rest of the launcher pipeline works unchanged. ## Verifying a built / deployed binary carries the full patch chain A dev `o/` build (`:make` + `:cuda`) emits **no** `MANIFEST.json` — only `package`/`release-cut` do. So when you have a binary of uncertain vintage (an `o/` artifact after a source reset, or a copy deployed to another host) and need to answer *"is this the full all-patches build, or vanilla/partial?"* **without** a 45–86 min rebuild, verify by symbols + arg names. The vendored `llama.cpp` tree is normally **clean** at rest (patches apply at build time and the tree is reset for upstream-syncability), so a clean tree + a feature-rich binary is the *expected* steady state — never infer "missing patches" from a clean source tree; inspect the artifact instead. **CUDA DSO** (`~/.llamafile/v//ggml-cuda.so` and `vendors/dist/llamafile/cuda//ggml-cuda.so` — must be byte-identical, G4): ```sh nm -D --defined-only ~/.llamafile/v//ggml-cuda.so | c++filt | \ grep -ciE 'dca_fused_case|ggml_flash_attn_ext_dca_fused|ggml_cuda_flash_attn_ext_lse' # DCA (0078+0081) # expect: DCA family ~150 syms incl ~144 dca_fused_case<…> (the 0082/#531 shard), # turbo ~24 (ggml_turbo_wht, turbo_cpu_fwht_inverse, ggml_cpy_turbo_f16_cuda), # tcq ~12, q6_0 present. Zero DCA/turbo syms ⇒ vanilla or pre-0078 DSO. ``` **Host binary** (cosmo APE — `nm` won't read it; use `strings`): ```sh strings -n4 vendors/sources/llamafile/o/llamafile/llamafile | \ grep -E '^(q6_0|turbo[2348]|turbo[23]_tcq|dca-chunk-size|QK_TURBO3_TCQ)$' # all present ⇒ arg.cpp kv_cache_types + DCA flag patches (0073/0078/0081) are in. ``` **Host vs DSO mtime split is benign.** `:make` (host) and `:cuda` (DSO) build independently; when the most recent patches are CUDA/build-only (e.g. 0082 build-parallelism, 0083 tcq-warp, 0084 cuda-graphs), only the DSO is rebuilt, so the host binary legitimately predates the DSO. Verify by *symbols*, not mtime. To make a build self-describing, run `package` (emits the sibling `MANIFEST.json` with per-artifact SHA256); the in-tree `o/` build has none. ## What's NOT in M6 These are tracked separately and intentionally out of M6 scope: - **GitHub Releases publish pipeline.** The artifacts produced by `package` are local; uploading them as a `mann1x/opencoti` release asset (and flipping the pin away from claude-hooks) is the M7 ticket. - **macOS / Metal backend.** `metal.c` is macOS-only; the M6 pipeline is Linux-first. Adding macOS needs a Darwin CI runner. - **Cross-arch builds.** cosmocc APE is dual-arch by design, but ggml-*.so are arch-specific. The `` token in the output filename is anticipatory; M6 ships x86_64 only. - **`build:llamafile:cuda --cublas`.** Mozilla-Ocho's cuda.sh optionally links against vendor cuBLAS instead of TinyBLAS. F4 defaults to TinyBLAS (works on every NVIDIA GPU without runtime CUDA install); cuBLAS is a follow-up if profile-guided choices ever need it. ## Verification checklist After a successful `package` run, the manifest tells you exactly what was embedded. To cross-check from the shell: ```bash # Confirm DSOs are inside the bare artifact. unzip -l dist/llamafile/opencoti-llamafile-0.10.1-x86_64.llamafile | grep "ggml-" # Confirm the manifest matches the artifact's actual SHA. sha256sum dist/llamafile/opencoti-llamafile-0.10.1-x86_64.llamafile jq -r .sha256 dist/llamafile/opencoti-llamafile-0.10.1-x86_64.llamafile.MANIFEST.json ``` The embedder smoke at `packages/opencoti-embedder/script/smoke.ts` runs the full launcher against whatever asset the downloader resolves. Set `OPENCOTI_EMBEDDER_LOCAL` to point at a freshly-built composite to exercise the new patches without a release upload. ## Incremental rebuilds, header dependencies, and ccache Iterating the vendored patch series (F4/F5) means editing C++ inside `vendors/sources/llamafile` and rebuilding with `bun run build:llamafile:make` (the incremental path — it does **not** reset the tree, so in-flight edits survive). Two build-infra facts make that loop work; both bit hard before being fixed (see buglog `bug-170`). ### Header dependency tracking (`0007-build-header-deps.patch`) The upstream build resolves header deps via cosmopolitan's `mkdeps` (`build/deps.mk` → `o//depend`). `mkdeps` only follows cosmopolitan's root-relative include style; llama.cpp sources use `#include "foo.h"` resolved through `-iquote` dirs, which `mkdeps` cannot follow — so it emitted **zero** dependency edges for `llama.cpp/src` headers. Editing a header therefore never marked its dependent objects stale, and `make` reported them "up to date" → the only way to push a header change through was to delete objects by hand and rebuild from scratch. Patch `0007` adds `-MMD -MF $@.d` to the `COMPILE.c`/`COMPILE.cc` commands so cosmocc emits an accurate per-translation-unit `.o.d` (project headers only), and `-include`s those files in `build/deps.mk`. clang resolves the real `-iquote` includes, so the header set is exact. A header edit now recompiles **exactly** its dependents (both fat-object arches, via the existing rule) and nothing else. The explicit `-MF $@.d` is **load-bearing, not cosmetic** — and it's what made ccache start working at all (see below). cosmocc defaults the dep file to `.d` by *appending* (`foo.cpp.o` → `foo.cpp.o.d`), but ccache derives the expected name from `-o foo.cpp.o` by *stripping* `.o` (→ `foo.cpp.d`). With bare `-MMD` the two disagree: ccache `stat`s a file cosmocc never wrote, logs *"failed to stat …foo.cpp.d"*, and aborts every compile with *"ccache internal error"* → **nothing is ever cached** (the cache sat at 0 files). Pinning `-MF` to cosmocc's actual name makes them agree, and ccache caches/restores normally. > The **first** build after this patch lands is a one-time full compile > that populates the `.o.d` set. Every build after that is correctly > incremental. Verify with: rebuild a TU, `touch` a header it includes → > only its dependents recompile; `touch` a header it does *not* include → > nothing recompiles. ### ccache + cosmocc fat objects ccache (enabled by default; cache at `.opencoti/ccache/`, disable with `OPENCOTI_NO_CCACHE=1`) wraps cosmocc's **fat** driver: one `cosmoc++ -c -o foo.o` writes **both** `foo.o` (x86-64) and `.aarch64/foo.o`, but ccache only knows the x86 `-o` output. So a cache **hit restores the x86 object only — never its `.aarch64` twin**. A hit is therefore safe only when the twin already exists on disk. Consequences, and how they're handled: - **Never delete a single-arch object** to force a rebuild. If you must force one, `touch` its source (or delete *both* `foo.o` **and** `.aarch64/foo.o`). Deleting only the x86 object and rebuilding makes ccache restore it while leaving no aarch64 twin → the APE link fails with *"concomitant .aarch64 file missing"*. With header tracking fixed, manual deletion is no longer needed at all. - **The full `build` resets the tree** (`git clean -fdx` wipes `o/`). A warm cache restoring x86-only objects into an empty tree would orphan every aarch64 twin. So `runBuild` sets `CCACHE_RECACHE=1`: cosmocc re-runs and writes both arches while still refreshing the cache for the next incremental `:make`. The incremental path runs against a populated `o/`, where hits are safe, so it uses the cache normally. ## Troubleshooting — DSO load & symbol resolution These two cost ~1h during M6-S2 (turbo3 KV). Both masquerade as logic bugs. **Do the diagnostic-first steps before touching any source.** ### 1. A CUDA change has "no effect" → confirm which `.so` is actually loaded The dev binary probes for `ggml-cuda.so` at **two** paths and loads the first that resolves: 1. `~/.llamafile/v//ggml-cuda.so` — the cache. `build:llamafile:cuda` mirrors the freshly-built DSO **here only**. 2. `$HOME/ggml-cuda.so` — `cuda.sh`'s default `OUTPUT=${HOME}/ggml-cuda.so`. A stale `$HOME/ggml-cuda.so` (left from an old manual `cuda.sh` run) silently shadows every rebuild that only updated the cache; the boot log gives no hint. Before theorizing about source/object/enum, confirm the loaded inode: ```sh pid=$(pgrep -f 'llamafile.*--server' | head -1) grep ggml-cuda.so /proc/$pid/maps # which path is mapped stat -c '%i %n' /root/ggml-cuda.so vendors/dist/llamafile/cuda/*/*/ggml-cuda.so ``` If the loaded inode isn't the file you just built, sync it (`cp` the fresh DSO over `$HOME/ggml-cuda.so`, or `rm` it to force the cache path). ### 2. "no usable GPU found" with a healthy GPU == DSO `dlopen` failure The dlopen error is swallowed at default verbosity. **Re-run with `--verbose`** to get the real line, e.g. `cuda: INFO: failed to load library .../ggml-cuda.so: undefined symbol: dequantize_row_turbo2_0`. ### Adding a new `ggml/src/*.c` core source — it must land in THREE places The DSO is **self-contained**: `llamafile/build-functions.sh` `ggml_core_sources` compiles `ggml.c` + `ggml-quants.c` into the `.so` as `ggml-core-*.o`. `ggml.c`'s `type_traits` table holds **data-relocation function pointers** to every type's CPU `to_float`/`from_float` impl — resolved **eagerly at dlopen**, so a missing impl is an undefined-symbol *load* failure even with `RTLD_LAZY`. A new core `.c` (e.g. `ggml-turbo-quant.c`) must therefore be added to: 1. `llamafile/build-functions.sh` → `ggml_core_sources` (DSO self-compiles it). 2. `llama.cpp.patches/llamafile-files/BUILD.mk` → `GGML_SRCS_C` (host links it). 3. **Force the host relink** — `make` does *not* propagate a `GGML_SRCS_C` addition into the final binary. `rm o/llamafile/llamafile o/llama.cpp/llama.cpp.a`, rebuild, and verify the binary mtime advanced. If a previously-working file goes missing from these lists, suspect a `build:llamafile:reset` having wiped outer-vendor edits ([DESTRUCTIVE_OPS.md](../protocols/DESTRUCTIVE_OPS.md)). ## #613 — the packaged engine as a zero-dependency single-file product Shipped 2026-07-05. The `package` subcommand's bare artifact is now the PRIMARY deliverable of the llamafile lane: a self-contained APE that runs the full opencoti inference engine (llama-server OpenAI-compatible HTTP API + CLI) with **zero external dependencies** — no opencoti framework, no side files. ### Versioning `opencoti--`, e.g. `opencoti-0.10.3-c2`. The tag lives in the committed `vendors/llamafile/OPENCOTI_TAG` (comment-tolerant, same format as the pin file; bump on every packaged cut whose binary or DSO bytes change). C-series lineage: `c1` = the un-packaged split deploy on bs2 (thin binary + external side-loaded DSO at `/srv/ml/opencoti-c1/`), `c2` = first packaged single-file cut. The MANIFEST.json records `versionString`, `opencotiTag`, `gitCommit`, the full patch list, and the sha256 of the artifact AND of every embedded DSO — the binary⇄DSO pairing is self-verifying (motivated by bug-937: a deleted external DSO silently fell back to CPU). ### Zero-dependency guarantee (what's compiled in) - **DSO**: zipalign-embedded; on first boot the APE extracts it to `~/.llamafile/v//` and side-loads from there. A fresh `$HOME` is sufficient — proven by the c2 smoke on bs2 (scratch `HOME`, extracted DSO sha == manifest sha `eec7436f…`, Blackwell GPU decode at 272.8 tps, dual-ctx MTP engaged at accept 0.728). - **TCQ codebooks**: decode AND encode use compiled-in `__constant__` codebooks in the DSO (`fattn-tcq.cuh` / `turbo-tcq-cuda.cuh`). The `TURBO_TCQ_CB` env is an opt-in override for codebook experiments only; if set but unreadable it warns on stderr and keeps the compiled-in default on BOTH encode and decode (consistent, no KV split). Exotic `turbo*_tcq` KV tiers therefore need no side file. Same story for `TURBO_INNERQ_*` — all opt-in calibration hooks, inert unless set. - **All opencoti features** are flag/env-activated on the standalone binary (`--spec-type`, `-ctk/-ctv` turbo tiers, `--dca-*`, `--vram-target`, `--kv-residency-mode`, …); nothing requires the TS/Bun or Go framework. ### Launch caveat (cross-platform) The APE is cosmocc cross-platform (Linux/Mac/Windows) but GPU support is CUDA-only today (Metal is a known gap; `runRocm`/`runVulkan` exist in the pipeline but aren't part of the c2 cut). On Linux without binfmt_misc, `exec` of the APE fails ENOEXEC (bug-156) — launch as `sh ./opencoti-llamafile-….llamafile`. Mac/Windows have not been smoke-tested for c2 (Linux-only exercised so far). ### Deploys The c2 artifact is deployed next to the split c1 binaries at `bs2:/srv/ml/opencoti-c1/opencoti-llamafile-0.10.3-c2-x86_64.llamafile`. Eval scripts still drive the split `llamafile.dualctx` + external DSO (fine for dev inner loop — a raw source edit only needs make, not packaging); the packaged artifact is the distribution/embedding vehicle. For packaged deploys the bug-224 DSO-cache-clear ritual is unnecessary: each cut extracts its own DSO keyed by version dir, and clearing `~/.llamafile/v` is harmless. ### Secondary native-ELF lane A plain-gcc/CMake build of the vendored `llama.cpp` tree (standard rpath `libggml-cuda.so`, clean gdb/nsys) for Linux/container embedders. Only one patch in the chain is cosmo-specific (`0071-state-io-cosmocc`), and the tree carries upstream CMakeLists (models/*.cpp is globbed), so the source is ~toolchain-agnostic. Note the CMake build uses cuBLAS where the cosmocc DSO uses TinyBLAS — a different perf profile (bug-2103/2104 class), so gate perf claims per-toolchain. **SHIPPED 2026-07-05 as patch `0097-native-elf-cmake`** (#619): five CMakeLists surgical hooks (opencoti TUs into the upstream source lists + `-Dmain=llama_server` for the overlay server.cpp) plus a CMake-lane-only stub TU (`ggml-cpu/opencoti-elf-stubs.cpp`) for the cosmocc-bound `llamafile/sgemm.cpp` CPU-fastpath dispatchers — return-false = fall back to upstream generic paths (Mozilla's own `fa_helpers_unsupported` contract), so the ELF build loses only CPU fastpaths, never correctness. Build (bs2-proven): ```sh cmake -S llama.cpp -B build-elf -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON \ -DCMAKE_CUDA_ARCHITECTURES=120 -DCMAKE_BUILD_TYPE=Release \ -DLLAMA_CURL=OFF -DGGML_NATIVE=OFF cmake --build build-elf -j24 --target llama-server ``` `-DGGML_CUDA_FA_ALL_QUANTS=ON` is REQUIRED: upstream's CMake only globs `fattn-vec*.cu` (all turbo/TCQ FA-VEC instances + codebook setters) under that flag; the cosmo build always sets it. Gate (bs2, RTX 6000 Blackwell): A4B `-ngl 99` boots, dual-ctx assistant-MTP accept 0.764, **310.2 tps** decode (vs 274 on the APE — native-arch cuBLAS build), coherent chat output, `ldd` shows rpath `libggml-cuda.so.0`. The cosmocc lane is untouched (BUILD.mk lists sources explicitly and never reads CMakeLists).