opencoti llamafile 0.10.3-c3
Third packaged cut (c series) of the opencoti single-file inference
engine (#613 lane). Built from
Mozilla-Ocho/llamafile v0.10.3
- the opencoti patch series (79 patches,
0001–0133) atvendors/patches/llamafile/.
Version string: opencoti-0.10.3-c3 — printed by --version and exposed
at runtime.
Artifacts
| File | Size | sha256 |
|---|---|---|
opencoti-llamafile-0.10.3-c3-x86_64.llamafile |
4,812,296,839 B | 8a35b1f87b7bb7f2915d8756450f09008f4cd8331cf4a763e40edbcb8c4cf7e9 |
opencoti-llamafile-0.10.3-c3-win-x86_64.llamafile.exe |
55,429,522 B | e5aa9a3d4e8274ca430f3a5557708955821cbc2d6d2c20f36a1c24228b0c642d |
dso/0.10.3-c3/ggml-cuda-x86_64.so (side-load DSO, HF only) |
4,756,838,104 B | 5e0342ba9b2ded4847280067dfd6ebc4257e1b83a984881988c528c163b23e4a |
The main artifact's embedded ggml-cuda.so is byte-identical to the
standalone side-load DSO above (glibc floor GLIBC_2.29 — runs on
Debian 11 / glibc 2.31 and newer).
Single-file Cosmopolitan APE. Fat binary: runs natively on x86_64 and aarch64 (Linux/macOS/Windows/BSD) — on ARM64 hosts (DGX Spark, GH200, Apple Silicon) inference is CPU-only in this cut; the CUDA DSO is x86_64-linux. GPU-enabled aarch64 (sbsa) is a planned separate artifact.
-win variant
The full artifact cannot run on Windows: Windows refuses to execute
binaries larger than 4 GB. …-win-x86_64.llamafile.exe is the same
host APE (identical patches/features, .exe extension Windows
requires) with no embedded GPU backend — 55 MB, CPU inference out
of the box on Windows/Linux/macOS/BSD. GPU:
- Windows: install CUDA Toolkit + MSVC; llamafile compiles
ggml-cuda.dllon the fly on first-nglrun (upstream mechanism). - Linux x86_64: side-load the standalone DSO:
mkdir -p ~/.llamafile/v/0.10.3/ && cp ggml-cuda-x86_64.so ~/.llamafile/v/0.10.3/ggml-cuda.so(verify sha againstSHA256SUMSfirst).
Run as a server:
sh ./opencoti-llamafile-0.10.3-c3-x86_64.llamafile --server --port 8080 \
-m <model>.gguf -ngl 99 --flash-attn on
On first GPU run the embedded ggml-cuda.so (CUDA 13.2, arches
sm_75/80/86/89/90/120) self-extracts to ~/.llamafile/v/0.10.3/.
Without --server the binary starts the interactive chat CLI.
What's new vs c2 (2026-07-05)
Packaging:
- NEW
-winvariant (Windows-runnable bare APE, GPU via side-load)- standalone side-load CUDA DSO published under
dso/on HF
- standalone side-load CUDA DSO published under
- Supported/target model families (Gemma-4 primary, Qwen secondary) now documented prominently in README + USAGE
Host binary (patches 0122–0133):
- RYS
--repeat-layersweight-shared layer duplication +--rys-probev2 task battery (#656–#673) - bug-858 CLOSED: MTP speculative decode at/above upstream-b9859 parity on both native pairs (spec-clone-cheap + verify GEMM lift, 0128)
- P2 mixed-KV auto-tail + parallel≥2 fit/checkpoint fixes (0121, 0130)
- Runtime introspection (#676/#677, 0132):
/propsgains an"opencoti"section (kv, residency, dca, speculative, sparse_attn, kv_reuse, repeat_layers, rest_kv);/slotsgains per-slot lifetime draft counters + session/pool binding - Full
--versionstringopencoti-<ver>-<tag>(#613, 0133)
CUDA DSO (rebuilt, 4.4 GB):
- #674: DCA-off scalar-MMA tensor-core decode generalized to D256/D512 (Gemma-4 head dims) — q8q4 KV decode ≈1.15× f16, up to 2.3× native VEC
- 0128 mmvf f16/bf16 TinyBLAS verify lift (Qwen NextN n3 parity)
- 0126/0127 P0 DCA-off scalar-MMA D128 foundation
Verification
MANIFEST.json ships alongside the artifact: versionString, gitCommit,
full patch list, and per-DSO sha256 of the embedded backend. Verify the
embedded DSO without running:
unzip -p opencoti-llamafile-0.10.3-c3-x86_64.llamafile ggml-cuda.so | sha256sum
Zero-dependency smoke performed on this cut: isolated-HOME --version,
DSO self-extraction sha-match, --server GPU boot on RTX 3090 with
/props opencoti introspection + greedy completion.
Build provenance
- Repo:
mann1x/opencoti@ commit4c3f2ec2e(dev branch) - Base: llamafile 0.10.3 (outer 7fca8b2, nested llama.cpp dbe9c0c) + Mozilla patch overlay + opencoti chain 0001–0133 (replay-proven byte-identical to the shipped tree)
- Host binary built on solidPC (cosmocc); CUDA DSO built on solidPC (Debian 11, CUDA 13.2) to pin the release glibc floor at 2.29 (bug-2190: a bs2-built DSO required glibc ≥2.32 and failed on older hosts)
License
Apache-2.0 (Mozilla-Ocho/llamafile) + MIT (llama.cpp and bundled projects). opencoti patches and packaging are MIT.