| # opencoti llamafile 0.10.3-c3 |
|
|
| Third packaged cut (`c` series) of the opencoti single-file inference |
| engine (#613 lane). Built from |
| [Mozilla-Ocho/llamafile v0.10.3](https://github.com/Mozilla-Ocho/llamafile/releases/tag/0.10.3) |
| + the opencoti patch series (79 patches, `0001`–`0133`) at |
| `vendors/patches/llamafile/`. |
|
|
| Version string: `opencoti-0.10.3-c3` — printed by `--version` and exposed |
| at runtime. |
|
|
| ## Artifacts |
|
|
| | File | Size | sha256 | |
| |---|---:|---| |
| | `opencoti-llamafile-0.10.3-c3-x86_64.llamafile` | 4,812,296,839 B | `8a35b1f87b7bb7f2915d8756450f09008f4cd8331cf4a763e40edbcb8c4cf7e9` | |
| | `opencoti-llamafile-0.10.3-c3-win-x86_64.llamafile.exe` | 55,429,522 B | `e5aa9a3d4e8274ca430f3a5557708955821cbc2d6d2c20f36a1c24228b0c642d` | |
| | `dso/0.10.3-c3/ggml-cuda-x86_64.so` (side-load DSO, HF only) | 4,756,838,104 B | `5e0342ba9b2ded4847280067dfd6ebc4257e1b83a984881988c528c163b23e4a` | |
|
|
| The main artifact's embedded `ggml-cuda.so` is byte-identical to the |
| standalone side-load DSO above (glibc floor GLIBC_2.29 — runs on |
| Debian 11 / glibc 2.31 and newer). |
| |
| Single-file Cosmopolitan APE. **Fat binary: runs natively on x86_64 and |
| aarch64** (Linux/macOS/Windows/BSD) — on ARM64 hosts (DGX Spark, GH200, |
| Apple Silicon) inference is CPU-only in this cut; the CUDA DSO is |
| x86_64-linux. GPU-enabled aarch64 (sbsa) is a planned separate |
| artifact. |
|
|
| ### `-win` variant |
|
|
| The full artifact cannot run on Windows: **Windows refuses to execute |
| binaries larger than 4 GB**. `…-win-x86_64.llamafile.exe` is the same |
| host APE (identical patches/features, `.exe` extension Windows |
| requires) with **no embedded GPU backend** — 55 MB, CPU inference out |
| of the box on Windows/Linux/macOS/BSD. GPU: |
|
|
| - **Windows:** install CUDA Toolkit + MSVC; llamafile compiles |
| `ggml-cuda.dll` on the fly on first `-ngl` run (upstream mechanism). |
| - **Linux x86_64:** side-load the standalone DSO: |
| `mkdir -p ~/.llamafile/v/0.10.3/ && cp ggml-cuda-x86_64.so |
| ~/.llamafile/v/0.10.3/ggml-cuda.so` (verify sha against |
| `SHA256SUMS` first). |
| |
| Run as a server: |
| |
| ```bash |
| sh ./opencoti-llamafile-0.10.3-c3-x86_64.llamafile --server --port 8080 \ |
| -m <model>.gguf -ngl 99 --flash-attn on |
| ``` |
| |
| On first GPU run the embedded `ggml-cuda.so` (CUDA 13.2, arches |
| sm_75/80/86/89/90/120) self-extracts to `~/.llamafile/v/0.10.3/`. |
| Without `--server` the binary starts the interactive chat CLI. |
| |
| ## What's new vs c2 (2026-07-05) |
| |
| **Packaging:** |
| - NEW `-win` variant (Windows-runnable bare APE, GPU via side-load) |
| + standalone side-load CUDA DSO published under `dso/` on HF |
| - Supported/target model families (Gemma-4 primary, Qwen secondary) |
| now documented prominently in README + USAGE |
|
|
| **Host binary (patches 0122–0133):** |
| - RYS `--repeat-layers` weight-shared layer duplication + `--rys-probe` |
| v2 task battery (#656–#673) |
| - bug-858 CLOSED: MTP speculative decode at/above upstream-b9859 parity |
| on both native pairs (spec-clone-cheap + verify GEMM lift, 0128) |
| - P2 mixed-KV auto-tail + parallel≥2 fit/checkpoint fixes (0121, 0130) |
| - Runtime introspection (#676/#677, 0132): `/props` gains an `"opencoti"` |
| section (kv, residency, dca, speculative, sparse_attn, kv_reuse, |
| repeat_layers, rest_kv); `/slots` gains per-slot lifetime draft |
| counters + session/pool binding |
| - Full `--version` string `opencoti-<ver>-<tag>` (#613, 0133) |
|
|
| **CUDA DSO (rebuilt, 4.4 GB):** |
| - #674: DCA-off scalar-MMA tensor-core decode generalized to D256/D512 |
| (Gemma-4 head dims) — q8q4 KV decode ≈1.15× f16, up to 2.3× native VEC |
| - 0128 mmvf f16/bf16 TinyBLAS verify lift (Qwen NextN n3 parity) |
| - 0126/0127 P0 DCA-off scalar-MMA D128 foundation |
|
|
| ## Verification |
|
|
| `MANIFEST.json` ships alongside the artifact: versionString, gitCommit, |
| full patch list, and per-DSO sha256 of the embedded backend. Verify the |
| embedded DSO without running: |
|
|
| ```bash |
| unzip -p opencoti-llamafile-0.10.3-c3-x86_64.llamafile ggml-cuda.so | sha256sum |
| ``` |
|
|
| Zero-dependency smoke performed on this cut: isolated-`HOME` `--version`, |
| DSO self-extraction sha-match, `--server` GPU boot on RTX 3090 with |
| `/props` opencoti introspection + greedy completion. |
|
|
| ## Build provenance |
|
|
| - Repo: `mann1x/opencoti` @ commit `4c3f2ec2e` (dev branch) |
| - Base: llamafile 0.10.3 (outer 7fca8b2, nested llama.cpp dbe9c0c) + |
| Mozilla patch overlay + opencoti chain 0001–0133 (replay-proven |
| byte-identical to the shipped tree) |
| - Host binary built on solidPC (cosmocc); CUDA DSO built on solidPC |
| (Debian 11, CUDA 13.2) to pin the release glibc floor at 2.29 |
| (bug-2190: a bs2-built DSO required glibc ≥2.32 and failed on |
| older hosts) |
|
|
| ## License |
|
|
| Apache-2.0 (Mozilla-Ocho/llamafile) + MIT (llama.cpp and bundled |
| projects). opencoti patches and packaging are MIT. |
|
|