license: apache-2.0
tags:
- llamafile
- llama.cpp
- gguf
- inference-engine
- cuda
- opencoti
opencoti-llamafile
Self-contained single-file inference engine from the opencoti project — a Mozilla-Ocho llamafile 0.10.3 base carrying the opencoti patch series: advanced KV residency/quantization (PolyKV/TurboQuant/TCQ), rolling-KV window with host-RAM spill, DCA long-context extension, MTP speculative decode, sparse attention, RYS layer duplication, and a runtime introspection/control API.
Start with USAGE.md — it explains exactly how this engine diverges from upstream llamafile, every added feature, its flags, defaults, limitations, and which features compose.
This repo hosts the packaged release artifacts (they exceed GitHub's
2 GiB release-asset cap) plus the full patch series under
patches/ and the project documentation under
docs/. Release notes and per-release SHA256SUMS live
under releases/ and are mirrored on the corresponding
GitHub release at
mann1x/opencoti.
Artifacts
| Artifact | Platform | GPU |
|---|---|---|
opencoti-llamafile-<ver>-<tag>-x86_64.llamafile |
x86_64 + aarch64 (fat APE; Linux/macOS/Windows/BSD) | CUDA 13.2, x86_64-linux (sm_75/80/86/89/90/120f) |
opencoti-llamafile-<ver>-<tag>-aarch64.llamafile (planned) |
same fat APE | CUDA 13.2, sbsa/aarch64 (sm_110f Jetson Thor, sm_121a DGX Spark GB10) |
The APE host binary runs natively on both x86_64 and aarch64; the two
artifacts differ only in the embedded ggml-cuda.so. On a platform
with no matching DSO, inference falls back to CPU.
Usage
chmod +x opencoti-llamafile-*.llamafile
# server mode
sh ./opencoti-llamafile-*.llamafile --server --port 8080 \
-m <model>.gguf -ngl 99 --flash-attn on
# without --server it starts the interactive chat CLI
On the first GPU run the embedded CUDA DSO self-extracts to
~/.llamafile/v/<ver>/. --version prints the opencoti version
string; GET /props exposes an opencoti introspection section at
runtime (see USAGE.md §introspection).
Verification
Each artifact ships with a MANIFEST.json recording the artifact
sha256, versionString, gitCommit, the full patch list, and the sha256
of every embedded backend DSO; releases/<tag>/SHA256SUMS is the
committed contract (CI verifies this repo's LFS blobs against it on
every release tag). Verify the embedded DSO without running:
unzip -p opencoti-llamafile-*.llamafile ggml-cuda.so | sha256sum
License
Apache-2.0 (llamafile) + MIT (llama.cpp and bundled projects). opencoti patches and packaging are MIT.