ManniX-ITA's picture
c3: release notes -win variant
4e674e8 verified
|
Raw
History Blame
4.7 kB

opencoti llamafile 0.10.3-c3

Third packaged cut (c series) of the opencoti single-file inference engine (#613 lane). Built from Mozilla-Ocho/llamafile v0.10.3

  • the opencoti patch series (79 patches, 00010133) at vendors/patches/llamafile/.

Version string: opencoti-0.10.3-c3 — printed by --version and exposed at runtime.

Artifacts

File Size sha256
opencoti-llamafile-0.10.3-c3-x86_64.llamafile 4,812,296,839 B 8a35b1f87b7bb7f2915d8756450f09008f4cd8331cf4a763e40edbcb8c4cf7e9
opencoti-llamafile-0.10.3-c3-win-x86_64.llamafile.exe 55,429,522 B e5aa9a3d4e8274ca430f3a5557708955821cbc2d6d2c20f36a1c24228b0c642d
dso/0.10.3-c3/ggml-cuda-x86_64.so (side-load DSO, HF only) 4,756,838,104 B 5e0342ba9b2ded4847280067dfd6ebc4257e1b83a984881988c528c163b23e4a

The main artifact's embedded ggml-cuda.so is byte-identical to the standalone side-load DSO above (glibc floor GLIBC_2.29 — runs on Debian 11 / glibc 2.31 and newer).

Single-file Cosmopolitan APE. Fat binary: runs natively on x86_64 and aarch64 (Linux/macOS/Windows/BSD) — on ARM64 hosts (DGX Spark, GH200, Apple Silicon) inference is CPU-only in this cut; the CUDA DSO is x86_64-linux. GPU-enabled aarch64 (sbsa) is a planned separate artifact.

-win variant

The full artifact cannot run on Windows: Windows refuses to execute binaries larger than 4 GB. …-win-x86_64.llamafile.exe is the same host APE (identical patches/features, .exe extension Windows requires) with no embedded GPU backend — 55 MB, CPU inference out of the box on Windows/Linux/macOS/BSD. GPU:

  • Windows: install CUDA Toolkit + MSVC; llamafile compiles ggml-cuda.dll on the fly on first -ngl run (upstream mechanism).
  • Linux x86_64: side-load the standalone DSO: mkdir -p ~/.llamafile/v/0.10.3/ && cp ggml-cuda-x86_64.so ~/.llamafile/v/0.10.3/ggml-cuda.so (verify sha against SHA256SUMS first).

Run as a server:

sh ./opencoti-llamafile-0.10.3-c3-x86_64.llamafile --server --port 8080 \
    -m <model>.gguf -ngl 99 --flash-attn on

On first GPU run the embedded ggml-cuda.so (CUDA 13.2, arches sm_75/80/86/89/90/120) self-extracts to ~/.llamafile/v/0.10.3/. Without --server the binary starts the interactive chat CLI.

What's new vs c2 (2026-07-05)

Packaging:

  • NEW -win variant (Windows-runnable bare APE, GPU via side-load)
    • standalone side-load CUDA DSO published under dso/ on HF
  • Supported/target model families (Gemma-4 primary, Qwen secondary) now documented prominently in README + USAGE

Host binary (patches 0122–0133):

  • RYS --repeat-layers weight-shared layer duplication + --rys-probe v2 task battery (#656–#673)
  • bug-858 CLOSED: MTP speculative decode at/above upstream-b9859 parity on both native pairs (spec-clone-cheap + verify GEMM lift, 0128)
  • P2 mixed-KV auto-tail + parallel≥2 fit/checkpoint fixes (0121, 0130)
  • Runtime introspection (#676/#677, 0132): /props gains an "opencoti" section (kv, residency, dca, speculative, sparse_attn, kv_reuse, repeat_layers, rest_kv); /slots gains per-slot lifetime draft counters + session/pool binding
  • Full --version string opencoti-<ver>-<tag> (#613, 0133)

CUDA DSO (rebuilt, 4.4 GB):

  • #674: DCA-off scalar-MMA tensor-core decode generalized to D256/D512 (Gemma-4 head dims) — q8q4 KV decode ≈1.15× f16, up to 2.3× native VEC
  • 0128 mmvf f16/bf16 TinyBLAS verify lift (Qwen NextN n3 parity)
  • 0126/0127 P0 DCA-off scalar-MMA D128 foundation

Verification

MANIFEST.json ships alongside the artifact: versionString, gitCommit, full patch list, and per-DSO sha256 of the embedded backend. Verify the embedded DSO without running:

unzip -p opencoti-llamafile-0.10.3-c3-x86_64.llamafile ggml-cuda.so | sha256sum

Zero-dependency smoke performed on this cut: isolated-HOME --version, DSO self-extraction sha-match, --server GPU boot on RTX 3090 with /props opencoti introspection + greedy completion.

Build provenance

  • Repo: mann1x/opencoti @ commit 4c3f2ec2e (dev branch)
  • Base: llamafile 0.10.3 (outer 7fca8b2, nested llama.cpp dbe9c0c) + Mozilla patch overlay + opencoti chain 0001–0133 (replay-proven byte-identical to the shipped tree)
  • Host binary built on solidPC (cosmocc); CUDA DSO built on solidPC (Debian 11, CUDA 13.2) to pin the release glibc floor at 2.29 (bug-2190: a bs2-built DSO required glibc ≥2.32 and failed on older hosts)

License

Apache-2.0 (Mozilla-Ocho/llamafile) + MIT (llama.cpp and bundled projects). opencoti patches and packaging are MIT.