Instructions to use prism-ml/Ternary-Bonsai-8B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-8B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-8B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-8B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-8B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-8B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-8B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-8B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-8B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-8B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-8B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-8B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-8B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-8B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-8B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-8B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-8B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-8B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-8B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-8B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-8B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-8B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-8B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-8B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-8B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-8B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-8B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-8B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-8B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-8B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-8B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Q2_0 g64 ternary is ~5Γ slower than Q4_K_M on Samsung S22 CPU - missing optimized kernel?
TL;DR: On a Snapdragon 8 Gen 1 phone (Termux, CPU-only, 4 threads), the Ternary-Bonsai-8B Q2_0 g64 GGUF (2.1 GB) generates at ~3.3 t/s, while LFM2.5-8B-A1B-GGUF model family at Q4_K_M (5.15 GB) does 15β21 t/s - measured the same morning on the same phone.
The 2-bit ternary file is 2.5Γ smaller but ~5Γ slower, which should be impossible if generation were just memory-bandwidth-bound. Output quality on my probes is fine.
My hypothesis: there is no optimized CPU path (repack layout / vecdot kernel) for the Q2_0 g64 block format, so it runs the generic fallback dequant. Is that expected, and is a kernel planned?
Device & setup
| Device | Samsung Galaxy S22 Ultra, Snapdragon 8 Gen 1 |
| CPU | 8Γ Cortex-A (1Γ 3.0 GHz + 3Γ 2.5 GHz + 4Γ 1.8 GHz), 4 threads used for inference |
| RAM | ~11.2 GB usable (12 GB part), 4 GB zram swap |
| OS / runtime | Android 14 (One UI 6), Termux |
| Engine | llama.cpp, custom Android build (Vulkan + CPU backends), commit 790cf51 |
| Ternary flags | -t 4 -c 16384 -b 512 -ub 256 --jinja --reasoning-budget 300 --device none --no-repack |
Both models were benchmarked the same morning, same phone, same thread count. Prompts: a 158-token and a ~550-token prompt, a 37-token math question with reasoning-budget 300, and a haiku request.
Results
| Model (file size) | Prompt t/s | Gen t/s | RSS (anon + file) | Swap used by model |
|---|---|---|---|---|
| LFM2.5-8B-A1B Q4_K_M (5.15 GB) | 33β42 | 15β21 | ~0.3 + 5.1 GB | 0 |
| Ternary-Bonsai-8B Q2_0 g64 (2.1 GB) | 4.0 | 3.2β3.4 | 2.6 + 2.1 GB | 0 |
Notes on fairness:
- The Q4_K_M run used the stock (repacked) layout; the ternary run used
--device none --no-repack. So the ternary got the less optimized path β the gap is, if anything, understated. - The phone was warmer during the ternary run (later in the morning, after more benchmarks), which shaves a few t/s off absolute numbers but doesn't come close to explaining a 5Γ gap.
Why I think it's a kernel issue, not a quant issue
- Speed doesn't scale with size. If token generation were memory-bandwidth-bound (as it is for every standard GGUF on CPU), the 2.1 GB ternary model should be the faster one β it streams ~2.5Γ less weight per token than Q4_K_M. Instead it's ~5Γ slower. That inverts the expected ordering, which points at per-token compute cost, i.e. the dequant/matmul path.
- Prompt processing is ~8β10Γ slow too (4.0 t/s vs 33β42 t/s). Both pp and tg going through the same weight-matrix path being uniformly slow fits "generic fallback kernel" rather than, say, a KV-cache or attention issue.
- Large anonymous allocation at load: the ternary model shows 2.6 GB RssAnon on top of its 2.1 GB file mapping (Q4_K_M shows ~0.3 GB anon). Something is being expanded/dequantized into an anonymous buffer at load time β consistent with a format that isn't being served in its native layout.
--no-repackwas applied to the ternary run. IfQ2_0 g64simply has no repack support and novecdot/matmulspecialization, it lands on the slowest generic code path, while Q4_K_M keeps its optimized block kernels.
Questions
- Is there an optimized CPU kernel (or repack layout) for the
Q2_0 g64ternary block type? If not, is that the expected explanation for the ~5Γ gap? - Is the 2.6 GB anonymous buffer at load expected (e.g. runtime expansion to an intermediate type)? It roughly doubles the effective footprint.
- Any guidance on what a good target is β e.g. should Q2_0 g64 on 4 CPU threads be within ~10β20% of Q4_K_M speed once a proper kernel exists (it should be faster, since it moves less data per token)?
Reproduce
llama-server -m Ternary-Bonsai-8B-Q2_0_g64.gguf \
-t 4 -c 16384 -b 512 -ub 256 --jinja \
--reasoning-budget 300 --device none --no-repack \
--port 8090 --host 127.0.0.1
# then any chat completion, e.g.:
curl -s localhost:8090/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"max_tokens":60,"messages":[{"role":"user","content":"What is the capital of France? Answer in one short sentence."}]}'
# check timings.predicted_per_second
Happy to run further A/B variants (repack on/off, other thread counts, -fa, KV cache types, both models under identical flags) if it helps isolate the bottleneck.
Hi, thanks for the note, yeah your observation makes sense. The Q2_0 is not as optimized for all hardware. Most likely some operations are going on slower fallback kernels.
That being said the other model you are comparing with is a mixture of expert with only 1B activate weights during token generation and that means a lot less operations per token needed, so that could be another reason for speed difference. Our 8B model is dense model so all the 8B weights are active.
My hypothesis: there is no optimized CPU path (repack layout / vecdot kernel) for the
Q2_0 g64block format, so it runs the generic fallback dequant. Is that expected, and is a kernel planned?
Understand when you're working on CPU you'll have a penalty not only from limited number of processors/cores, but also a penalty based on alignment.
For speed purposes CPU's prefer to be aligned by their register size, which is going to be 32bit or 64bit. So something that may take say 10 cycles to grab the operation for a 4byte aligned block, may be 20 cycles longer for a single byte when the address isn't divisible by 4. Add shifts, AND, multiply, then xor and shifting back to put it back into the same sized space before re-writing the byte block, what may be say 50 cycles total for read multiply/add write for 1 bytes could be 10x longer depending on how it is written, nevermind if the instructions can't be run in parallel/out of order (based on unrelated registers having work that doesn't need to wait), which is how a number of modern CPU's work which is an indirect penalty as well. And if the transformer relies on FPU processing which in theory only takes a few cycles, doing FPU instruction calls you have to load the numbers in via instructions, tell it the operation to do, then do an FWAIT instruction before you can get it back, then convert it down vs just using built-in integer instructions.
I don't have the concrete numbers at the moment to give a real example of how many cycles for a single read/process/write but if you're going to run on CPU and you don't have super limited RAM, then probably use the Q8_0 models if possible. Though with Q4 it's shift 4 or AND 0xf to grab the two different halves so that's quite a bit faster.
Most programs compiled will align data in the register/size bounds for speed purposes, even if you lose a lot of space to padding, in which case you have to override structures to follow specific size/types if it's for a legacy process/format.