Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Real KLD testing against Unsloth's gguf BF16 of Qwen 3.8-27B
Ternary Bonsai 2 27B vs Qwen3.8-27B quants: benchmark summary
Hardware: Tesla PG500-216 (V100-class, Volta, sm_70, 32 GB), Dell PowerEdge, Unraid
Runtime: PrismML llama.cpp fork, build bdc23b56b (10728), CUDA 12.8.1, CUDA_DOCKER_ARCH=70
Models tested:
Qwen3.8-27B-BF16(Unsloth): referenceQwen3.8-27B-Q8_0(Unsloth)Qwen3.8-27B-UD-Q4_K_XL(Unsloth)Ternary-Bonsai-2-27B-PQ2_0(PrismML): 2.13 bpw, group-128 ternaryTernary-Bonsai-2-27B-PTQ1_0(PrismML): 1.75 bpw dense ternary (speed only)
Summary
Q8_0 is effectively identical to BF16, and Q4_K_XL is an excellent quant (within about 1% on every measure). Bonsai PQ2_0 is about 38% faster at generation and uses 10 GB less VRAM than Q4_K_XL, but it diverges from the original model far more than normal quantization does: roughly 40× the Q4's KL divergence, and it picks a different top token about 1 in 4.5 times.
Bonsai is a retrained ternary model, not a rounded copy of Qwen3.8-27B, so these numbers show it behaves differently from the original, not necessarily worse on real tasks. Chat-template task evals (gsm8k, mmlu_pro) are the remaining test.
Current pick: Q4_K_XL as the default for fidelity; Bonsai PQ2_0 only where speed or VRAM matters more than matching Qwen's behavior. Avoid PTQ1_0 on Volta.
Speed (llama-bench)
-ngl 99 -fa 1 -p 512 -n 128 -d 0,8192 -r 5
| Model | Size | pp512 (t/s) | tg128 (t/s) | pp512 @ 8K | tg128 @ 8K |
|---|---|---|---|---|---|
| Qwen3.8 Q4_K_XL | 16.34 GiB | 751.6 ± 37.2 | 36.09 ± 0.06 | 654.2 ± 33.0 | 34.04 ± 0.17 |
| Bonsai PQ2_0 | 6.70 GiB | 700.7 ± 46.2 | 49.87 ± 0.28 | 634.3 ± 16.3 | 46.49 ± 0.21 |
| Bonsai PTQ1_0 | 5.53 GiB | 757.8 ± 22.1 | 32.89 ± 0.16 | 664.8 ± 15.4 | 31.42 ± 0.14 |
Notes:
- Prompt processing is roughly tied across all three; decode speed is the only real difference.
- Bonsai PQ2_0 is 2.4× smaller than Q4_K_XL but only 1.38× faster. Effective decode bandwidth is about 335 GB/s vs about 590 GB/s for Q4_K, so the ternary kernels are unpacking-bound on Volta rather than memory-bound.
- PTQ1_0's denser trit packing costs more ALU per weight, making it slower than both others on this card. Its only advantage is 1.2 GB less VRAM.
- Both models slow down very little at 8K context (hybrid attention).
Perplexity (wikitext-2 test, full 145 chunks, -c 2048)
| Model | PPL |
|---|---|
| Qwen3.8 Q4_K_XL | 6.3439 ± 0.0404 |
| Bonsai PQ2_0 | 8.3536 ± 0.0576 |
Bonsai is about 32% higher. All models share the same tokenizer (gpt2 / qwen35, 248,320 tokens, BOS 248044, EOS 248046), so these are directly comparable.
KL divergence vs BF16 (wikitext-2, 50 chunks, -c 2048)
Reference logits from BF16 with partial offload (-ngl 34, about 26.9 GB VRAM, 51.5 s/chunk, about 43 min). BF16 PPL over these 50 chunks: 6.1345 ± 0.0654.
| Metric | Q8_0 | Q4_K_XL | Bonsai PQ2_0 |
|---|---|---|---|
| Mean PPL | 6.1396 | 6.1003 | 8.0420 |
| PPL ratio vs BF16 | 1.0008 | 0.9944 | 1.3110 |
| Cor(ln PPL) vs BF16 | 99.96% | 99.73% | 92.16% |
| Mean KLD | 0.000644 | 0.008685 | 0.340281 |
| Median KLD | 0.000101 | 0.001830 | 0.175541 |
| 90% KLD | 0.000467 | 0.010595 | 0.752498 |
| 95% KLD | 0.000784 | 0.019142 | 1.185927 |
| 99% KLD | 0.002857 | 0.065884 | 2.886118 |
| 99.9% KLD | 0.041576 | 0.558065 | 7.351860 |
| Max KLD | 4.84 | 21.93 | 21.24 |
| Mean Δp | −0.003% | −0.029% | −4.376% |
| RMS Δp | 0.670% | 2.582% | 15.952% |
| Same top token | 99.284% | 96.882% | 77.769% |
Interpretation
- Q8_0: Indistinguishable from BF16 for practical purposes (mean KLD 0.0006, 99.3% top-token agreement).
- Q4_K_XL: A strong quant. Its slightly lower PPL than BF16 is a common artifact, not an improvement; KLD shows small real drift. Greedy output differs from BF16 about once every 32 tokens.
- Bonsai PQ2_0: Divergence is broad, not limited to a few outlier tokens: its median KLD (0.18) exceeds the Q4's 99th percentile (0.066). It disagrees with BF16's top pick about once every 4.5 tokens. Its negative mean Δp shows it is systematically less confident in the correct next token. The lower per-chunk PPL correlation (92%) means it finds different passages hard, consistent with a retrained sibling model rather than a lossy copy.
- Ratios: Bonsai mean KLD is about 39× Q4_K_XL and about 530× Q8_0.
Caveats
- Perplexity and KLD score raw Wikipedia text without a chat template or reasoning. They measure closeness to Qwen3.8-27B's distribution, not task ability.
- PrismML's "~98% of benchmark performance" claim is based on task benchmarks in thinking mode, which this suite doesn't test.
- Bonsai's kernels are not tuned or benchmarked by PrismML on Volta; speed results may not reflect newer GPUs.
- The BF16 base run disabled fused Gated Delta Net ops because of the CPU/GPU split. This affects speed only, not results.
Next steps
- Run task evals through
llama-serverwith the chat template, e.g.lm_eval --model local-chat-completions --tasks gsm8k,mmlu_pro --apply_chat_template --limit 200, with matching sampling and reasoning effort for Q4_K_XL and Bonsai PQ2_0. - Optionally, HellaSwag / Winogrande via
llama-perplexity --hellaswag/--winograndefor quick accuracy-based checks.
Commands used
# prism helper
prism() { docker run --rm --init --gpus '"device=1"' -v /mnt/user/appdata/llama_cpp/model:/models local/llama.cpp-prism:full "$@"; }
# speed
prism --bench -m <model> -ngl 99 -fa 1 -p 512 -n 128 -d 0,8192 -r 5
# perplexity
prism --perplexity -m <model> -f /models/wiki.test.raw -ngl 99 -c 2048
# KLD base (BF16, partial offload)
prism --perplexity -m <bf16-00001-of-00002.gguf> -f /models/wiki.test.raw -ngl 34 -c 2048 --chunks 50 \
--kl-divergence-base /models/qwen-bf16.kld
# KLD compare
prism --perplexity -m <model> -f /models/wiki.test.raw -ngl 99 -c 2048 --chunks 50 \
--kl-divergence-base /models/qwen-bf16.kld --kl-divergence
Very interesting result. I wonder if the large KL divergence here may connect directly to what was observed in #44.
In #44, plain rotate+RTN already reproduced roughly 92% of the ternary symbols, yet replacing the remaining ~8% with those apparently natural RTN choices caused a catastrophic perplexity collapse. GPTQ/Hessian-based attempts also failed to reproduce that residual structure, which suggested that the important difference was not simply a better post-training quantizer, but trained weight movement before the final ternary projection.
Seen from that perspective, the large KLD here may not be “quantization error” in the usual PTQ sense at all.
It may instead be measuring how far the retrained ternary model has moved into a different low-bit-friendly solution basin while preserving much of the useful task behavior.
That would also explain why the model can differ substantially from BF16 at the token-distribution level, while still retaining surprisingly strong reasoning/coding performance.
Of course, the exact Bonsai 2 training recipe is not public, so this is only an interpretation. But the combination of:
- ~92% symbol agreement with simple rotate+RTN,
- extreme sensitivity to the remaining ~8% in #44,
- failure of ordinary PTQ/Hessian methods to reproduce those choices,
- and the much larger KLD measured here,
seems consistent with QAT/KD or some other training process moving selected weights across ternary decision boundaries rather than merely minimizing local rounding error.
In other words, perhaps the interesting quantity here is not just “how much information was lost by quantization,” but “how far training intentionally moved the model before quantization became cheap.”
That distinction may be important when interpreting Bonsai against conventional Q4/Q8 quantization.
Yes, that is a good point mktnhr.
I have also seen though "real word tasks" do diverge from Q4 quantizations.
Please see linked: https://www.youtube.com/watch?v=NZMUtVPxJvQ
Could you please include the comparison with UD-IQ2_XXS or UD-IQ2_S as well? Since they are in the same "weight class"?



