Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
I made a ternary 35B-A3B based on the results of my previous Bonsai-2 forensics & research work.
Update to my previous post (#62): the MoE side is finished.
Scion-35B-A3B is empero-ai/Qwen3.8-35B-A3B-Distill quantized to 2.61 bpw: 11.34 GB, one file, and it fits a single 20 GB card with room for context.
The recipe: the expert banks go into a ternary PQ2_0 container (2-bit codes, one fp16 scale per 128 weights), and rank-512 low-rank corrections are trained on the residual stream (attention output, MoE block output, plus router deltas) by output-KD against the BF16 teacher, with the deployed quantizer in the loop. The corrections are embedded in the GGUF, so there is no --lora step at load time. Training took about 4,000 steps on rented H100 time ($7.26).
Numbers, all under one protocol (wikitext-2 PPL, KLD against BF16 logits, HellaSwag and Winogrande at 400 tasks):
- HellaSwag 79.00 and Winogrande 76.25, against 80.00 / 76.00 for Q4_K_M and 81.25 / 76.00 for BF16, at about half of Q4_K_M's file size.
- PPL 8.354, the best of the 2-bit class (IQ2_M 8.413, Q2_K 8.473).
- KLD 0.269 mean, still 2-bit-class against Q4_K_M's 0.031. That gap is the open item and the card says so up front.
Three findings from the build worth writing down:
- Rotation is not the MoE lever. Rotated and raw RTN on OLMoE experts came out even (relative error 0.511 vs 0.516). The loss is routing drift, and the corrections have to sit on the residual stream: per-expert branches were 25x worse PPL at 8x the parameters.
- The deployable Lloyd quantizer beats one-shot absmean on MoE experts (0.442 vs 0.517), and the GGUF expert codes reproduce the torch rule to 0.001 relative error.
- Router-KD is a no-op: it matched a zero-weight control within about 1.5%. What recovers router agreement is repairing the state the router reads.
Credits: the container, kernels and the branch this builds on come from Prism ML's llama.cpp and the TAARDIS fork conventions, engine only, no Bonsai weights. Base weights from empero-ai and Qwen; BF16 GGUF conversion from MrFuzzihead. Not affiliated with or endorsed by any of them. The runtime is my own fork of llama.cpp, sky-is-green/prism-ml-llama.cpp branch moe-corr-runtime (embedded adapters, about 90 lines); stock llama.cpp cannot load the container types, and I am happy to upstream the loader bits if that is useful.
If anyone wants to try it or poke holes in the numbers, the model repo's discussions are open, and I am happy to share any part of the recipe in more detail.
Model card: https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B