Text Generation
GGUF
llama.cpp
ternary
2-bit
llama-cpp
cuda
metal
on-device
hybrid-attention
prismml
bonsai
conversational
Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
−2 logit bias on PTQ1_0, RTX 5060 Laptop: MATH-500 44/50 → 43/50
#83 opened 2 days ago
by
7dollarbooks
Add RTX 3080 (10 GB) to the cross-platform throughput table
#82 opened 2 days ago
by
Siphion
Works on Tesla V100-16GB (SM 7.0): 52 tok/s @ 64K context
#81 opened 3 days ago
by
Vyach59
nobody told us, but Pelikan can be ANIMATEEED!
1
#80 opened 3 days ago
by
KottCh
could gemma4-31B-it be better?
👍 1
#79 opened 5 days ago
by
luckystar199746
Intel Arc: a faster SYCL path for PTQ1_0 (about 88 t/s with MTP on a B580)
1
#78 opened 5 days ago
by
torchit1
Sitting frozen like a nuclear bomb, the 8G laptop speeds up from 50 tokens straight to 70 tokens,8G再次从50Token/S大幅升到恐怖的70Token/s
3
#75 opened 6 days ago
by
SuperLogic
I made a ternary 35B-A3B based on the results of my previous Bonsai-2 forensics & research work.
🔥 4
#74 opened 7 days ago
by
SkyIsNotGreen
Think too much
1
#73 opened 7 days ago
by
jezzza1401
Which Jinja Template?
1
#70 opened 9 days ago
by
Omnomynous
image to SVG test: Ichigo :)
🔥 3
3
#69 opened 9 days ago
by
KottCh
Does not work in LM Studio (NB - update CUDA, use NOXZURE1's answer)
10
#68 opened 9 days ago
by
SoulFireMage
Memory performance
👀 1
2
#67 opened 9 days ago
by
trombonator
One 27B model, two jobs — chat:50token/s + JEV decision engine:0.5s — on an 8 GB laptop
🧠 3
#65 opened 10 days ago
by
SuperLogic
Can you create Ternary versions for 35b-a3b models?
#64 opened 10 days ago
by
naildirect
Pull coding capability, the Tetris.html which Bonsai-2-27b generated cannot play - console error
#63 opened 10 days ago
by
mind0n
I took Bonsai 2 27B apart, part 2: the format is public, the last 8% is ordinary ternary QAT.
👀 1
6
#62 opened 10 days ago
by
SkyIsNotGreen
Set the model name and min_p in the GGUF metadata (weights unchanged)
👍 1
#59 opened 11 days ago
by
bri-prism
for 16gb vram context size <=114688 is fast. Pelican is gorgeous :)
👍🤯 4
8
#56 opened 11 days ago
by
KottCh
intel alchemist kernel
#55 opened 11 days ago
by
Demilenos
Real KLD testing against Unsloth's gguf BF16 of Qwen 3.8-27B
👀👍 3
3
#54 opened 11 days ago
by
mrumel
When will they release the official version with ROCM or Vulkan support for the 6900XT?
1
#53 opened 12 days ago
by
aidoluiz
AMD works just wonderfully, here is how:
#52 opened 12 days ago
by
RegisteredWednesday
issue : 6800 (rdna 2) gpu card on PrismML-Eng llama.cpp . cant run model by unsloth Studio and ..
👍 1
#51 opened 12 days ago
by
myhugginfacegacc
SYCL backend: any speculative type collapses performance (even target prefill drops ~200x) - draft model itself is healthy
#50 opened 13 days ago
by
Yoo00ooOO
How does the 2080 Ti 22G perform with this model?
2
#49 opened 13 days ago
by
coresen
PLEASE stop lying about the "intelligence" of the model.
🔥👍 24
7
#47 opened 13 days ago
by
Splarkszter
在4060笔记本,8G下,优化到45Token/S之后,再提10%到50Token/S,但这还不是上限...理论上限可能高达65Token/S
👍 2
1
#46 opened 14 days ago
by
SuperLogic
Solved "xhigh" loop and include MTP. Test on 4080 12G VRAM laptop, reach ~60 t/s
👍🚀 3
#45 opened 14 days ago
by
zhijin123
Built a multi agent orchestrator with it
🚀 2
#43 opened 14 days ago
by
anubhav200
dspark will improve intelligence/bit
#42 opened 14 days ago
by
john1248
3060 rtx 8 gb vram here ... share your preset !
4
#41 opened 14 days ago
by
oytaub
First time used... Failed basic tool calls immediately. (Q2 variant)
👍 5
1
#40 opened 14 days ago
by
laser50
在4060笔记本,8G下,优化到45Token/S,已经达到大厂的Token接口速度,彻底实现Token自由,大家再也不需要去大厂订阅。
6
#39 opened 14 days ago
by
SuperLogic
Ternary-Bonsai-2-27B on an RTX 2060 SUPER 8GB — Japanese supplied-context reasoning is the biggest surprise
1
#38 opened 14 days ago
by
mktnhr
LM Studio says no
4
#37 opened 14 days ago
by
Vort
No sure why all the hate?
❤️ 2
#36 opened 15 days ago
by
Mogsie
It is fighting with the harness, chat template and blaming the human operator
10
#35 opened 15 days ago
by
gbuzhf
OMG
#34 opened 15 days ago
by
darkmatter2222
Density table divides Bonsai by the ideal 5.80 GB; competitors by shipped files
#33 opened 15 days ago
by
makerportal
Really useful quantified small models PQ2_0. Have good intelligence and pretty fast. (With evalscope testing result)
🧠🚀 4
#31 opened 15 days ago
by
zhijin123
confused on sampling parameters
👀 1
#30 opened 15 days ago
by
Jcamacho05
Bonsai 2 27B Uncensored
3
#29 opened 15 days ago
by
e-RHM-e
Joining your group
1
#28 opened 15 days ago
by
Mohammedkarimi
Surprisingly good PTQ1_0 quantization – and ~39.5 tok/s on an RTX 4070
👍🔥 7
#27 opened 15 days ago
by
TheWegemann
出一个Q6的试试?看看能力能不能保留FP16的99.99%?
1
#26 opened 15 days ago
by
kelei999999
Update README.md
1
#25 opened 15 days ago
by
rikunarita-3
Something is wrong
6
#24 opened 16 days ago
by
kashish4u