Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
How does the 2080 Ti 22G perform with this model?
I'm experiencing the 2080 Ti 22G in this model, and would appreciate insights on its performance.
Specifically, I'm looking for:
VRAM usage – Will this model fit within 22 GB of VRAM, and what quantization/precision is recommended?
Inference speed – What performance (e.g. tokens/sec or images/sec) can be expected on this GPU?
Compatibility – Are there any known issues or recommended settings (batch size, attention implementation, etc.) for the 2080 Ti series?
If anyone has run this model on a 2080 Ti 22G before, please share your results or tips in the comments.
For reference, I tested Bonsai-2-27B on a much smaller RTX 2060 SUPER 8GB, so my results may be useful as a baseline for your 2080 Ti 22GB.
My full test and results are here:
https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/discussions/38
The main configuration I used was:
- RTX 2060 SUPER 8GB (Turing, compute capability 7.5)
- CUDA 12.4
- PrismML llama.cpp fork
- full GPU offload (
-ngl 99) - Flash Attention enabled
- 32K context
- K cache:
Q5_1 - V cache:
Q4_0 --parallel 1
With PTQ1_0 (1.75 bpw), the whole model + 32K context fit comfortably in 8GB VRAM.
Observed VRAM usage was about 6.2–6.5GB.
Performance on the 2060 SUPER was:
pp512: ~178 tok/spp2048: ~100 tok/stg128: ~16.8 tok/s- a real long generation (1747-token prompt + 2975 generated tokens): 11.0 tok/s average
Decode speed gradually dropped from about 13.5 tok/s near the beginning to about 9.4 tok/s after ~3000 generated tokens as the active context grew.
I also tested PQ2_0 on the same 8GB GPU. It was much tighter on memory, around 7.6 GiB / 8 GiB, but still fit with the same 32K context and quantized KV cache.
Interestingly, PQ2_0 was faster:
- prompt processing: ~97 tok/s for the test prompt
- long-generation average: 13.7 tok/s
- roughly ~20% faster decoding than PTQ1_0 in my tests
So from a VRAM capacity standpoint, 22GB should give you a very large amount of headroom for either PTQ1_0 or PQ2_0, even with a fairly large context.
The 2080 Ti is also Turing / compute capability 7.5, like the 2060 SUPER I tested, so I would expect the same PrismML CUDA path to be applicable.
I don't have a direct 2080 Ti 22GB benchmark, so I don't want to guess an exact tokens/sec figure. But given that the model already runs fully on an 8GB 2060 SUPER, your card should certainly be an interesting one to benchmark.
If you try it, I'd be very interested to see your llama-bench results, especially pp512, pp2048, and tg128, since they would make a nice comparison with the 2060 SUPER results.
@mktnhr Thank you for your warm reply. Here's my simple test results for share with you.
Environment: Ubuntu 26.04 LTS (GNU/Linux 7.0.0-31-generic x86_64) - RTX 2080Ti 22G - NVIDIA-SMI 595.91.07 - Driver Version: 595.91.07 - CUDA Version: 13.2
=========================== PTQ1_0 ================================
sudo ./llama-bench -m Qwen3.8-27B-abliterated-Ternary-Bonsai-PTQ1_0-Huihui.gguf -ngl 99 -fa on -t 4 -p 512,2048 -n 128 -r 3
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 22001 MiB):
Device 0: NVIDIA GeForce RTX 2080 Ti, compute capability 7.5, VMM: yes, VRAM: 22001 MiB
model size params backend ngl threads fa test t/s qwen35 27B PTQ1_0 - 1.75 bpw ternary (group 128) 6.14 GiB 26.90 B CUDA 99 4 1 pp512 477.62 ± 3.29 qwen35 27B PTQ1_0 - 1.75 bpw ternary (group 128) 6.14 GiB 26.90 B CUDA 99 4 1 pp2048 474.02 ± 0.55 qwen35 27B PTQ1_0 - 1.75 bpw ternary (group 128) 6.14 GiB 26.90 B CUDA 99 4 1 tg128 34.23 ± 0.04
build: 3ae4f5108 (10725)
=========================== PQ2_0 ================================
sudo ./llama-bench -m Qwen3.8-27B-abliterated-Ternary-Bonsai-PQ2_0-Huihui.gguf -ngl 99 -fa on -t 4 -p 512,2048 -n 128 -r 3
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 22001 MiB):
Device 0: NVIDIA GeForce RTX 2080 Ti, compute capability 7.5, VMM: yes, VRAM: 22001 MiB
model size params backend ngl threads fa test t/s qwen35 27B PQ2_0 - 2.13 bpw (group 128) 7.17 GiB 26.90 B CUDA 99 4 1 pp512 724.96 ± 7.07 qwen35 27B PQ2_0 - 2.13 bpw (group 128) 7.17 GiB 26.90 B CUDA 99 4 1 pp2048 721.74 ± 1.09 qwen35 27B PQ2_0 - 2.13 bpw (group 128) 7.17 GiB 26.90 B CUDA 99 4 1 tg128 38.61 ± 0.02
build: 3ae4f5108 (10725)
As you can see, the test results match your expectations. PQ2_0 faster than PTQ1_0 on compute capability 7.5.