Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
SYCL backend: any speculative type collapses performance (even target prefill drops ~200x) - draft model itself is healthy
Environment
- GPU: Intel Arc A770 16GB (DG2/ACM-G10), driver 32.0.101.8991
- Stack: oneAPI 2026.1, Level Zero backend (confirmed
[level_zero:gpu:0]in banner) - OS: Windows 11, Python-free CLI build
- Source: PrismML fork source snapshot (tarball, build banner reports
b0-unknown) - Model: Ternary-Bonsai-2-27B PTQ1_0 (qwen35 hybrid arch, 1.75 bpw)
- Draft: Qwen3.8-4B-Distill Q4_K_M (empero-ai, same vocab, qwen35 hybrid arch)
Summary
With any speculative type enabled (draft-simple, ngram-simple), performance
collapses far below the no-spec baseline β and notably the target model's own
prefill collapses too, which points at a synchronization/pacing problem in the
speculative driver rather than draft quality or GPU offload.
Measurements (same machine, same binaries)
| Config | pp512 | tg128 / generation |
|---|---|---|
Target alone (llama-bench) |
215 t/s | 15.9 t/s |
Draft alone (llama-bench, Qwen3.8-4B Q4_K_M) |
1308 t/s | 22.1 t/s |
spec draft-simple, draft-n-max 16 (llama-cli) |
5.9 t/s | 0.5β0.9 t/s |
spec ngram-simple, draft-n-max 16 (llama-cli, no draft model) |
6.4 t/s | 9.0 t/s |
Notes:
- The draft model is fully healthy in isolation:
offloaded 34/34 layers to GPU,
1308 t/s prefill. Vocab is compatible (accepted tokens decode correctly). - Under
draft-simple, verbose timings showdraft_n: 6β13, draft_n_accepted: 1β3
(~20β30% accept rate) andpredicted_per_token_msof 1285β2183 ms β vs
~63 ms/token for the target alone. The per-step cost scales with draft_n, and
the draft alone should cost ~50 ms per drafted token at its standalone speed. - Under
ngram-simplethere is no draft model at all, yet target prefill
still drops from ~1300-class to 6.4 t/s. This is the strongest hint that the
spec driver enforces a per-step (or per-op) sync on the target queue on SYCL,
e.g. waiting on every verification batch instead of pipelining.
Reproduce
# baseline (healthy)
llama-bench -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 999 -p 512 -n 128
# collapsed
llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
-md Qwen3.8-4B-Q4_K_M.gguf -ngl 999 -ngld 999 \
--spec-type draft-simple --spec-draft-n-max 16 \
-st -p "say OK" -n 32
# same collapse with no draft model
llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 999 \
--spec-type ngram-simple --spec-draft-n-max 16 -st -p "say OK" -n 32
Secondary issue: default type inference crashes on plain drafts
Without an explicit --spec-type, a plain (non-sidecar) draft GGUF is routed
into the draft-mtp implementation and aborts:
speculative.cpp:2082: GGML_ASSERT(n_embd == llama_model_n_embd_out(...) &&
"MTP input row width must match the target h_nextn width") failed
For a draft model with no MTP metadata, falling back to draft-simple (or
failing with a clear "no sidecar found, pass --spec-type" message) would be
friendlier than an assert.
Question
Is SYCL a supported backend for the rewritten speculative driver? If known-broken,
a docs note would save others the debugging session; if not known, happy to
provide verbose logs / tracing output on request.