Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Sitting frozen like a nuclear bomb, the 8G laptop speeds up from 50 tokens straight to 70 tokens,8G再次从50Token/S大幅升到恐怖的70Token/s
Model: Ternary-Bonsai-2-27B, ternary 1.75 bop (5.79 GB)
Speedup Effect Source
Ternary quantization + matching CUDA kernels 27B down to 5.79 GB, runs on one card professorpalmer/llama.cpp-ada-ternary
CUDA Graph 1.5× llama.cpp / ggml
MTP speculative decoding, depth 3 3 extra tokens per forward pass llama.cpp --spec-type draft-mtp; depth tuned by us
Draft vocab 248k → 24k Draft step reads 25.6 MB, 90% less Our own work
27B 模型,8GB 笔记本显卡,关思考 70 t/s。
模型:Ternary-Bonsai-2-27B,三值量化 1.75 bpw(5.79 GB)
提速技术 效果 来源
三值量化 + 配套 CUDA 内核 27B 压到 5.79 GB,单卡可跑 professorpalmer/llama.cpp-ada-ternary
CUDA Graph 提速 1.5× llama.cpp / ggml
MTP 推测解码,深度 3 一次前向多出 3 个 token llama.cpp --spec-type draft-mtp,深度由我们实测定
起草词表 248k → 24k 起草一步读 25.6 MB,省 90% 我们自己做的
稍后补上实测的速度及输出效果
具体怎么弄?能开多大上下文? 谢谢
Ternary-Bonsai-2-27B @ RTX 4060 Laptop 8GB:63~70 t/s 完整复现方案
采样 --temp 0.6 --top-k 20 --top-p 0.95。
一、下载什么
| 组件 | 内容 |
|---|---|
| 主模型 | Ternary-Bonsai-2-27B-PTQ1_0-mtp.gguf(~6.2 GB),来自 prism-ml 的 Bonsai-2 HuggingFace 仓库(基座 = Qwen/Qwen3.8-27B 的三值量化版)。关键:必须用带 -mtp 后缀的这份——MTP 草稿头内嵌在文件里(blk.64.nextn),不需要单独的 draft 模型 |
| 推理程序 | professorpalmer 的 bonsai-q8-product llama.cpp fork(我们基于 build 462,commit c8b8993a6,sm_89)。它带 PTQ1_0 三值量化 CUDA 内核和 draft-mtp 推测解码支持;原版 llama.cpp 不认这种量化,会直接加载失败 |
| 视觉(可选) | mmproj-Qwen3.8-27B-Q8_0.gguf,不挂就是纯文本 |
二、fork 上必须有的改动(serving recipe 在源码里,不在命令行)
拿到 fork 后核对/补上这几处源码级配置(都是我们实测出来的最优,每条都有 A/B 数据):
- KV 缓存 q8_0(目标 context 和 draft context 都是):速度与 q4_0 打平,但接受率 0.789→0.823——赢在接受率上
- ubatch 256(
-ub 256):512/1024 都是负收益(−5~10%) --spec-draft-window 16384:起草侧上下文窗口,配合 KV q8_0 拉接受率--spec-draft-n-max-tail 4- PTQ1_0 的 PT-vs-MMQ 内核切换点 4→5(
ggml-cuda/mmvq.cu):depth≥4 反正是悬崖,这条让小 batch 走更快的 PT 路径 - 草稿词表 shortlist 支持:
--spec-draft-vocab-map+ 加载时构建 compact 头副本(见下节) - 采样默认:
temp 0.6 / top-k 20 / min-p 0(写在common.h)
三、自建词表头(shortlist)——最大的单点优化
原理:MTP 草稿头的 output.weight 是 [5120, 248320] 全词表(265 MB),每个起草步整读一遍。shortlist 把起草打分限制在自己语料实际用到的 24,000 行上,每步只读 ~26 MB。目标模型验证时仍用全词表,所以输出分布一个 bit 都不变——只省字节,不伤质量(实测接受率持平)。
怎么建(fork 内 vocab/mixbuild.py):
- 备三个语料池,全部用线上同一个 tokenizer 切分:
code:本地仓库树挖掘的代码,~1790 万 token(主力)talk:让模型自己回答 46 条中文 + 8 条中文代码题的输出,~4.1 万 tokenreal:真实服务流量,~6.2 万 token
- 池权重
real=3, talk=2, code=2,按出现频次选出 top-24,000 个 token,保留 25% holdout 自测 - 表长 24k 是实测最优:4k/8k 漏太多,64k 不回本;CJK 大约 3,600 行就饱和
- 产出
shortlist_24k_mix.vocab.txt,加载时 fork 会自动把它拷成头自己的量化格式
四、启动脚本关键配置项
set "SLOTS=1" rem 必须 1!
set "CTX=65536" rem 见下节上下文
set "REA=on" rem 推理模式开关
set "EFFORT=medium" rem medium 是唯一不向 system prompt 注入文字的档位
set "DRAFT_MAP=maps\shortlist_24k_mix.vocab.txt"
set "TIER=0"
if %CTX% GTR 32768 set "TIER=24000" rem 分级 KV 的 VRAM cell 数
bin\llama-server.exe -m models\Ternary-Bonsai-2-27B-PTQ1_0-mtp.gguf ^
--host 127.0.0.1 --port 18080 --parallel %SLOTS% -rea %REA% ^
-c %CTX% -b 4096 -t 8 -ngl 99 -fa on %TIERARG% ^
--spec-type draft-mtp --spec-draft-vocab-map "%DRAFT_MAP%"
要点:
--spec-type draft-mtp,起草深度 3(代码默认)- 客户端
cache_prompt: true,采样 temp 0.6 / top-k 20
五、编译
Windows + MSVC 14.44 + CUDA(sm_89),标准 cmake 流程:
cmake -S . -B build -DLLAMA_BUILD_UI=OFF -DLLAMA_USE_PREBUILT_UI=OFF # OFF 掉内嵌 UI,免 npm/网络
ninja -C build -j 6 llama-server
产物 bin/ 需要:llama-server.exe + llama-common.dll + llama-server-impl.dll + ggml*.dll(ggml-cuda 241MB 那个是本体)+ CUDA 运行时 dll 放旁边。
六、上下文能开多大
是的,128K 能开。 机制是分级 KV:q8_0 每位置 34.8 KB,8GB 卡 VRAM 只装得下 ~2.8 万位置,超出部分进 pinned host 内存的“尾段”(--kv-vram-cells 24000 控制 VRAM 头槽位数),一个 VMM 地址段内自动调度,业务无感:
建议:8GB 卡日常 32K~64K,128K 当能力边界用,别当性能口径宣传。
两点说明:-mtp 模型和 bonsai-q8-product fork 都来自 prism-ml/professorpalmer 的发布,社区复现者拿这两个上游 + 本文第三、四节即可对齐;数字口径是我们卡上的实测(4060 Laptop 8GB/256GB/s),别的卡按带宽比例估算即可。
过程稍微麻烦但一般的AI编程大概可按这个步骤复现,代码和数据比较乱,稍后有时间会整理下放Git给大家参考。

