Instructions to use ukisai/Swift-Bonsai-2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ukisai/Swift-Bonsai-2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0 # Run inference directly in the terminal: ./llama-cli -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0
Use Docker
docker model run hf.co/ukisai/Swift-Bonsai-2-GGUF:Q2_0
- LM Studio
- Jan
- vLLM
How to use ukisai/Swift-Bonsai-2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ukisai/Swift-Bonsai-2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-Bonsai-2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ukisai/Swift-Bonsai-2-GGUF:Q2_0
- Ollama
How to use ukisai/Swift-Bonsai-2-GGUF with Ollama:
ollama run hf.co/ukisai/Swift-Bonsai-2-GGUF:Q2_0
- Unsloth Desktop
- Pi
How to use ukisai/Swift-Bonsai-2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-Bonsai-2-GGUF:Q2_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ukisai/Swift-Bonsai-2-GGUF with Docker Model Runner:
docker model run hf.co/ukisai/Swift-Bonsai-2-GGUF:Q2_0
- Lemonade
How to use ukisai/Swift-Bonsai-2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ukisai/Swift-Bonsai-2-GGUF:Q2_0
Run and chat with the model
lemonade run user.Swift-Bonsai-2-GGUF-Q2_0
List all available models
lemonade list
- Hermes Agent
How to use ukisai/Swift-Bonsai-2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-Bonsai-2-GGUF:Q2_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-Bonsai-2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-Bonsai-2-GGUF:Q2_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default ukisai/Swift-Bonsai-2-GGUF:Q2_0Run Hermes
hermesSwift Bonsai 2 GGUF
Swift Bonsai 2 is UkisAI's reasoning-efficient derivative of Prism ML's Ternary Bonsai 2 27B. It uses 39.8% fewer thinking tokens while scoring 0.19% higher than the base.
Both the 1-bit and 2-bit quantizations are available in this repository; each file is a plain Bonsai 2 pack.
Quantizations
| Quantization | File | Download size |
|---|---|---|
| 1-bit / PTQ1_0 | Swift-Bonsai-2-PTQ1_0.gguf | 5.947 GB |
| 2-bit / PQ2_0 | Swift-Bonsai-2-PQ2_0.gguf | 7.206 GB |
Each file is a complete model with the Swift correction merged into the ternary weights: no adapter file, patch, or extra flag is needed, and the files are the same size as the base Bonsai 2 packs. PTQ1_0 remains the earlier Swift release; PQ2_0 now contains updated merged weights.
Training approach
We built Swift by identifying reasoning-marker tokens that, in our analysis, trigger overthinking in the model's reasoning rollouts. We then fine-tuned the model by penalizing usage of those tokens while it reasons.
Swift produces shorter reasoning traces while keeping accuracy in line with the base model.
Evaluation
GPQA and C-Eval compare base Ternary Bonsai 2 27B with the historical Swift runtime correction under the saved protocols. IFBench and AIME compare base PQ2_0 with the current Swift PQ2_0 file. See the note below. Scores are the percentage of correct responses across all scored repetitions: three complete GPQA-Diamond runs and five runs each of C-Eval, IFBench, and AIME 2025. Token reductions are relative to the base model.
| Benchmark | Score | Mean tokens | Median tokens | |||
|---|---|---|---|---|---|---|
| Base | Swift | Base | Swift | Reduction | Reduction | |
| General reasoning | ||||||
| GPQA-Diamond | 84.18% | 84.34% | 20,066 | 16,245 | ↓ 19.0% | ↓ 39.8% |
| C-Eval | 81.37% | 81.47% | 2,123 | 1,732 | ↓ 18.4% | ↓ 7.1% |
| IFBench | 82.13% | 82.60% | 8,670 | 8,519 | ↓ 1.7% | ↑ 0.4% |
| Mathematics | ||||||
| AIME 2025 | 92.00% | 93.33% | 22,158 | 23,101 | ↑ 4.3% | ↑ 5.1% |
Token statistics measure thinking tokens, except AIME, where they measure the full completion. IFBench uses the official loose scorer on the full 1,500 responses. An up arrow means the current 2-bit model used more tokens. Token changes are not a direct measurement of latency or cost changes.
Benchmark methodology and reproduction settings
Sampling: temperature 1.0, top-p 0.95, top-k 20, min-p 0, repetition penalty 1, presence penalty 0, and no additional inference-time logit penalty.
| Benchmark | Questions | Repetitions | Scored responses | Output cap |
|---|---|---|---|---|
| GPQA-Diamond | 198 | 3 | 594 | 81,920 |
| C-Eval validation, 5-shot | 1,346 | 5 | 6,730 | 16,384 |
| IFBench | 300 | 5 | 1,500 | 81,920 |
| AIME 2025 | 30 | 5 | 150 | 81,920 |
- C-Eval covers the complete validation split, not the hidden test set. One duplicated prompt is cached per seed; all question IDs are scored and weighted separately.
- AIME scoring: the current base and Swift PQ2_0 files used identical corrected prompts (30 questions × 5 seeds, output cap 81,920) and the same archived Math-Verify 0.9.0 scorer, which checks the final response and falls back to reasoning. That scorer gives 92.00% base / 93.33% Swift. Final-answer-only scoring gives 90.67% / 90.67%. The AIME token columns report full completion tokens.
- IFBench reports official loose scoring for the current base and Swift PQ2_0 files (300 questions × 5 seeds). The current run completed all 1,500 responses without request errors; prompt and sampler settings were spot-checked against the base run.
- Model selection: GPQA informed the choice of training settings and is not an untouched holdout. Named benchmark exclusions and exact-match checks were performed; a comprehensive near-overlap audit was not completed.
- GPQA and C-Eval are the original measured base/Swift runtime results. IFBench and AIME were refreshed on the current PQ2_0 file against base PQ2_0. PTQ1_0 remains the earlier merged release. Reproducibility can vary with runtime builds and hardware.
- Small score differences are not established improvements. Exact aggregates, truncation counts, and confidence intervals where computed are available in
benchmark_results.json.
Research status and limitations
This is an experimental research release. In our internal evaluations and practical testing, benchmark scores did not consistently translate into reliable general-purpose behavior. Instruction following, tool use, and open-ended coding or agent tasks remain uneven. These observations concern the specific models, runtimes, and tests we used; they do not establish a general conclusion about ternary models.
We see ternary quantization as a promising direction for making larger models more accessible. Further advances in training, low-bit adaptation, and inference support may make the approach increasingly useful. This release shares a research result and its current limitations, rather than presenting a production-ready model.
How to use
Download and run
These files use Prism's PQ2_0 and PTQ1_0 ternary packs and run on the PrismML-Eng/llama.cpp fork (tested at revision 1a07bfa5f). Stock llama.cpp does not support these tensor types. No patch, adapter file, or extra flag is needed.
Linux prerequisites: Git, CMake, a C++ compiler, and the NVIDIA CUDA toolkit.
git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release -j
hf download ukisai/Swift-Bonsai-2-GGUF Swift-Bonsai-2-PQ2_0.gguf --local-dir .
./build/bin/llama-server -m Swift-Bonsai-2-PQ2_0.gguf --host 127.0.0.1 --port 8080 \
-ngl 99 -c 32768 --flash-attn on --jinja \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --repeat-penalty 1.0
To run 1-bit / PTQ1_0 instead, download Swift-Bonsai-2-PTQ1_0.gguf and pass it to -m.
The command above serves a 32,768-token context on 127.0.0.1:8080. Use -c 98304 for the longer evaluation caps if memory permits. Runtime memory includes model weights and context/state caches; download size is not total VRAM usage.
OpenAI-compatible API
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "bonsai-2-swift",
"messages": [{"role": "user", "content": "Explain the difference between correlation and causation."}],
"temperature": 1.0,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"max_tokens": 4096,
"chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"}
}'
The launcher flags above are the intended model settings. Custom integrations should preserve them.
License and attribution
Created using Bonsai by Prism ML. This is an independent UkisAI release, not an official Prism ML release or endorsement.
The model is distributed under Apache-2.0. The upstream notice and attributions are retained.
Citation
@misc{swift-bonsai-2-gguf,
title = {Swift Bonsai 2 GGUF},
author = {UkisAI},
year = {2026},
url = {https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF}
}
- Downloads last month
- 3,808
1-bit
2-bit

Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf ukisai/Swift-Bonsai-2-GGUF:Q2_0