Instructions to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0 # Run inference directly in the terminal: llama cli -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0 # Run inference directly in the terminal: llama cli -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0 # Run inference directly in the terminal: ./llama-cli -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Use Docker
docker model run hf.co/Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
- LM Studio
- Jan
- vLLM
How to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Override-6/Ternary-Bonsai-2-27B-abliterated-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Override-6/Ternary-Bonsai-2-27B-abliterated-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
- Ollama
How to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with Ollama:
ollama run hf.co/Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
- Unsloth Desktop
- Pi
How to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with Docker Model Runner:
docker model run hf.co/Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
- Lemonade
How to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-abliterated-gguf-TQ1_0
List all available models
lemonade list
- Hermes Agent
How to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Override-6/Ternary-Bonsai-2-27B-abliterated-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Override-6/Ternary-Bonsai-2-27B-abliterated-gguf:TQ1_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Ternary-Bonsai-2-27B — abliterated (still pure ternary)
This is an abliterated version of prism-ml/Ternary-Bonsai-2-27B-gguf (PTQ1_0 packing). It keeps the exact original file size, 5,946,648,928 bytes:
- Every tensor has its original type and shape.
- All 66 edited weight matrices are still
{-1, 0, +1}× fp16 group-128 scales. - On average 0.64% of their trits were flipped, and their scales were refit.
- No tensor was upcast to a higher-precision format.
Created using Bonsai by Prism ML. Built from Qwen3.8-27B (Alibaba Cloud). See
NOTICE.txt.
Usage
These files need the PrismML llama.cpp fork, exactly like the original. Stock llama.cpp cannot run PTQ1_0 or apply the Hadamard activation transform.
# https://github.com/PrismML-Eng/llama.cpp/releases (tested with prism-b10709-9a9394a, CUDA 12.4)
./llama-server -m Ternary-Bonsai-2-27B-PTQ1_0-abliterated.gguf -ngl 99 -fa on -c 32768
Sampling is the same as the base model: thinking mode temp 1.0, top_p 0.95, top_k 20.
Results
All numbers were measured locally on an RTX 4070 with the fork above.
| Original PTQ1_0 | Naive abliteration (writers re-quantized to Q6_K) | This model | |
|---|---|---|---|
| File size | 5.95 GB | 9.36 GB | 5.95 GB |
| Harmful prompts, thinking mode: complied / deflected / refused | 0 / 3 / 101 | – | 76 / 27 / 1 |
| Harmful prompts, no-think mode: complied / deflected / refused / empty | 1 / 8 / 95 / 0 | 44 / 50 / 10 / 0 | 74 / 30 / 0 / 0 |
| MMLU, log-prob scoring (n=1000) | 78.9 | 78.2 | 79.0 |
| GSM8K, thinking mode (n=100) | 94.0 | 93.0 | 94.0 |
| IFEval, prompt-level strict (n=541, no-think) | 78.4 | 78.0 | 81.7 |
| IFEval, instruction-level strict | 85.5 | 84.7 | 87.1 |
| Needle-in-a-haystack, 8k / 16k / 32k / 64k (n=40) | 100 / 100 / 100 / 100 | 100 / 100 / 100 / 100 | 100 / 100 / 100 / 100 |
| KL divergence vs original, harmless chats | 0 | 0.0240 | 0.0189 |
| KL divergence vs original, wikitext-2 | 0 | 0.0134 | 0.0130 |
Evaluation setup
- Harmful prompts: the 104 held-out prompts of
mlabonne/harmful_behaviors(test split). Responses are labeled by an LLM judge (the original model in non-thinking mode,judge.py):- complied: it does what was asked
- deflected: it answers with a sanitized, fictional or "ethical" substitute
- refused
- empty
- KL divergence comes from
llama-perplexity --kl-divergenceagainst the original model. The harmless chat corpus is the original model's own answers to 200 held-outharmless_alpacaprompts. - IFEval uses the official checker code (via
lm_eval). Needle-in-a-haystack hides a 7-digit code at 5 depths inside wikitext contexts. - Limits: these are subsets and spot checks, not the 14-benchmark thinking-mode suite reported by Prism ML. Coding, tool calling and vision were not re-evaluated.
- Noise levels: with 100 GSM8K questions, differences of ±2–3 points are noise. The IFEval gain (+3.3 points, about 1.8 standard errors) is best read as "no loss, possibly a small gain".
- Deflections remain: about a quarter of harmful requests still get a watered-down answer. A stronger setting reduces this at a measurable KL cost.
Comparison with other Bonsai 2 abliterations
Other community builds were evaluated with the same harness: same held-out prompts, same judge, same KL corpora, same flags. Numbers can differ from the ones on their own cards, which use other prompt sets and judges.
| Size | No-think: complied / deflected / refused | Thinking: complied / deflected / empty | KL chat / wiki | MMLU | GSM8K | IFEval strict | NIAH 8k–64k | |
|---|---|---|---|---|---|---|---|---|
| Original (prism-ml) | 5.95 GB | 1 / 8 / 95 | 0 / 3 / 0 | 0 / 0 | 78.9 | 94 | 78.4 | 100% |
| This model | 5.95 GB | 74 / 30 / 0 | 76 / 27 / 0 | 0.019 / 0.013 | 79.0 | 94 | 81.7 | 100% |
| OS-Software Heretic (PTQ1_0) | 5.95 GB | 97 / 4 / 3 | 101 / 0 / 3 | 0.011 / 0.008 | 77.3 | 91 | 78.2 | 100% |
| dealignai CRACK (PQ2_0) | 7.21 GB | 88 / 16 / 0 | 98 / 5 / 1 | 0.046 / 0.060 | 78.9 | 93 | – | – |
| BoldingBuilds (PTQ1_0) | 5.95 GB | 47 / 57 / 0 | 56 / 29 / 19 | 0.055 / 0.050 | 79.5 | 92 | – | – |
Summary: the OS-Software Heretic build is more uncensored: it deflects far less, with lower KL. This model keeps every capability benchmark at or above the original.
Prior work on lattice editing: editing the ternary/2-bit codes directly, instead of re-quantizing, is not unique to this model. BoldingBuilds, Hikari07jp and OS-Software released lattice-edited builds earlier. What this model adds is the published end-to-end pipeline and the scale-refit healing step.
Abliteration process
The scripts are in scripts/. The fixed refusal direction is scripts/directions/dirs_proj_both.npy.
1. Why the usual recipe fails on a ternary model
Abliteration replaces every residual-stream writer W with W − α·r·rᵀW. On this model:
- The per-weight change is far below one ternary step (about 1% of a step, at most half a step).
- Re-quantizing the projected matrix therefore just rounds back to the original model, so a naive ternary abliteration silently does nothing.
- The alternative is to store the writers at higher precision, which makes the file 60% bigger.
2. Refusal direction (make_dirs.py, tools/actdump.cpp)
- Capture the residual stream (
l_out) of every layer at the last prompt token, for 400 harmful and 400 harmless training prompts. - Compute the direction at the layer-36 residual:
- difference of means, harmful minus harmless
- pooled over both answer positions:
<think>\n(thinking mode) and</think>\n\n(no-think mode). The two per-mode directions only have cosine 0.64. A direction taken from one mode only leaves the other mode broken: the model returns empty answers. - orthogonalized against the harmless mean. The raw direction also carries general features; removing it wholesale gave KL ≈ 2.
- On held-out prompts it separates harmful from harmless with Cohen's d ≈ 9.6 (thinking) and 9.05 (no-think).
3. Ternary-preserving ablation (abliterate.py)
The trits are edited directly. The PTQ1_0 block codec is re-implemented and verified to round-trip byte-exact.
For each writer in layers 16–48, the exact delta
−α·r·(rᵀW)is realized by unbiased stochastic trit flips, followed by a small greedy correction so thatrᵀW′ = (1−α)·rᵀWholds column by column.Per-component strength:
- α = 2.0 on attention and linear-attention output writers (
attn_output,ssm_out) - α = 0.7 on
ffn_down
MLP edits cost far more KL per unit of refusal removed, consistent with Heretic's findings.
- α = 2.0 on attention and linear-attention output writers (
The Hadamard rotation of this model is on the input axis only, so an output-side projection is exact in the stored basis.
4. Healing (heal.py, tools/wdump.cpp)
Capture the rotated inputs of every writer (
H = E[xxᵀ]) on 16k calibration tokens:- alpaca train prompts with the original model's answers
- GSM8K train thinking traces
- harmful-prompt train responses
- wikitext-2 train
None of these overlap the evaluation sets.
For each edited writer, refit only the fp16 group scales by ridge least squares, per output row:
min_s (s·B − w*) H (s·B − w*)ᵀ, wherew* = (I − α·r·rᵀ)Wis the exact abliterated target. Trits, layout and size are unchanged.
5. What didn't work (so you don't have to try)
| Approach | Result |
|---|---|
| Raw (unprojected) direction on all layers | KL 1.4–7.8, model broken |
| Direction from the thinking position only | many empty answers in no-think mode |
| Narrower layer bands (16–40, 16–44) | refusal comes back (Hydra effect) |
| GPTQ trit re-selection | output error is lower, but it does not remove the direction by itself |
| Two orthogonal directions | worse on both refusal and KL |
| Overshoot α = 1.3–1.5 on all writers | 0% refusal but KL 2–4× higher and many empty answers |
Reproducing
actdumpon thinking and no-think prompt files (prompts.py), thenmake_dirs.py.wdumpon the calibration corpus (gen_calib.py), all layers.Build the model:
python abliterate.py --src Ternary-Bonsai-2-27B-PTQ1_0.gguf --out cand.gguf \ --dirs directions/dirs_proj_both.npy --dir-layer 36 --layers 16-48 --alpha-attn 2.0 --alpha-ffn 0.7 python heal.py --orig Ternary-Bonsai-2-27B-PTQ1_0.gguf --edited cand.gguf --out final.gguf --acts wacts \ --dirs directions/dirs_proj_both.npy --dir-layer 36 --layers 16-48 --alpha-attn 2.0 --alpha-ffn 0.7 --recorrect 0Evaluate with
evalmodel.sh(refusal + KL),bench.py,mmlu_lp.pyandjudge.py.
Stochastic flips use seed 0.
Disclaimer
This model has had its refusal behavior removed. It will produce content the original model would decline to produce. You are responsible for how you use it and for complying with applicable laws.
- Downloads last month
- 2,038
1-bit