How to use from the
Use from the
MLX library
# Make sure mlx-vlm is installed
# pip install --upgrade mlx-vlm

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

# Load the model
model, processor = load("pipenetwork/Ternary-Bonsai-2-27B-MLX-8bit")
config = load_config("pipenetwork/Ternary-Bonsai-2-27B-MLX-8bit")

# Prepare input
image = ["http://images.cocodataset.org/val2017/000000039769.jpg"]
prompt = "Describe this image."

# Apply chat template
formatted_prompt = apply_chat_template(
    processor, config, prompt, num_images=1
)

# Generate output
output = generate(model, processor, formatted_prompt, image)
print(output)

Ternary-Bonsai-2-27B-MLX-8bit

Stock-runtime MLX build of Ternary-Bonsai-2-27B — prism-ml's ternarized Qwen3.8-27B (64-layer qwen3_5 hybrid: Gated-DeltaNet + every-4th full attention, with the official vision tower) — with the blockwise Hadamard rotation unfolded back into the standard weight basis, so it loads in unmodified mlx-vlm (≥ 0.7): no custom runtime, no forked kernels.

8-bit affine (group 64): statistically indistinguishable from bf16 at 54% of the size.

from mlx_vlm import load, generate
model, processor = load("pipenetwork/Ternary-Bonsai-2-27B-MLX-8bit")

How this relates to prism-ml's own MLX release

prism-ml ships an excellent 8.6 GB 2-bit pack whose weights are the exact ternary values — it is the efficiency frontier for this model, and it requires their bundled runtime (the stored weights are Hadamard-rotated; activations are transformed to match). This set serves the complementary case: standard-basis weights for stock tooling, fine-tuning, and downstream conversion. The unfold is exact — refold(unfold(W)) is bitwise identical in fp32, and the fold contract was verified against their runtime, not assumed (applying the sign vector in the wrong order moves logits by 7.4; the test catches it).

Fidelity

Against prism's 2-bit pack running under their runtime, same prompts, 81 positions: max |Δlogit| 0.22 on a ±20 scale, cosine 0.99999, argmax 96.7–100% (the flips are ties; in fp16 — their activation dtype — agreement is 100%). Paired over 145 shared wikitext-2 windows the perplexity ratio is 0.9992 [0.9990, 0.9994]: a bf16-rounding-level difference.

Perplexity (wikitext-2 test, 296,815 tokens, identical windows through stock mlx-vlm):

build size ppl
prism 2-bit, their runtime 8.6 GB 8.9607
bf16 (this set's unquantized) 54.7 GB 8.9679
8-bit 29.5 GB 8.9636
6-bit 22.8 GB 8.9548
4-bit 16.1 GB 9.1497

8-bit and 6-bit are statistically indistinguishable from bf16; 4-bit costs +2.1% — the only build with a measurable loss, and still the smallest stock-loadable one. Vision verified end-to-end. Requires mlx-vlm ≥ 0.7 (earlier versions double-shift qwen3_5 norms).

License

Apache-2.0, as upstream (prism-ml/Ternary-Bonsai-2-27B-gguf); their NOTICE.txt is included. Conversion code: https://github.com/PipeNetwork/bonsai2-mlx.

Downloads last month
688
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pipenetwork/Ternary-Bonsai-2-27B-MLX-8bit

Base model

Qwen/Qwen3.8-27B
Quantized
(31)
this model

Collection including pipenetwork/Ternary-Bonsai-2-27B-MLX-8bit