How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
# Run inference directly in the terminal:
llama cli -hf darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
# Run inference directly in the terminal:
llama cli -hf darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
# Run inference directly in the terminal:
./llama-cli -hf darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
# Run inference directly in the terminal:
./build/bin/llama-cli -hf darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
Use Docker
docker model run hf.co/darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF:Q2_0
Quick Links

AJAN-SIMIT Ternary-Bonsai Q2_0 (GGUF)

Q2_0 GGUF quantizations of Prism ML's Ternary-Bonsai model family, produced by DarkStar Initiative for the Ajan Simit offline mobile AI project (an on-device Android voice assistant).

We did not train or create the base model. This repository contains only our own re-quantization of Prism ML's publicly released weights, done so the model is small enough to run entirely on-device on a phone.

Provenance / credits

  • Base model: Qwen3-1.7B (Apache 2.0, Qwen team)
  • Ternary fine-tune: prism-ml/Ternary-Bonsai (Apache 2.0, Prism ML)
  • This Q2_0 re-quantization: DarkStar Initiative / Ajan Simit, produced with a self-built llama-quantize from Prism ML's own fork, PrismML-Eng/llama.cpp (prism branch) -- credited per their own model card's request.
  • License: Apache 2.0, inherited unchanged from the base model and the fine-tune. No added restrictions. A full copy of the license text is included in this repo as LICENSE.

If you use the original Ternary-Bonsai model itself (not just this re-quantization), please cite Prism ML directly:

@techreport{ternarybonsai,
  title  = {Ternary Bonsai: 1.58-bit Language Models},
  author = {Prism ML},
}

What's different from the upstream Prism ML release

Only the quantization level and the file/repo naming:

  • Quantization: Q2_0 (2.125 bpw), quantized from Prism ML's F16 release. No retraining, no architecture changes, no fine-tuning of our own.
  • Naming: the AJAN-SIMIT- prefix marks these files as our own community re-quantization and build, produced for the Ajan Simit app specifically -- not an official Prism ML release, and not endorsed by Prism ML or the Qwen team.

Files

File Size
AJAN-SIMIT-Ternary-Bonsai-1.7B-Q2_0.gguf ~463 MB
AJAN-SIMIT-Ternary-Bonsai-4B-Q2_0.gguf ~1.07 GB
AJAN-SIMIT-Ternary-Bonsai-8B-Q2_0.gguf ~2.18 GB
AJAN-SIMIT-Ternary-Bonsai-27B-Q2_0.gguf ~8.25 GB

Usage

Compatible with any recent llama.cpp-based inference engine. Q2_0 is a first-class quant type in Prism ML's llama.cpp fork; mainline llama.cpp support may vary by version.

Downloads last month
354
GGUF
Model size
2B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for darkstarinitiative/AJAN-SIMIT-Ternary-Bonsai-Q2_0-GGUF