Instructions to use here-be-dragons-ai/Kolibri-1-MLX-3bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("here-be-dragons-ai/Kolibri-1-MLX-3bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "here-be-dragons-ai/Kolibri-1-MLX-3bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "here-be-dragons-ai/Kolibri-1-MLX-3bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default here-be-dragons-ai/Kolibri-1-MLX-3bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "here-be-dragons-ai/Kolibri-1-MLX-3bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Run an OpenAI-compatible server
# Install MLX LM
uv tool install mlx-lm# Start the server
mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
# Calling the OpenAI-compatible server with curl
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "here-be-dragons-ai/Kolibri-1-MLX-3bit",
"messages": [
{"role": "user", "content": "Hello"}
]
}'Kolibri-1 MLX 3-bit
A mixed 3/6-bit MLX quantization of Aleph-Alpha/Kolibri-1, Aleph Alpha's 78B-A3.5B mixture-of-experts reasoning model for German and English.
This is a community conversion by here-be-dragons.ai, not an official Aleph Alpha release. For the model itself (training, evaluations, intended use, limitations) see the original model card and the tech report.
Quantization
| Part | Precision |
|---|---|
| Routed experts (75.5B of 78.1B parameters) | 3 bit affine, group size 64 |
| Attention, shared expert, embedding, LM head | 6 bit affine, group size 64 |
MoE router (mlp.gate) |
bf16, as in the release; expert_bias stays fp32 |
3.61 bits per weight, 33 GiB on disk. The block-FP8 release weights were dequantized to bf16 and quantized once, with no intermediate format.
Requirements
The
kolibri1architecture is not yet part of a released mlx-vlm or mlx-lm. Until the port is merged upstream, install mlx-vlm from thekolibri1branch of our fork:pip install git+https://github.com/here-be-dragons-ai/mlx-vlm@kolibri1mlx-lm support is pending upstream review.
The same weights load in both mlx-vlm and mlx-lm.
Usage (mlx-vlm)
from mlx_vlm import load, stream_generate
model, processor = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
tokenizer = processor.tokenizer
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
for chunk in stream_generate(model, processor, prompt, max_tokens=2048,
temperature=1.0, top_p=0.97, top_k=128):
print(chunk.text, end="", flush=True)
OpenAI-compatible server:
python -m mlx_vlm.server --model here-be-dragons-ai/Kolibri-1-MLX-3bit --port 8080
Usage (mlx-lm)
from mlx_lm import load, stream_generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
sampler = make_sampler(temp=1.0, top_p=0.97, top_k=128)
for chunk in stream_generate(model, tokenizer, prompt, max_tokens=2048, sampler=sampler):
print(chunk.text, end="", flush=True)
Server notes (mlx-vlm)
Reasoning is controlled with the top-level request fields reasoning_effort or
enable_thinking; the server ignores them inside chat_template_kwargs. The thinking
comes back in reasoning, separate from content.
Recommended sampling, from the original release: temperature=1.0, top_p=0.97, top_k=128
(also set in generation_config.json).
The chat template supports Kolibri's reasoning mode: pass reasoning_effort
(none, low, medium, high) to apply_chat_template. Without it, the model does not think.
Measurements
M5 Pro, 48 GB, iogpu.wired_limit_mb=40960, mlx 0.32.2, mlx-vlm 0.7.4. Method,
scripts and raw data:
local-sovereign-mlx/docs/kolibri-quality.
Quality (2026-10-09)
Measured against the FP8 release; all arms run the same layer-streamed forward pass, logits in fp32 from each arm's own LM head. Two extra arms: the noise floor is the FP8 release run against itself with the prefill split into 512-token pieces (same weights, different rounding), and uniform 3-bit is a control built for this comparison, not a release.
The noise floor is high for this model. Rounding differences of about 1e-4 after the first layer tip its top-6-of-384 expert routing and grow over 50 layers, so the FP8 release already differs from itself by mean KL 0.036 on Wikipedia. Read the numbers below against that floor.
KL divergence to FP8 per token, mean / p99.9, and the share of positions with the same most likely next token:
| text | noise floor | this build (3.61 bpw) | uniform 3-bit (3.51 bpw) |
|---|---|---|---|
| Wikipedia de/en (76k tokens) | 0.036 / 5.7 / 95.0% | 0.114 / 10.3 / 88.9% | 0.376 / 13.1 / 76.5% |
| Calibration v5 (120k) | 0.094 / 8.6 / 91.8% | 0.208 / 11.0 / 85.3% | 0.525 / 12.8 / 73.1% |
| chat, oasst2 de/en (65k) | 0.011 / 1.3 / 97.0% | 0.398 / 9.7 / 72.6% | 0.527 / 10.5 / 68.9% |
| tool calling (59k) | 0.042 / 6.8 / 97.4% | 0.112 / 10.1 / 94.0% | 0.310 / 13.5 / 88.0% |
| FLORES, 23 EU languages (91k) | 0.061 / 4.5 / 89.9% | 0.167 / 6.4 / 81.0% | 0.515 / 8.3 / 65.8% |
This build sits at 2 to 3 times the noise floor on all texts except chat,
where it moves clearly further from FP8 (perplexity 11.6 → 13.5). About half
of its heavy tail on Wikipedia (p99.9 10.3) is already in the noise floor
(5.7). The KL columns are the ones llama-perplexity --kl-divergence prints.
Multiple choice, reasoning_effort=none, scored by letter logits. Δ in
points against FP8 with a paired 95% interval; flips are answers that turn
right → wrong / wrong → right:
| benchmark | FP8 | this build | Δ (95% CI) | flips | uniform 3-bit | Δ (95% CI) | flips |
|---|---|---|---|---|---|---|---|
| Belebele de (900) | 92.9% | 93.1% | +0.2 (−0.7 to +1.2) | 7/9 | 92.0% | −0.9 (−2.2 to +0.4) | 21/13 |
| Belebele en (900) | 95.2% | 94.8% | −0.4 (−1.4 to +0.5) | 10/6 | 93.6% | −1.7 (−2.9 to −0.5) | 21/6 |
| Global-MMLU-Lite de (400) | 72.5% | 71.2% | −1.2 (−3.2 to +0.7) | 10/5 | 68.5% | −4.0 (−7.0 to −1.0) | 26/10 |
| Global-MMLU-Lite en (400) | 73.8% | 72.0% | −1.8 (−3.9 to +0.3) | 12/5 | 72.0% | −1.8 (−4.2 to +0.7) | 15/8 |
For this build no difference is detectable on any set, but equivalence within ±1 point cannot be shown at these sample sizes. The noise floor changes no answer. These are likelihood scores without reasoning: they show what the quantization changes and are not comparable with the scores on the original model card. Method: quality-method.md.
Time to first token (2026-10-08)
On Apple Silicon at long context, the prefill is what you wait for.
| context | cold | prefill rate | follow-up turn, prefix-cache hit |
|---|---|---|---|
| 1k | 0.8 s | 1,620 t/s | 0.35 s |
| 8k | 5.3 s | 1,590 t/s | 0.4 s |
| 32k | 24 s | 1,380 t/s | 0.6 s |
| 64k | 59 s | 1,090 t/s | 1.4 s* |
| 96k | 109 s | 885 t/s | – |
mlx-vlm server, PREFILL_STEP 2048, exact APC, reasoning_effort=low. About
two minutes before the first token at 96k cold. * With
APC_MEMORY_RESERVE_GB=1.5: mlx-vlm's automatic 4 GiB reserve is too large
next to 33 GB of weights on a 48 GB Mac, so follow-up turns from ~32k tokens
up missed the cache; 96k not re-measured. With reasoning_effort=none
follow-up turns miss the cache; see the linked notes.
Decode (2026-10-03)
~70 tokens/s at short context, 57 t/s at 23k, 40 t/s at 96k.
Memory
Peak 35.3 GB on short prompts, 38.0 GiB at 96k tokens.
Correctness
The port's forward pass matches a reference implementation of the vLLM
semantics to 1e-5 (on CPU). Needle retrieval succeeded at 23k and 96k tokens;
tool calls work through <tool_call> + JSON.
License
Apache 2.0, same as the original model. See LICENSE. Kolibri 1 was developed by Aleph Alpha Research GmbH.
- Downloads last month
- 1,334
3-bit
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm# Interactive chat REPL mlx_lm.chat --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"