How to use from
MLX LM
Generate or start a chat session
# Install MLX LM
uv tool install mlx-lm
# Interactive chat REPL
mlx_lm.chat --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Run an OpenAI-compatible server
# Install MLX LM
uv tool install mlx-lm
# Start the server
mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
# Calling the OpenAI-compatible server with curl
curl -X POST "http://localhost:8000/v1/chat/completions" \
   -H "Content-Type: application/json" \
   --data '{
     "model": "here-be-dragons-ai/Kolibri-1-MLX-3bit",
     "messages": [
       {"role": "user", "content": "Hello"}
     ]
   }'
Quick Links

Kolibri-1 MLX 3-bit

A mixed 3/6-bit MLX quantization of Aleph-Alpha/Kolibri-1, Aleph Alpha's 78B-A3.5B mixture-of-experts reasoning model for German and English.

This is a community conversion by here-be-dragons.ai, not an official Aleph Alpha release. For the model itself (training, evaluations, intended use, limitations) see the original model card and the tech report.

Quantization

Part Precision
Routed experts (75.5B of 78.1B parameters) 3 bit affine, group size 64
Attention, shared expert, embedding, LM head 6 bit affine, group size 64
MoE router (mlp.gate) bf16, as in the release; expert_bias stays fp32

3.61 bits per weight, 33 GiB on disk. The block-FP8 release weights were dequantized to bf16 and quantized once, with no intermediate format.

Requirements

The kolibri1 architecture is not yet part of a released mlx-vlm or mlx-lm. Until the port is merged upstream, install mlx-vlm from the kolibri1 branch of our fork:

pip install git+https://github.com/here-be-dragons-ai/mlx-vlm@kolibri1

mlx-lm support is pending upstream review.

The same weights load in both mlx-vlm and mlx-lm.

Usage (mlx-vlm)

from mlx_vlm import load, stream_generate

model, processor = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
tokenizer = processor.tokenizer
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
for chunk in stream_generate(model, processor, prompt, max_tokens=2048,
                             temperature=1.0, top_p=0.97, top_k=128):
    print(chunk.text, end="", flush=True)

OpenAI-compatible server:

python -m mlx_vlm.server --model here-be-dragons-ai/Kolibri-1-MLX-3bit --port 8080

Usage (mlx-lm)

from mlx_lm import load, stream_generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
sampler = make_sampler(temp=1.0, top_p=0.97, top_k=128)
for chunk in stream_generate(model, tokenizer, prompt, max_tokens=2048, sampler=sampler):
    print(chunk.text, end="", flush=True)

Server notes (mlx-vlm)

Reasoning is controlled with the top-level request fields reasoning_effort or enable_thinking; the server ignores them inside chat_template_kwargs. The thinking comes back in reasoning, separate from content.

Recommended sampling, from the original release: temperature=1.0, top_p=0.97, top_k=128 (also set in generation_config.json).

The chat template supports Kolibri's reasoning mode: pass reasoning_effort (none, low, medium, high) to apply_chat_template. Without it, the model does not think.

Measurements

M5 Pro, 48 GB, iogpu.wired_limit_mb=40960, mlx 0.32.2, mlx-vlm 0.7.4. Method, scripts and raw data: local-sovereign-mlx/docs/kolibri-quality.

Quality (2026-10-09)

Measured against the FP8 release; all arms run the same layer-streamed forward pass, logits in fp32 from each arm's own LM head. Two extra arms: the noise floor is the FP8 release run against itself with the prefill split into 512-token pieces (same weights, different rounding), and uniform 3-bit is a control built for this comparison, not a release.

The noise floor is high for this model. Rounding differences of about 1e-4 after the first layer tip its top-6-of-384 expert routing and grow over 50 layers, so the FP8 release already differs from itself by mean KL 0.036 on Wikipedia. Read the numbers below against that floor.

KL divergence to FP8 per token, mean / p99.9, and the share of positions with the same most likely next token:

text noise floor this build (3.61 bpw) uniform 3-bit (3.51 bpw)
Wikipedia de/en (76k tokens) 0.036 / 5.7 / 95.0% 0.114 / 10.3 / 88.9% 0.376 / 13.1 / 76.5%
Calibration v5 (120k) 0.094 / 8.6 / 91.8% 0.208 / 11.0 / 85.3% 0.525 / 12.8 / 73.1%
chat, oasst2 de/en (65k) 0.011 / 1.3 / 97.0% 0.398 / 9.7 / 72.6% 0.527 / 10.5 / 68.9%
tool calling (59k) 0.042 / 6.8 / 97.4% 0.112 / 10.1 / 94.0% 0.310 / 13.5 / 88.0%
FLORES, 23 EU languages (91k) 0.061 / 4.5 / 89.9% 0.167 / 6.4 / 81.0% 0.515 / 8.3 / 65.8%

This build sits at 2 to 3 times the noise floor on all texts except chat, where it moves clearly further from FP8 (perplexity 11.6 → 13.5). About half of its heavy tail on Wikipedia (p99.9 10.3) is already in the noise floor (5.7). The KL columns are the ones llama-perplexity --kl-divergence prints.

Multiple choice, reasoning_effort=none, scored by letter logits. Δ in points against FP8 with a paired 95% interval; flips are answers that turn right → wrong / wrong → right:

benchmark FP8 this build Δ (95% CI) flips uniform 3-bit Δ (95% CI) flips
Belebele de (900) 92.9% 93.1% +0.2 (−0.7 to +1.2) 7/9 92.0% −0.9 (−2.2 to +0.4) 21/13
Belebele en (900) 95.2% 94.8% −0.4 (−1.4 to +0.5) 10/6 93.6% −1.7 (−2.9 to −0.5) 21/6
Global-MMLU-Lite de (400) 72.5% 71.2% −1.2 (−3.2 to +0.7) 10/5 68.5% −4.0 (−7.0 to −1.0) 26/10
Global-MMLU-Lite en (400) 73.8% 72.0% −1.8 (−3.9 to +0.3) 12/5 72.0% −1.8 (−4.2 to +0.7) 15/8

For this build no difference is detectable on any set, but equivalence within ±1 point cannot be shown at these sample sizes. The noise floor changes no answer. These are likelihood scores without reasoning: they show what the quantization changes and are not comparable with the scores on the original model card. Method: quality-method.md.

Time to first token (2026-10-08)

On Apple Silicon at long context, the prefill is what you wait for.

context cold prefill rate follow-up turn, prefix-cache hit
1k 0.8 s 1,620 t/s 0.35 s
8k 5.3 s 1,590 t/s 0.4 s
32k 24 s 1,380 t/s 0.6 s
64k 59 s 1,090 t/s 1.4 s*
96k 109 s 885 t/s –

mlx-vlm server, PREFILL_STEP 2048, exact APC, reasoning_effort=low. About two minutes before the first token at 96k cold. * With APC_MEMORY_RESERVE_GB=1.5: mlx-vlm's automatic 4 GiB reserve is too large next to 33 GB of weights on a 48 GB Mac, so follow-up turns from ~32k tokens up missed the cache; 96k not re-measured. With reasoning_effort=none follow-up turns miss the cache; see the linked notes.

Decode (2026-10-03)

~70 tokens/s at short context, 57 t/s at 23k, 40 t/s at 96k.

Memory

Peak 35.3 GB on short prompts, 38.0 GiB at 96k tokens.

Correctness

The port's forward pass matches a reference implementation of the vLLM semantics to 1e-5 (on CPU). Needle retrieval succeeded at 23k and 96k tokens; tool calls work through <tool_call> + JSON.

License

Apache 2.0, same as the original model. See LICENSE. Kolibri 1 was developed by Aleph Alpha Research GmbH.

Downloads last month
1,334
Safetensors
Model size
78B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for here-be-dragons-ai/Kolibri-1-MLX-3bit

Quantized
(20)
this model