Kolibri-1 MLX 3-bit

A mixed 3/6-bit MLX quantization of Aleph-Alpha/Kolibri-1, Aleph Alpha's 78B-A3.5B mixture-of-experts reasoning model for German and English, sized to run on a 48 GB Apple Silicon Mac.

This is a community conversion by here-be-dragons.ai, not an official Aleph Alpha release. For the model itself (training, evaluations, intended use, limitations) see the original model card and the tech report.

Quantization

Part Precision
Routed experts (75.5B of 78.1B parameters) 3 bit affine, group size 64
Attention, shared expert, embedding, LM head 6 bit affine, group size 64
MoE router (mlp.gate) bf16, as in the release; expert_bias stays fp32

3.61 bits per weight, 33 GiB on disk. The block-FP8 release weights were dequantized to bf16 and quantized once, with no intermediate format. A uniform 4-bit version would be about 44 GB and does not fit on a 48 GB machine.

Requirements

The kolibri1 architecture is not yet part of a released mlx-vlm or mlx-lm. Until the port is merged upstream, install mlx-vlm from the kolibri1 branch of our fork:

pip install git+https://github.com/here-be-dragons-ai/mlx-vlm@kolibri1

mlx-lm support is pending upstream review.

The same weights load in both mlx-vlm and mlx-lm.

  • Apple Silicon with 48 GB unified memory or more
  • On 48 GB, raise the GPU wired-memory limit, since the weights alone are 32.8 GiB: sudo sysctl -w iogpu.wired_limit_mb=40960
  • Do not run another large model at the same time.

Usage (mlx-vlm)

from mlx_vlm import load, stream_generate

model, processor = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
tokenizer = processor.tokenizer
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
for chunk in stream_generate(model, processor, prompt, max_tokens=2048,
                             temperature=1.0, top_p=0.97, top_k=128):
    print(chunk.text, end="", flush=True)

OpenAI-compatible server:

python -m mlx_vlm.server --model here-be-dragons-ai/Kolibri-1-MLX-3bit --port 8080

Usage (mlx-lm)

from mlx_lm import load, stream_generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
sampler = make_sampler(temp=1.0, top_p=0.97, top_k=128)
for chunk in stream_generate(model, tokenizer, prompt, max_tokens=2048, sampler=sampler):
    print(chunk.text, end="", flush=True)

Server notes (mlx-vlm)

Reasoning is controlled with the top-level request fields reasoning_effort or enable_thinking; the server ignores them inside chat_template_kwargs. The thinking comes back in reasoning, separate from content.

Recommended sampling, from the original release: temperature=1.0, top_p=0.97, top_k=128 (also set in generation_config.json).

The chat template supports Kolibri's reasoning mode: pass reasoning_effort (none, low, medium, high) to apply_chat_template. Without it, the model does not think.

Measurements

M5 Pro, 48 GB, iogpu.wired_limit_mb=40960, mlx 0.32.2, mlx-vlm 0.7.4. Method, scripts and raw data: local-sovereign-mlx/docs/kolibri-quality.

Quality (2026-10-08)

Measured against the FP8 release through the same harness: all arms run the same layer-streamed forward pass, logits in fp32 from each arm's own LM head. The uniform 3-bit row is a control built for this comparison, not a release.

this build (3/6-bit) uniform 3-bit
bits per weight 3.61 3.51
KL vs FP8, mean / median / p99 0.114 / 0.021 / 2.07 0.376 / 0.138 / 5.42
top-1 token agreement with FP8 88.9% 76.5%
perplexity (FP8: 16.98) 17.11 19.38

Text: 76k tokens of German and English Wikipedia prose at pinned revisions, part of it written after the model's knowledge cutoff.

Multiple choice, reasoning_effort=none, scored by letter logits (paired against FP8, exact McNemar test):

benchmark FP8 this build uniform 3-bit
Belebele de (900) 92.9% 93.1% (+0.2, p=0.80) 92.0% (−0.9, p=0.23)
Belebele en (900) 95.2% 94.8% (−0.4, p=0.45) 93.6% (−1.7, p=0.006)
Global-MMLU-Lite de (400) 72.5% 71.2% (−1.2, p=0.30) 68.5% (−4.0, p=0.011)
Global-MMLU-Lite en (400) 73.8% 72.0% (−1.8, p=0.14) 72.0% (−1.8, p=0.21)

None of this build's differences is significant at these sample sizes; a loss of 1-2 points on Global-MMLU-Lite cannot be ruled out. These are likelihood scores without reasoning: they show what the quantization changes, and are not comparable with the scores on the original model card.

Time to first token (2026-10-08)

On Apple Silicon at long context, the prefill is what you wait for.

context cold prefill rate follow-up turn, prefix-cache hit
1k 0.8 s 1,620 t/s 0.35 s
8k 5.3 s 1,590 t/s 0.4 s
32k 24 s 1,380 t/s 0.6 s
64k 59 s 1,090 t/s –
96k 109 s 885 t/s –

mlx-vlm server, PREFILL_STEP 2048, exact APC, reasoning_effort=low. About two minutes before the first token at 96k cold. In mlx-vlm 0.7.4 the prefix cache stops hitting after a prefill of 64k tokens or more until the server restarts, and with reasoning_effort=none follow-up turns miss it; see the linked notes.

Decode (2026-10-03)

~70 tokens/s at short context, 57 t/s at 23k, 40 t/s at 96k.

Memory

Peak 35.3 GB on short prompts, 38.0 GiB at 96k tokens.

Correctness

The port's forward pass matches a reference implementation of the vLLM semantics to 1e-5 (on CPU). Needle retrieval succeeded at 23k and 96k tokens; tool calls work through <tool_call> + JSON.

License

Apache 2.0, same as the original model. See LICENSE. Kolibri 1 was developed by Aleph Alpha Research GmbH.

Downloads last month
1,262
Safetensors
Model size
78B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for here-be-dragons-ai/Kolibri-1-MLX-3bit

Quantized
(19)
this model