How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "Sawfwair/Kolibri-1-MLX-Mixed-2bit"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default Sawfwair/Kolibri-1-MLX-Mixed-2bit
Run Hermes
hermes
Quick Links

Kolibri-1 MLX Mixed 2-bit

This is a native mere.run Swift/MLX artifact for Aleph Alpha's Kolibri-1. It contains 24.66 GB of logical tensor storage and is the candidate for a 36 GB Apple Silicon machine. Full-checkpoint Apple memory fit and throughput remain unverified. It requires a development build with the native Kolibri runtime (source snapshot 8bfa3f5d6d9f23097cd1446935acc4589ce6fa72).

Quantization and source

The source is Aleph-Alpha/Kolibri-1-BF16 at 7a8f290e7858825c3cf5e4c447ba68345de9f1d3. Routed experts use MLX affine 2-bit/group-128 weights. Attention and shared experts use 8-bit/group-64. Embeddings, learned norms, routers, and the vocabulary head retain source precision. Router correction biases are converted losslessly to FP32; routing and logits compute in FP32. No retraining or instruction fine-tuning occurred.

Expert banks are stacked in numeric order, and config.json records every projection's policy. Use the native mere.run Kolibri loader. This custom layout is not a generic mlx-lm or Transformers checkpoint and is not the upstream vLLM tensor layout.

Measured diagnostic

Paired native BF16 and quantized scoring ran on a Runpod NVIDIA B200 with the same token sequences, score boundaries, and teacher-forced continuations. The suite contains 16 cases, including 3 calibration cases and 13 heldout cases. The heldout split has only 198 target tokens; it is a small diagnostic rather than a general capability benchmark.

Heldout measurement Result
Mean full-vocabulary KL 0.026801
Top-token agreement with BF16 97.47%
Target perplexity increase 3.46%
Peak CUDA MLX allocation 26.57 GB

Overall, English, and German diagnostic gates passed independently. Calibration KL selected this baseline over the activation-weighted fitted candidate. The fit improved target perplexity and top-token agreement but worsened full-distribution KL. Peak MLX allocation excludes other process/system memory and does not prove Apple memory fit. All diagnostic cases disable thinking; reasoning quality and large-context quality remain unmeasured here.

KOLIBRI_NATIVE_DIAGNOSTIC.json contains source/binary hashes, exact sequences, conversion manifests, fitting statistics, and per-case results for all variants. KOLIBRI_QUALIFICATION.json binds this artifact to its scoring receipt. The conversion manifest retains quality_qualified: false; public availability does not establish general accuracy or register a managed mere.run download.

Run

Download this repository, then pass the local directory to a mere.run build containing the native Kolibri runtime:

hf download Sawfwair/Kolibri-1-MLX-Mixed-2bit \
  --local-dir ./Kolibri-1-MLX-Mixed-2bit

mere.run text chat --model ./Kolibri-1-MLX-Mixed-2bit \
  --prompt "Erkl盲re, wie ein Regenbogen entsteht." \
  --context-size 8192 --max-tokens 256

mere.run api serve --engine text-chat-kolibri \
  --model ./Kolibri-1-MLX-Mixed-2bit --context-size 8192

The runtime uses the source tokenizer/chat template, supports tool messages and opt-in token logprobs, and defaults native chat to an 8,192-token budget. The 262,144-token limit in the pinned configuration is not a memory-fit claim. LoRA, constrained JSON, KV quantization, prefix reuse, and continuous batching are not implemented for this family.

License and provenance

Apache 2.0; see LICENSE, MODIFICATIONS.md, and UPSTREAM_MODEL_CARD.md. Aleph Alpha's upstream model card describes the original model's intended use, training, and limitations. KOLIBRI_CONVERSION.json records pinned source and output-shard SHA-256 values; SHA256SUMS covers the published bundle.

Downloads last month
423
Safetensors
Model size
7B params
Tensor type
BF16
路
U32
路
F32
路
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for Sawfwair/Kolibri-1-MLX-Mixed-2bit

Quantized
(14)
this model