Instructions to use Sawfwair/Kolibri-1-MLX-Mixed-2bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Sawfwair/Kolibri-1-MLX-Mixed-2bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Sawfwair/Kolibri-1-MLX-Mixed-2bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Sawfwair/Kolibri-1-MLX-Mixed-2bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Sawfwair/Kolibri-1-MLX-Mixed-2bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Sawfwair/Kolibri-1-MLX-Mixed-2bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Sawfwair/Kolibri-1-MLX-Mixed-2bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Sawfwair/Kolibri-1-MLX-Mixed-2bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Sawfwair/Kolibri-1-MLX-Mixed-2bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sawfwair/Kolibri-1-MLX-Mixed-2bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Sawfwair/Kolibri-1-MLX-Mixed-2bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Sawfwair/Kolibri-1-MLX-Mixed-2bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Sawfwair/Kolibri-1-MLX-Mixed-2bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Sawfwair/Kolibri-1-MLX-Mixed-2bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Sawfwair/Kolibri-1-MLX-Mixed-2bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Sawfwair/Kolibri-1-MLX-Mixed-2bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure the model in Pi
# Install Pi:
npm install -g @earendil-works/pi-coding-agent# Add to ~/.pi/agent/models.json:
{
"providers": {
"mlx-lm": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "none",
"models": [
{
"id": "Sawfwair/Kolibri-1-MLX-Mixed-2bit"
}
]
}
}
}Run Pi
# Start Pi in your project directory:
piKolibri-1 MLX Mixed 2-bit
This is a native mere.run Swift/MLX artifact for Aleph Alpha's Kolibri-1.
It contains 24.66 GB of logical tensor storage and is the candidate for a 36 GB
Apple Silicon machine. Full-checkpoint Apple memory fit and throughput remain
unverified. It requires a development build with the native Kolibri runtime
(source snapshot 8bfa3f5d6d9f23097cd1446935acc4589ce6fa72).
Quantization and source
The source is Aleph-Alpha/Kolibri-1-BF16 at
7a8f290e7858825c3cf5e4c447ba68345de9f1d3. Routed experts use MLX affine
2-bit/group-128 weights. Attention and shared experts use 8-bit/group-64.
Embeddings, learned norms, routers, and the vocabulary head retain source
precision. Router correction biases are converted losslessly to FP32; routing
and logits compute in FP32. No retraining or instruction fine-tuning occurred.
Expert banks are stacked in numeric order, and config.json records every
projection's policy. Use the native mere.run Kolibri loader. This custom layout
is not a generic mlx-lm or Transformers checkpoint and is not the upstream
vLLM tensor layout.
Measured diagnostic
Paired native BF16 and quantized scoring ran on a Runpod NVIDIA B200 with the same token sequences, score boundaries, and teacher-forced continuations. The suite contains 16 cases, including 3 calibration cases and 13 heldout cases. The heldout split has only 198 target tokens; it is a small diagnostic rather than a general capability benchmark.
| Heldout measurement | Result |
|---|---|
| Mean full-vocabulary KL | 0.026801 |
| Top-token agreement with BF16 | 97.47% |
| Target perplexity increase | 3.46% |
| Peak CUDA MLX allocation | 26.57 GB |
Overall, English, and German diagnostic gates passed independently. Calibration KL selected this baseline over the activation-weighted fitted candidate. The fit improved target perplexity and top-token agreement but worsened full-distribution KL. Peak MLX allocation excludes other process/system memory and does not prove Apple memory fit. All diagnostic cases disable thinking; reasoning quality and large-context quality remain unmeasured here.
KOLIBRI_NATIVE_DIAGNOSTIC.json contains source/binary hashes, exact sequences,
conversion manifests, fitting statistics, and per-case results for all variants.
KOLIBRI_QUALIFICATION.json binds this artifact to its scoring receipt.
The conversion manifest retains quality_qualified: false; public availability
does not establish general accuracy or register a managed mere.run download.
Run
Download this repository, then pass the local directory to a mere.run build containing the native Kolibri runtime:
hf download Sawfwair/Kolibri-1-MLX-Mixed-2bit \
--local-dir ./Kolibri-1-MLX-Mixed-2bit
mere.run text chat --model ./Kolibri-1-MLX-Mixed-2bit \
--prompt "Erkl盲re, wie ein Regenbogen entsteht." \
--context-size 8192 --max-tokens 256
mere.run api serve --engine text-chat-kolibri \
--model ./Kolibri-1-MLX-Mixed-2bit --context-size 8192
The runtime uses the source tokenizer/chat template, supports tool messages and opt-in token logprobs, and defaults native chat to an 8,192-token budget. The 262,144-token limit in the pinned configuration is not a memory-fit claim. LoRA, constrained JSON, KV quantization, prefix reuse, and continuous batching are not implemented for this family.
License and provenance
Apache 2.0; see LICENSE, MODIFICATIONS.md, and UPSTREAM_MODEL_CARD.md.
Aleph Alpha's upstream model card describes the original model's intended use,
training, and limitations. KOLIBRI_CONVERSION.json records pinned source and
output-shard SHA-256 values; SHA256SUMS covers the published bundle.
- Downloads last month
- 423
2-bit
Model tree for Sawfwair/Kolibri-1-MLX-Mixed-2bit
Base model
Aleph-Alpha/Kolibri-1-BF16
Start the MLX server
# Install MLX LM: uv tool install mlx-lm# Start a local OpenAI-compatible server: mlx_lm.server --model "Sawfwair/Kolibri-1-MLX-Mixed-2bit"