Instructions to use imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Mellum2.1 Thinking, MLX 6-bit (group 64), benchmarked
JetBrains' Mellum2.1 Thinking (12B MoE, 2.5B active) converted to native MLX 6-bit affine, group size 64, with the upstream 8-bit router policy. JetBrains created and trained the model. This release adds a measured comparison against the BF16 original and other quantizations, plus serving settings tested for coding agents. No fine-tuning or calibration.
What the measurements show
164 HumanEval+ tasks (expanded tests), same RTX 5090, same prompts, upstream chat template, greedy decoding, 8,192 completion tokens, pinned EvalPlus container. One sample per task.
| Candidate | Weights | HumanEval+ | Solved within 4,096 tokens | Median completion tokens | Reasoning length vs BF16 (paired) |
|---|---|---|---|---|---|
| BF16 original | 24.3 GB | 153/164 (93.3%) | 148 | 1,571 | 1.00x |
| This 6-bit | 9.87 GB | 146/164 (89.0%) | 141 | 1,521 | 1.04x |
| randmaru MXFP4 | 6.46 GB | 144/164 (87.8%) | 139 | 1,544 | 0.98x |
| Our naive affine 4-bit (not released) | 6.84 GB | 133/164 (81.1%) | 91 | 3,242 | 2.22x |
How to read this:
- 6-bit and MXFP4 are statistically tied: 9 tasks passed only 6-bit, 7 only MXFP4 (exact McNemar p = 0.80). This card does not claim 6-bit is better than MXFP4.
- Both are measurably below BF16 (6-bit: 1 vs 8 discordant tasks, p = 0.039).
- Both keep BF16's reasoning length. Our plain affine 4-bit did not: it reasoned 2.2x longer on the same tasks and ran out of budget mid-thought on 18. That is why it isn't released. The cause is under investigation.
- These are single-run regression results. Training overlap with HumanEval+ is possible. They do not measure agent reliability.
Raw responses, per-task outcomes, provenance and scripts: github.com/imaddde867/mellum-mlx.
Which Mac
- 24 GB unified memory or more: recommended.
- 16 GB: not recommended. Measured MLX peak: 10.31 GB at 4K and 10.51 GB at 16K. Swap used: 282.81 MB before the 6-bit sequence, 1,947.31 MB before 16K and 1,883.31 MB after 16K. Swap grew materially across the overall sequence, but decreased during the sampled 16K stage. Editor-open status was not recorded.
- Apple M4 (16 GB), warmup excluded, three measured trials: decode 48.0 / 47.1 / 43.5 tok/s, TTFT 1.609 / 7.099 / 32.450 s at 1K / 4K / 16K, respectively. Greedy decoding, 256 output tokens, AC power.
- Long context is cheap on memory. 21 of 28 layers use a 1,024-token sliding window, so the KV cache is about 14 KB per token (about 0.5 GB at 32K).
Use it
python -m pip install 'mlx==0.32.3' 'mlx-lm==0.32.0'
# Chat (JetBrains' recommended sampling)
python -m mlx_lm generate --model imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 \
--prompt 'Find the bug: def mean(xs): return sum(xs) / len(xs) - 1' \
--temp 0.6 --top-p 0.95 --top-k 20 --max-tokens 16384
# OpenAI-compatible server for coding agents
python -m mlx_lm server --model imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 \
--temp 1.0 --max-tokens 16384
- Set
--max-tokens. The MLX-LM server defaults to 512 tokens. That cuts a thinking model off before it answers or finishes a tool call. Reasoning and the answer share the budget. - Don't use greedy (
--temp 0) for real work. JetBrains uses temperature 0.6 / top-p 0.95 / top-k 20 for chat, and temperature 1.0 with up to 16K tokens per turn for agents.
Known issues with MLX-LM 0.32.0 serving
- Tool calls that fail to parse are dropped silently. That covers invalid JSON, and also valid-looking JSON with a literal newline inside a string. The response still says
finish_reason: "tool_calls"buttool_callsis empty, so agents can't see or repair the mistake. In our trials this happened twice, both times from an unescaped quote in code. A tested patch that returns the failed call text to the client is in the GitHub repo. - Reasoning isn't returned to the template. Mellum's template replays prior reasoning in tool loops only from
reasoning_content. The server returns it asreasoning. Clients that echo assistant messages back unchanged lose the model's earlier thinking between tool steps.
Repository-agent trial (one repo, not a reliability claim)
Five bounded trials on one small Python repository, run on CUDA through the MLX-LM 0.32.0 server. Greedy decoding, 8,192 tokens per turn, up to 12 turns, fixed tests in an offline container.
- Solved (1): a retrieval-limit type check. A 2-line patch; 36/36 fixed tests pass.
- Not solved (4): a mutation-routing fix. What went wrong:
- In two turns, the model emitted invalid tool-call JSON (an unescaped quote inside code). The server discarded it silently, still reporting
finish_reason: "tool_calls". One of those ended a trial. - One edit introduced a syntax error the model didn't repair.
- One fix used substring matching, which misses "changing" and "deleting".
- Two turns spent the whole 8,192-token budget thinking.
- In two turns, the model emitted invalid tool-call JSON (an unescaped quote inside code). The server discarded it silently, still reporting
- Harness limitations in these runs: greedy decoding, prior reasoning not passed back between steps, and file contents double-escaped in the early trials. The post-publication rerun used a separate patched MLX-LM environment, temperature 1.0, 16K tokens per turn and
reasoning_contentreplay: 1 success out of 3 seeds (0, 1, 2). Seed 1 passed all 36 fixed tests; seed 0 exhausted its turn budget and seed 2 left a syntax error followed by an invalid repair call. All rerun transcripts and outcomes are retained. This is one task, not a general reliability estimate.
Every transcript is retained in the GitHub repo. This is not an agent success rate.
Provenance
- Source:
JetBrains/Mellum2.1-12B-A2.5B-Thinking@92ddae9fc7665e9f801d141d2e5a6b2caf2460c4, Apache-2.0. - Quantization: affine 6-bit, group 64; routers 8-bit; other weights BF16. Produced with
mlx_lm.convert, the standard recipe; no novel method. - Serialized weights: 9,873,101,418 bytes. Conversion receipt with SHA-256 for every file: conversion.json.
- Runtime tested: MLX 0.32.3, MLX-LM 0.32.0, Python 3.12.
- Downloads last month
- -
6-bit
Model tree for imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64
Base model
JetBrains/Mellum2-12B-A2.5B-Base
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True)