Mellum2.1 Thinking, MLX 6-bit (group 64), benchmarked

JetBrains' Mellum2.1 Thinking (12B MoE, 2.5B active) converted to native MLX 6-bit affine, group size 64, with the upstream 8-bit router policy. JetBrains created and trained the model. This release adds a measured comparison against the BF16 original and other quantizations, plus serving settings tested for coding agents. No fine-tuning or calibration.

What the measurements show

164 HumanEval+ tasks (expanded tests), same RTX 5090, same prompts, upstream chat template, greedy decoding, 8,192 completion tokens, pinned EvalPlus container. One sample per task.

Candidate Weights HumanEval+ Solved within 4,096 tokens Median completion tokens Reasoning length vs BF16 (paired)
BF16 original 24.3 GB 153/164 (93.3%) 148 1,571 1.00x
This 6-bit 9.87 GB 146/164 (89.0%) 141 1,521 1.04x
randmaru MXFP4 6.46 GB 144/164 (87.8%) 139 1,544 0.98x
Our naive affine 4-bit (not released) 6.84 GB 133/164 (81.1%) 91 3,242 2.22x

How to read this:

  • 6-bit and MXFP4 are statistically tied: 9 tasks passed only 6-bit, 7 only MXFP4 (exact McNemar p = 0.80). This card does not claim 6-bit is better than MXFP4.
  • Both are measurably below BF16 (6-bit: 1 vs 8 discordant tasks, p = 0.039).
  • Both keep BF16's reasoning length. Our plain affine 4-bit did not: it reasoned 2.2x longer on the same tasks and ran out of budget mid-thought on 18. That is why it isn't released. The cause is under investigation.
  • These are single-run regression results. Training overlap with HumanEval+ is possible. They do not measure agent reliability.

Raw responses, per-task outcomes, provenance and scripts: github.com/imaddde867/mellum-mlx.

Which Mac

  • 24 GB unified memory or more: recommended.
  • 16 GB: not recommended. Measured MLX peak: 10.31 GB at 4K and 10.51 GB at 16K. Swap used: 282.81 MB before the 6-bit sequence, 1,947.31 MB before 16K and 1,883.31 MB after 16K. Swap grew materially across the overall sequence, but decreased during the sampled 16K stage. Editor-open status was not recorded.
  • Apple M4 (16 GB), warmup excluded, three measured trials: decode 48.0 / 47.1 / 43.5 tok/s, TTFT 1.609 / 7.099 / 32.450 s at 1K / 4K / 16K, respectively. Greedy decoding, 256 output tokens, AC power.
  • Long context is cheap on memory. 21 of 28 layers use a 1,024-token sliding window, so the KV cache is about 14 KB per token (about 0.5 GB at 32K).

Use it

python -m pip install 'mlx==0.32.3' 'mlx-lm==0.32.0'

# Chat (JetBrains' recommended sampling)
python -m mlx_lm generate --model imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 \
  --prompt 'Find the bug: def mean(xs): return sum(xs) / len(xs) - 1' \
  --temp 0.6 --top-p 0.95 --top-k 20 --max-tokens 16384

# OpenAI-compatible server for coding agents
python -m mlx_lm server --model imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64 \
  --temp 1.0 --max-tokens 16384
  • Set --max-tokens. The MLX-LM server defaults to 512 tokens. That cuts a thinking model off before it answers or finishes a tool call. Reasoning and the answer share the budget.
  • Don't use greedy (--temp 0) for real work. JetBrains uses temperature 0.6 / top-p 0.95 / top-k 20 for chat, and temperature 1.0 with up to 16K tokens per turn for agents.

Known issues with MLX-LM 0.32.0 serving

  • Tool calls that fail to parse are dropped silently. That covers invalid JSON, and also valid-looking JSON with a literal newline inside a string. The response still says finish_reason: "tool_calls" but tool_calls is empty, so agents can't see or repair the mistake. In our trials this happened twice, both times from an unescaped quote in code. A tested patch that returns the failed call text to the client is in the GitHub repo.
  • Reasoning isn't returned to the template. Mellum's template replays prior reasoning in tool loops only from reasoning_content. The server returns it as reasoning. Clients that echo assistant messages back unchanged lose the model's earlier thinking between tool steps.

Repository-agent trial (one repo, not a reliability claim)

Five bounded trials on one small Python repository, run on CUDA through the MLX-LM 0.32.0 server. Greedy decoding, 8,192 tokens per turn, up to 12 turns, fixed tests in an offline container.

  • Solved (1): a retrieval-limit type check. A 2-line patch; 36/36 fixed tests pass.
  • Not solved (4): a mutation-routing fix. What went wrong:
    • In two turns, the model emitted invalid tool-call JSON (an unescaped quote inside code). The server discarded it silently, still reporting finish_reason: "tool_calls". One of those ended a trial.
    • One edit introduced a syntax error the model didn't repair.
    • One fix used substring matching, which misses "changing" and "deleting".
    • Two turns spent the whole 8,192-token budget thinking.
  • Harness limitations in these runs: greedy decoding, prior reasoning not passed back between steps, and file contents double-escaped in the early trials. The post-publication rerun used a separate patched MLX-LM environment, temperature 1.0, 16K tokens per turn and reasoning_content replay: 1 success out of 3 seeds (0, 1, 2). Seed 1 passed all 36 fixed tests; seed 0 exhausted its turn budget and seed 2 left a syntax error followed by an invalid repair call. All rerun transcripts and outcomes are retained. This is one task, not a general reliability estimate.

Every transcript is retained in the GitHub repo. This is not an agent success rate.

Provenance

  • Source: JetBrains/Mellum2.1-12B-A2.5B-Thinking @ 92ddae9fc7665e9f801d141d2e5a6b2caf2460c4, Apache-2.0.
  • Quantization: affine 6-bit, group 64; routers 8-bit; other weights BF16. Produced with mlx_lm.convert, the standard recipe; no novel method.
  • Serialized weights: 9,873,101,418 bytes. Conversion receipt with SHA-256 for every file: conversion.json.
  • Runtime tested: MLX 0.32.3, MLX-LM 0.32.0, Python 3.12.

JetBrains original · MLX-LM · randmaru MXFP4

Downloads last month
-
Safetensors
Model size
12B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-6bit-g64