randmaru/Mellum2.1-12B-A2.5B-Thinking-mlx-mxfp4

This is an MXFP4 MLX quantization of JetBrains/Mellum2.1-12B-A2.5B-Thinking for Apple Silicon inference.

The base model is a 12B mixture-of-experts decoder (2.5B active parameters, 64 experts / top-8, 28 layers, 128K context). Weights are quantized to 4-bit MXFP4 with group size 32; the router gates are kept at 8-bit (group size 64). The result is a single model.safetensors of 6.46 GB, averaging ≈4.25 bits per weight.


MXFP4 vs 4Bit Quantization Comparison

Parameter MXFP4 4Bit (GGUF Q4_K_M)
Reference this repo (MLX) JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF
Quantization format 4‑bit floating point with microscaling, group 32, shared exponent E8M0 4‑bit k‑quant (mixed 4/6‑bit blocks)
Tensor types U32 (packed weights), U8 (scales), BF16 GGUF k‑quant blocks with FP scales
Weights file size ≈6.46 GB (6,457,111,871 bytes) ≈8.07 GB (8,071,295,264 bytes)
Total download size ≈6.46 GB (6,464,259,055 bytes) ≈8.07 GB (8,071,295,264 bytes)
Effective bits per weight 4.25 ≈5.3
Router gates kept at 8‑bit (group 64) folded into the k‑quant
Runtime mlx-lm, etc. llama.cpp, etc.
Hardware support Apple Silicon GPU via MLX (Metal) Metal, CUDA, ROCm, CPU
Quality good for its bit‑rate (4.25 bpw) higher bit‑rate (5.3 bpw) tends to preserve accuracy better on outliers

Key takeaways:

  • Parameter size: The MXFP4 safetensors is ≈1.6 GB smaller than the 4‑bit GGUF Q4_K_M (≈6.46 GB vs ≈8.07 GB). Most of the difference is bit‑rate, not packing magic: MXFP4 averages 4.25 bits/weight while Q4_K_M keeps several tensors at higher precision (≈5.3 bits/weight overall). Treat it as a size/quality trade‑off.
  • Total download: MXFP4 ships as a few small files (config, tokenizer, one model.safetensors) totalling ≈6.46 GB, versus the single ≈8.07 GB GGUF.
  • Apple Silicon: MXFP4 runs natively on the Apple GPU through MLX (Metal); the GGUF build runs through llama.cpp's Metal backend.
  • Quality: the higher bit‑rate of Q4_K_M generally retains a little more accuracy; MXFP4 trades some precision for a smaller footprint. Evaluate on your own task.
  • GGUF alternative: the same base model is also available as MXFP4_MOE (7.03 GB, ≈4.6 bits/weight) in the GGUF repo.

Actual speed depends on the backend, GPU, batch size, and quantization implementation.

Downloads last month
36
Safetensors
Model size
12B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for randmaru/Mellum2.1-12B-A2.5B-Thinking-mlx-mxfp4