--- license: apache-2.0 base_model: - JetBrains/Mellum2.1-12B-A2.5B-Thinking library_name: mlx tags: - text-generation-inference - edit-prediction - next-edit-suggestion pipeline_tag: text-generation --- # randmaru/Mellum2.1-12B-A2.5B-Thinking-mlx-mxfp4 This is an MXFP4 MLX quantization of [JetBrains/Mellum2.1-12B-A2.5B-Thinking](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking) for Apple Silicon inference. The base model is a 12B mixture-of-experts decoder (2.5B active parameters, 64 experts / top-8, 28 layers, 128K context). Weights are quantized to 4-bit MXFP4 with group size 32; the router gates are kept at 8-bit (group size 64). The result is a single `model.safetensors` of 6.46 GB, averaging ≈4.25 bits per weight. --- **MXFP4 vs 4Bit Quantization Comparison** | Parameter | `MXFP4` | `4Bit` (GGUF `Q4_K_M`) | |----------|---------|--------| | Reference | this repo (MLX) | [JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF) | | Quantization format | 4‑bit floating point with microscaling, group 32, shared exponent E8M0 | 4‑bit k‑quant (mixed 4/6‑bit blocks) | | Tensor types | U32 (packed weights), U8 (scales), BF16 | GGUF k‑quant blocks with FP scales | | Weights file size | **≈6.46 GB** (6,457,111,871 bytes) | ≈8.07 GB (8,071,295,264 bytes) | | Total download size | **≈6.46 GB** (6,464,259,055 bytes) | ≈8.07 GB (8,071,295,264 bytes) | | Effective bits per weight | **4.25** | ≈5.3 | | Router gates | kept at 8‑bit (group 64) | folded into the k‑quant | | Runtime | `mlx-lm`, etc. | `llama.cpp`, etc. | | Hardware support | Apple Silicon GPU via MLX (Metal) | Metal, CUDA, ROCm, CPU | | Quality | good for its bit‑rate (4.25 bpw) | higher bit‑rate (5.3 bpw) tends to preserve accuracy better on outliers | **Key takeaways:** - **Parameter size:** The MXFP4 `safetensors` is ≈1.6 GB smaller than the 4‑bit GGUF `Q4_K_M` (≈6.46 GB vs ≈8.07 GB). Most of the difference is bit‑rate, not packing magic: MXFP4 averages 4.25 bits/weight while `Q4_K_M` keeps several tensors at higher precision (≈5.3 bits/weight overall). Treat it as a size/quality trade‑off. - **Total download:** MXFP4 ships as a few small files (config, tokenizer, one `model.safetensors`) totalling ≈6.46 GB, versus the single ≈8.07 GB GGUF. - **Apple Silicon:** MXFP4 runs natively on the Apple GPU through MLX (Metal); the GGUF build runs through llama.cpp's Metal backend. - **Quality:** the higher bit‑rate of `Q4_K_M` generally retains a little more accuracy; MXFP4 trades some precision for a smaller footprint. Evaluate on your own task. - **GGUF alternative:** the same base model is also available as `MXFP4_MOE` (7.03 GB, ≈4.6 bits/weight) in the GGUF repo. > [!Note] Actual speed depends on the backend, GPU, batch size, and quantization implementation.