--- base_model: JetBrains/Mellum2.1-12B-A2.5B-Thinking base_model_relation: quantized library_name: gguf pipeline_tag: text-generation language: - en tags: - mellum - gguf - llama.cpp - quantized - moe - thinking license: apache-2.0 --- # Mellum2.1 Thinking — GGUF This repository contains **GGUF** builds of [`JetBrains/Mellum2.1-12B-A2.5B-Thinking`](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking) in every quantization we publish, ready to run with [`llama.cpp`](https://github.com/ggml-org/llama.cpp), Ollama, LM Studio, and other GGUF-compatible runtimes. Each variant is one file; download the one you need. Mellum2.1 Thinking is a Mixture-of-Experts reasoning model (64 experts, 8 activated per token, 131,072-token context) that emits its chain of thought inside `...` blocks before the final answer. For the full model description, evaluation results, and architecture details, see the original model card: **[JetBrains/Mellum2.1-12B-A2.5B-Thinking](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking)**. ## Available quantizations | File | Quantization | Description | Size | KLD vs BF16 ↓ | Top-token match ↑ | |---|---|---|---|---|---| | [`Mellum2.1-12B-A2.5B-Thinking-BF16.gguf`](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking-BF16.gguf) | `BF16` | 16-bit, no quantization (reference) | 24.3 GB | — | — | | [`Mellum2.1-12B-A2.5B-Thinking-Q8_0.gguf`](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking-Q8_0.gguf) | `Q8_0` | 8-bit, effectively lossless | 12.9 GB | 0.008 | 96.1% | | [`Mellum2.1-12B-A2.5B-Thinking-Q6_K.gguf`](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking-Q6_K.gguf) | `Q6_K` | 6-bit k-quant, very high quality | 10.9 GB | 0.020 | 93.9% | | [`Mellum2.1-12B-A2.5B-Thinking-Q4_K_M.gguf`](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking-Q4_K_M.gguf) | **`Q4_K_M`** | 4-bit k-quant, balanced (recommended) | 8.1 GB | 0.075 | 88.0% | | [`Mellum2.1-12B-A2.5B-Thinking-MXFP4_MOE.gguf`](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking-MXFP4_MOE.gguf) | `MXFP4_MOE` | MXFP4 4-bit on MoE experts, smallest | 7.0 GB | 0.115 | 85.6% | KL divergence and top-token agreement are measured against the BF16 logits on Wikitext-2 (`n_ctx=512`); lower KLD / higher agreement means closer to the unquantized model. ## Download ```sh hf download JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF Mellum2.1-12B-A2.5B-Thinking-Q4_K_M.gguf --local-dir . ``` ## Run with llama.cpp ```sh # Pull and serve in one step (downloads the GGUF automatically; the suffix picks the quantization) llama-server -hf JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M --ctx-size 131072 --temp 0.6 --top-p 0.95 --top-k 20 # Or run a one-off prompt with a local file llama-cli -m Mellum2.1-12B-A2.5B-Thinking-Q4_K_M.gguf --ctx-size 131072 --temp 0.6 --top-p 0.95 --top-k 20 -p "Is 1024 a power of 2? Explain your reasoning." ``` The server exposes an OpenAI-compatible API on `http://localhost:8080/v1`: ```python from openai import OpenAI client = OpenAI(base_url="http://localhost:8080/v1", api_key="llama.cpp") chat_response = client.chat.completions.create( model="JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF", messages=[ {"role": "user", "content": "Is 1024 a power of 2? Explain your reasoning."}, ], max_tokens=81920, temperature=0.6, top_p=0.95, extra_body={"top_k": 20}, ) print(chat_response.choices[0].message.content) ``` ## Run with Ollama ```sh ollama run hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M ``` ## License Released under the Apache 2.0 license. --- *For the full model card, evaluation results, and architecture details, refer to the original model: [JetBrains/Mellum2.1-12B-A2.5B-Thinking](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking).*