prithivMLmods commited on
Commit
d8de093
·
verified ·
1 Parent(s): e2ee80f

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +33 -1
README.md CHANGED
@@ -1,8 +1,40 @@
1
  ---
2
  base_model:
3
  - JetBrains/Mellum2.1-12B-A2.5B-Thinking
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  ---
5
 
6
  # **JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF**
7
 
8
- > **Mellum2.1 Thinking** is JetBrains' updated reasoning model, a 12B-parameter mixture-of-experts with 2.5B active parameters (28 layers, 64 experts with 8 active, 131,072-token context, Apache 2.0). It is the successor to Mellum2 Thinking with the architecture unchanged, and nearly all of the improvement comes from post-training, where reinforcement learning grew from a short final stage into the main part of training. It uses new RL tasks in math, competitive programming, science, tool use, and software engineering, each source filtered before training. For software engineering, it trained in real repositories with a shell and file-editing tools, rewarded when tests pass, over millions of sandboxed runs. The gains are largest on agentic work: SWE-bench Verified rises from 2.0 to 47.0, Terminal-Bench 2.1 from 0.6 to 17.4, and SWE-bench Pro from 0.0 to 28.0. LiveCodeBench v6 climbs from 69.4 to 82.0, AIME 25/26 from 60.1 to 83.3, and BFCL v4 from 49.6 to 62.3. Against Qwen3.5-9B it leads on LiveCodeBench v6 (82.0 vs 75.4), HumanEval+, MBPP+, BFCL v4 and WorkBench, but trails on AIME, GPQA Diamond (64.6 vs 77.8), SWE-bench Verified and Pro, Terminal-Bench 2.1, and IFEval. Safety results are mixed: HarmBench harmful rate improves from 21.5 to 8.5, but XSTest safe compliance slips from 91.2 to 88.8. All numbers are self-reported by JetBrains, from a shared pipeline in thinking mode. It is intended for complex agentic tasks and hard coding, math, and reasoning problems, and is served via vLLM with the `qwen3` reasoning parser and optional Hermes-style tool calling.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  base_model:
3
  - JetBrains/Mellum2.1-12B-A2.5B-Thinking
4
+ license: apache-2.0
5
+ language:
6
+ - en
7
+ pipeline_tag: text-generation
8
+ library_name: transformers
9
+ tags:
10
+ - text-generation-inference
11
+ - JetBrains
12
+ - llama-cpp
13
+ - agentic
14
+ - agent
15
+ - reasoning
16
+ - math
17
+ - coding
18
+ - tool-calling
19
  ---
20
 
21
  # **JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF**
22
 
23
+ > **Mellum2.1 Thinking** is JetBrains' updated reasoning model, a 12B-parameter mixture-of-experts with 2.5B active parameters (28 layers, 64 experts with 8 active, 131,072-token context, Apache 2.0). It is the successor to Mellum2 Thinking with the architecture unchanged, and nearly all of the improvement comes from post-training, where reinforcement learning grew from a short final stage into the main part of training. It uses new RL tasks in math, competitive programming, science, tool use, and software engineering, each source filtered before training. For software engineering, it trained in real repositories with a shell and file-editing tools, rewarded when tests pass, over millions of sandboxed runs. The gains are largest on agentic work: SWE-bench Verified rises from 2.0 to 47.0, Terminal-Bench 2.1 from 0.6 to 17.4, and SWE-bench Pro from 0.0 to 28.0. LiveCodeBench v6 climbs from 69.4 to 82.0, AIME 25/26 from 60.1 to 83.3, and BFCL v4 from 49.6 to 62.3. Against Qwen3.5-9B it leads on LiveCodeBench v6 (82.0 vs 75.4), HumanEval+, MBPP+, BFCL v4 and WorkBench, but trails on AIME, GPQA Diamond (64.6 vs 77.8), SWE-bench Verified and Pro, Terminal-Bench 2.1, and IFEval. Safety results are mixed: HarmBench harmful rate improves from 21.5 to 8.5, but XSTest safe compliance slips from 91.2 to 88.8. All numbers are self-reported by JetBrains, from a shared pipeline in thinking mode. It is intended for complex agentic tasks and hard coding, math, and reasoning problems, and is served via vLLM with the `qwen3` reasoning parser and optional Hermes-style tool calling.
24
+
25
+ ## Model Files
26
+
27
+ | File Name | Quant Type | File Size | File Link | Description |
28
+ |-----------|------------|-----------|-----------|-------------|
29
+ | Mellum2.1-12B-A2.5B-Thinking.BF16.gguf | BF16 | 24.3 GB | [Link](https://huggingface.co/prithivMLmods/JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking.BF16.gguf) | Full BF16 weights. Highest quality, largest file size. |
30
+ | Mellum2.1-12B-A2.5B-Thinking.Q3_K_L.gguf | Q3_K_L | 6.59 GB | [Link](https://huggingface.co/prithivMLmods/JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking.Q3_K_L.gguf) | Lower quality but usable, good for low RAM availability. |
31
+ | Mellum2.1-12B-A2.5B-Thinking.Q3_K_M.gguf | Q3_K_M | 6.33 GB | [Link](https://huggingface.co/prithivMLmods/JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking.Q3_K_M.gguf) | Low quality. |
32
+ | Mellum2.1-12B-A2.5B-Thinking.Q4_K_M.gguf | Q4_K_M | 8.07 GB | [Link](https://huggingface.co/prithivMLmods/JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking.Q4_K_M.gguf) | Good quality, default size for most use cases, *recommended*. |
33
+ | Mellum2.1-12B-A2.5B-Thinking.Q4_K_S.gguf | Q4_K_S | 7.4 GB | [Link](https://huggingface.co/prithivMLmods/JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking.Q4_K_S.gguf) | Slightly lower quality with more space savings, *recommended*. |
34
+ | Mellum2.1-12B-A2.5B-Thinking.Q5_K_M.gguf | Q5_K_M | 9.21 GB | [Link](https://huggingface.co/prithivMLmods/JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking.Q5_K_M.gguf) | High quality, *recommended*. |
35
+ | Mellum2.1-12B-A2.5B-Thinking.Q5_K_S.gguf | Q5_K_S | 8.63 GB | [Link](https://huggingface.co/prithivMLmods/JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking.Q5_K_S.gguf) | High quality, *recommended*. |
36
+ | Mellum2.1-12B-A2.5B-Thinking.Q6_K.gguf | Q6_K | 10.9 GB | [Link](https://huggingface.co/prithivMLmods/JetBrains-Mellum2.1-12B-A2.5B-Thinking-GGUF/blob/main/Mellum2.1-12B-A2.5B-Thinking.Q6_K.gguf) | Very high quality, near perfect, *recommended*. |
37
+
38
+ ## llama.cpp
39
+
40
+ LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp