Instructions to use JetBrains/Mellum2.1-12B-A2.5B-Thinking with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JetBrains/Mellum2.1-12B-A2.5B-Thinking with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="JetBrains/Mellum2.1-12B-A2.5B-Thinking") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("JetBrains/Mellum2.1-12B-A2.5B-Thinking") model = AutoModelForCausalLM.from_pretrained("JetBrains/Mellum2.1-12B-A2.5B-Thinking", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use JetBrains/Mellum2.1-12B-A2.5B-Thinking with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "JetBrains/Mellum2.1-12B-A2.5B-Thinking" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JetBrains/Mellum2.1-12B-A2.5B-Thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking
- SGLang
How to use JetBrains/Mellum2.1-12B-A2.5B-Thinking with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "JetBrains/Mellum2.1-12B-A2.5B-Thinking" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JetBrains/Mellum2.1-12B-A2.5B-Thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "JetBrains/Mellum2.1-12B-A2.5B-Thinking" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JetBrains/Mellum2.1-12B-A2.5B-Thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use JetBrains/Mellum2.1-12B-A2.5B-Thinking with Docker Model Runner:
docker model run hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking
Mellum2.1 Thinking
Mellum2.1 is a thinking model. Use it for complex agentic tasks, such as working in a repository, running commands, and calling tools, and for hard non-agentic problems in coding, math, and reasoning.
Mellum2.1 Highlights
Mellum2.1 is the next version of Mellum2 Thinking. The architecture is unchanged (a 12B mixture-of-experts model with 2.5B active parameters), and almost all of the work for this version went into post-training, primarily reinforcement learning (RL):
- Reinforcement learning at a new scale: RL went from a short final stage to the main part of training, after many experiments on both the methods and the data.
- More data, filtered harder: new RL tasks in math, competitive programming, science, tool use, and software engineering, combining open RL datasets with tasks we built ourselves and filtering every source before training.
- Real environments for agentic skills: for software engineering, the model trains inside real repositories with a shell and file-editing tools and is rewarded when the tests pass.
As a result, Mellum2.1 handles agentic tasks much better than Mellum2. After millions of sandboxed runs in real environments during training, it explores a codebase, edits files, and checks its own changes.
Model Overview
Mellum2.1 Thinking has the following features:
- Number of Parameters: 12B total, 2.5B active
- Number of Layers: 28
- Hidden Size: 2304
- Intermediate Size: 7168
- MoE Intermediate Size: 896
- Number of Experts: 64
- Number of Activated Experts: 8
- Number of Attention Heads (GQA): 32 for Q and 4 for KV
- Context Length: 131,072
- Sliding Window: 1,024 (3 of every 4 layers)
- Vocabulary Size: 98,304
- Precision: bfloat16
- License: Apache 2.0
Serving with vLLM
# Without tool calling
vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
--max-model-len 131072 \
--reasoning-parser qwen3
# With tool calling
vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
--max-model-len 131072 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser hermes
GGUF builds for llama.cpp, Ollama, and LM Studio, as well as the multi-token prediction (MTP) head for speculative decoding in vLLM, are coming soon.
Quickstart
from openai import OpenAI
# Configured by environment variables
client = OpenAI()
messages = [
{"role": "user", "content": "Find the bug in this function and explain the fix: def mean(xs): return sum(xs) / len(xs) - 1"},
]
chat_response = client.chat.completions.create(
model="JetBrains/Mellum2.1-12B-A2.5B-Thinking",
messages=messages,
max_tokens=81920,
temperature=0.6,
top_p=0.95,
extra_body={"top_k": 20},
)
print("Chat response:", chat_response)
Evaluation
All values are percentages; higher is better except HarmBench, where lower is better. All models were evaluated by JetBrains with the same pipeline in thinking mode. All values are self-reported by JetBrains.
| Benchmark | Mellum2.1 Thinking | Mellum2 Thinking | Gemma 4 (E4B) | Qwen3.5 (9B) |
|---|---|---|---|---|
| Coding | ||||
| LiveCodeBench v6 | 82.0 | 69.4 | 69.4 | 75.4 |
| HumanEval+ | 91.5 | 90.9 | 89.1 | 89.6 |
| MBPP+ | 79.4 | 75.4 | 70.9 | 69.8 |
| Math | ||||
| AIME 25/26 | 83.3 | 60.1 | 45.0 | 86.7 |
| GSM-Plus | 88.3 | 87.1 | 87.4 | 91.4 |
| Agentic | ||||
| SWE-bench Verified | 47.0 | 2.0 | 23.0 | 50.0 |
| Terminal-Bench 2.1 | 17.4 | 0.6 | 3.4 | 21.7 |
| SWE-bench Pro | 28.0 | 0.0 | 4.0 | 38.0 |
| Tool Use | ||||
| BFCL v4 | 62.3 | 49.6 | 52.5 | 58.5 |
| WorkBench | 44.6 | 45.1 | 46.1 | 39.7 |
| ToolHop | 49.1 | 46.7 | 39.9 | 52.0 |
| Conversational | ||||
| IFEval | 90.6 | 79.5 | 90.8 | 92.4 |
| Knowledge | ||||
| GPQA Diamond | 64.6 | 51.0 | 53.1 | 77.8 |
| MMLU-Redux | 87.8 | 86.0 | 84.9 | 89.5 |
| MixEval-Hard | 46.4 | 41.6 | 41.4 | 50.3 |
| Safety | ||||
| XSTest | 88.8 | 91.2 | 95.2 | 95.6 |
| HarmBench (↓) | 8.5 | 21.5 | 38.1 | 6.6 |
Notes:
- Non-agentic benchmarks use greedy decoding.
- Agentic benchmarks use the same open-source agent harness (Pi v0.73.1, shell and file tools) for every model, with each model's default sampling (temperature 1.0 for Mellum2.1), a 114K-token context, and up to 16K tokens per turn.
- AIME is the mean of AIME 2025 and AIME 2026 (30 questions each).
- BFCL v4 is the macro-average of five subtasks: v1, v2, v3, web search, memory.
- Mellum2 Thinking was re-evaluated with this pipeline, so its numbers differ slightly from the Mellum2 Technical Report.
License
Released under the Apache 2.0 license.
- Downloads last month
- 198
Model tree for JetBrains/Mellum2.1-12B-A2.5B-Thinking
Base model
JetBrains/Mellum2-12B-A2.5B-BaseSpaces using JetBrains/Mellum2.1-12B-A2.5B-Thinking 3
Collection including JetBrains/Mellum2.1-12B-A2.5B-Thinking
Paper for JetBrains/Mellum2.1-12B-A2.5B-Thinking
Evaluation results
- pass@1 on LiveCodeBench v6self-reported82.000
- pass@1 on HumanEval+self-reported91.500
- pass@1 on MBPP+self-reported79.400
- exact match on AIME 25/26self-reported83.300
- exact match on GSM-Plusself-reported88.300
- resolved rate on SWE-bench Verifiedself-reported47.000
- resolved rate on Terminal-Bench 2.1self-reported17.400
- resolved rate on SWE-bench Proself-reported28.000