Instructions to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32") model = AutoModelForCausalLM.from_pretrained("blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32
- SGLang
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with Docker Model Runner:
docker model run hf.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32
Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32
This is an experimental community weight-only quantization of JetBrains' Mellum2.1 Thinking, frozen at revision 92ddae9fc7665e9f801d141d2e5a6b2caf2460c4. The checkpoint has 12,149,923,072 parameters in total and approximately 2.44 billion active parameters per token. It retains the upstream 28-layer architecture, 64 experts per layer, eight active experts, tokenizer, chat template and generation settings. JetBrains created the original model; this repository supplies its quantized derivative and validation information. It is released with the documented limitations, including failed Hindsight semantic admission and unreliable thinking-off compliance.
Quantization
The export uses compressed-tensors pack-quantized format with symmetric INT4 weights, group size 32, positive scales and no exported activation quantization (input_activations: null). All 5,376 expert projections and 112 attention projections are quantized. Routers, embeddings, LM head and normalization tensors retain BF16 precision. AWQ smoothing may change BF16 router and normalization values; embeddings and the LM head are checked byte-for-byte against the source.
Calibration combines AWQ smoothing (duo_scaling: both, 20 grid points) and an MSE weight observer. llmcompressor 0.14.0 substitutes its memoryless_mse implementation for the requested mse observer. Explicit mappings include the BF16 router when balancing expert input normalization, and omit the GQA-incompatible attention v_proj -> o_proj mapping. See quantization-recipe.yaml.
The sequential quantization callback computes weight quantization parameters for every targeted projection, including experts not activated by the calibration samples. Unrouted experts can therefore receive data-free MSE weight scales while skipping activation-dependent AWQ smoothing. Packed coverage of all 64 experts per layer does not establish representative activation coverage for every expert.
The 256 public calibration samples use seed 1234 and at most 2,048 tokens per sample: 128 UltraChat conversations, 64 Hermes function-calling examples and 64 Hermes agentic JSON examples. Calibration renders the upstream chat template with enable_thinking=False. No private production memories, customer prompts or runtime traces are included. Dataset pins are:
HuggingFaceH4/ultrachat_200k@8049631c405ae6576f93f445c6b8166f76f5505a,default/train_sft.NousResearch/hermes-function-calling-v1@dae3e1d28cfbcf4b915c04ea1e072030529b4bda,func_calling/trainandjson_mode_agentic/train.
The CUDA quantizer uses llmcompressor 0.14.0, compressed-tensors 0.19.0, Transformers 5.17.0, datasets 5.0.1, NumPy 2.5.3 and inherited Torch 2.13.0+cu130. The base image is ghcr.io/avesed/vllm-ampere-optimized@sha256:6744eac95c9d0ee426c5c60ebaf72437da551e8a1786e8d2270e5164fb17dbfc; the successful calibration ran in immutable local image sha256:e4f0dbe60e59050796a15b7b0c4646130a93acfd4e2431cde05923e401470b10. This local image ID records the run; it is not a claim that this derivative image is downloadable from a public registry.
Temporary supported PyTorch forward hooks discard only the unused eager attention probability tuple item during calibration; attention calculation and the primary output are unchanged. A tiny-model preflight checked bit-identical BF16 logits before, during and after the hook; this does not establish full-checkpoint quantization fidelity. The hooks are removed before export. No upstream packages were modified. Reproduction files, provenance and the completed run receipt describe the actual run, exact frozen source hashes, timestamps, bounded RAM/swap adjustments and sampled GPU usage. Those calibration measurements are not inference benchmarks. SHA-256 hashes of the published files appear in publication-manifest.json.
Serving
This artifact is W4A16, not a portable activation-calibrated W4A8 checkpoint. A compatible Ampere vLLM fork can choose INT8 computation for eligible Marlin projections using VLLM_MARLIN_INPUT_DTYPE=int8; BF16 tensors and other operations remain BF16. Do not add W4A8 activation metadata: it changes backend dispatch. Compatibility with other runtimes must be verified separately.
The tested serving environment is the Ampere fork with vLLM 0.29.0 and Transformers 5.17.0, local immutable image sha256:68d4a7a5303358be89f9cc6e035fa4d719a5fac554d912e6b31d16998b35a268. That image ID identifies the tested installation, not a public registry download. Broader compatibility has not been tested.
That installation requires serving/serve_compat.py: it passes a supported callable hf_overrides before model-config validation and restores the checkpoint's exact full/sliding RoPE dictionaries. It also supplies a temporary tokenizer-only directory containing byte-identical original tokenizer files, with --tokenizer-mode hf, for the server's lifetime. It changes neither checkpoint configuration/weights nor installed packages. A plain CLI --hf-overrides dictionary is applied too late for this installation and does not fix the problem. Actual engine/tokenizer preflight passed for source and quantized configurations; full BF16 source loading and a short generation also passed. Those source checks are not quantized-model quality evidence.
Use the upstream qwen3 reasoning parser and hermes tool parser. Run inside the compatible serving environment with an explicitly reserved GPU. The example below uses this Thinking checkpoint's thinking-on mode and the tested experimental INT8/FP8 configuration. It is not an approved Hindsight deployment. Download the immutable release locally, then launch:
hf download blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 \
--revision PUBLISHED_REVISION --local-dir mellum21
VLLM_MARLIN_INPUT_DTYPE=int8 python mellum21/serving/serve_compat.py \
--model mellum21 \
--served-model-name Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 \
--quantization compressed-tensors --dtype bfloat16 \
--kv-cache-dtype fp8_e4m3 --attention-backend FLASHINFER \
--no-disable-hybrid-kv-cache-manager \
--kv-cache-memory-bytes 10737418240 \
--kv-transfer-config '{"kv_connector":"OffloadingConnector","kv_role":"kv_both","kv_connector_extra_config":{"cpu_bytes_to_use":17179869184,"blocks_per_chunk":1,"self_describing_kv_events":true,"eviction_policy":"lru","offload_prompt_only":false}}' \
--max-model-len 131072 --max-num-seqs 8 --max-num-batched-tokens 8192 \
--enable-chunked-prefill --enable-prefix-caching --scheduling-policy priority \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser hermes \
--default-chat-template-kwargs '{"enable_thinking":true}'
Replace PUBLISHED_REVISION with the immutable Hub commit returned by publication. Reserve the GPU and enough host memory for the 16 GiB CPU KV cache. The context flag requests the architecture's limit, and eight scheduler slots do not establish eight-request admission at that length. The command requests eligible INT8 Marlin computation and FP8 KV with uncalibrated unit scales; verify selected kernels and keep runtime settings matched to the evidence below. Serving instructions include a BF16-KV config/tokenizer preflight that does not load weights. A client may set the thinking template control explicitly:
response = client.chat.completions.create(
model="Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32",
messages=[{"role": "user", "content": "Return a JSON object containing 2 + 2."}],
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
max_tokens=2048,
)
Direct tokenizer use accepts tokenizer.apply_chat_template(..., enable_thinking=False). This requests the template's nonthinking mode; it does not guarantee that generation contains no reasoning prose or literal reasoning tags. The source is a Thinking checkpoint, and JetBrains' reported benchmarks used thinking mode. The setting also does not establish native support for reasoning_effort=low or a thinking-token budget. The source checkpoint contains no MTP head. This quantization does not calibrate KV-cache scales; FP8 KV quality requires separate evidence.
Validation and limitations
Experimental release authorized with known limitations. Hindsight semantic admission has not passed.
The full 256-sample CUDA calibration completed, packing/export exited successfully, and the structural audit passed for all 5,376 expert and 112 attention projections, 365,395,968 finite positive scales, 28 BF16 routers and no MTP tensors. BF16 embeddings and LM head matched the source bytes; tokenizer vocabulary/special IDs, generation settings and chat template were preserved. This is structural evidence, separate from semantic quality. The complete model.safetensors SHA-256 is 81c89dc887629530cefcb1a8eb2362da4ca74df4988bcdf951a86efc0de84927.
The full quantized export loaded in the tested Ampere vLLM installation. Bounded profiling confirmed signed INT8 activation inputs in both MoE projections and eligible dense Marlin operations. These observations establish actual loading/dispatch, not quality or performance admission.
Short direct-model diagnostics
A 12-case synthetic conversation/structured-output/tool profile used the stock template requesting thinking-off mode and a 2,048-token output cap. These were direct model calls, with no Hindsight bank writes or BEAM run. The custom completion heuristic is a diagnostic with stronger requirements and known validator false positives; it is not a general accuracy benchmark.
| Arm | Custom heuristic passes | Observed limitations |
|---|---|---|
| Original BF16 source, BF16 KV | 5/12 | Coherent outputs, but fact-type, history and persisted-owner errors |
| INT4 weights, INT8 compute, BF16 KV | 3/12 | Coherent core facts, observation-update errors and one empty parsed final tool turn |
| INT4 weights, INT8 compute, FP8 E4M3 KV | 5/12 | Coherent core facts and final tool turns; observation-update errors remain |
The original BF16 source used TP2 and Triton experts; the quantized arms used TP1 and Marlin INT8 computation in the same vLLM implementation. Parallelism, kernels, weight precision and activation precision therefore changed together. This control does not isolate quantization loss or exclude an implementation-specific weakness. Manual source-faithfulness review found plan/date attribution problems in all arms. Valid streams, JSON schemas or tool progression do not establish factual fidelity. Two bounded repeats of the empty BF16-KV final turn returned valid tool calls; the original failure remains part of the evidence.
Matched long-context diagnostics
Identical synthetic fixtures exercised nominal 16K, 32K, 64K and 128K inputs in fresh and repeated-prefix phases. The largest actual prompt was 130,878 tokens with 128 output tokens reserved. Both arms used the same INT4 weights, INT8 computation, a 9 GiB GPU KV pool and 16 GiB CPU KV LRU. Some fresh cases reused 32 header tokens, so these were not strictly cold performance measurements. FP8 KV used uncalibrated unit scales.
| KV type | Secret-reference turns passed | Entity/date/abstention turns passed |
|---|---|---|
| BF16 | 25/32 | 8/16 |
| FP8 E4M3 | 29/32 | 8/16 |
The partitions are separate fixture outcomes, not aggregate model-accuracy scores. FP8 improved reference retrieval in this set but did not resolve factual quality. Both arms missed end-of-context references at 128K and hallucinated dates or distractor entities when the requested entity/date was absent. Present-reference cases could return the case identifier instead of the person. These failures limit long-context and abstention use; architecture support for 131,072 tokens is not a semantic-quality guarantee.
A separate BF16-activation diagnostic kept the same INT4 weights and BF16 KV and retested six failing turns. It passed 2/6: the 64K end-reference pair passed, while the 128K end-reference and 64K entity/date pairs still failed. This subset does not establish an alternative serving policy or show that disabling INT8 computation resolves the problem.
Thinking control and extraction sanity
Four public/synthetic prompts were tested in thinking-off and thinking-on modes, for eight calls on the final experimental INT8/FP8, 10 GiB configuration. Thinking-on returned correct finals with clean EOS/tool termination and separately parsed reasoning for all four: grounded lookup, plain arithmetic, strict JSON arithmetic and a named tool call. This vLLM 0.29 installation exposes reasoning through message.reasoning, not the legacy reasoning_content field; a null legacy field is not proof that reasoning was absent.
Thinking-off was clean for three prompts. The plain-arithmetic case emitted deliberative prose and a literal </think> before its correct final, and an identical repeat reproduced the leak. Adding an explicit no-reasoning instruction removed the marker but still produced calculation steps. The original BF16 source on the same vLLM, with BF16 KV and INT8/FAMP disabled, reproduced the leak twice. Quantization, INT8 computation and FP8 KV are therefore not necessary for this observed failure; native Transformers generation was not tested, so model behavior cannot be separated from that serving implementation. enable_thinking=False is a supported template request, not a demonstrated hard switch. This is a Thinking checkpoint; the small paired sanity test is not general accuracy evidence.
Two unchanged extraction/retain requests were also replayed with thinking-on, identical messages/schemas, temperature 0.1 and a 16,384-token output cap. Both passed JSON Schema validation and ended at EOS with reasoning separated from clean final content. One omitted important user-authored component groups and scheduled phase ranges; the other preserved imperative tasks and corrected some date attribution, but still turned scheduled milestones into completed events. Neither established reliable extraction fidelity or Hindsight admission. Raw inputs, reasoning and final outputs remain private and are not included here.
Native weight-fidelity control
Eight independently authored holdouts covered 5,342 teacher-forced next-token positions. The original source and this exact exported checkpoint used identical frozen token IDs, native Transformers, BF16 weights/activations, eager attention/experts, no KV cache and no TF32, INT8 or FP8 serving kernels. The export was dequantized through supported package APIs, with no custom unpacker or upstream patch. Both used identical all-GPU 14/14-layer placement; actual parameters/execution hooks had offloading disabled, and missing/unexpected/mismatched/loading-error lists were empty.
| Native pooled measure | Original BF16 source | Dequantized export |
|---|---|---|
| Mean negative log likelihood | 2.369141 | 2.337466 |
| Perplexity | 10.688207 | 10.354963 |
The quant/source pooled perplexity ratio was 0.968821, mean NLL delta was -0.031675, source/quant top-1 agreement was 86.447% and mean top-five overlap was 80.019%. Six holdouts had lower perplexity; JSON-schema and Python-source holdouts worsened by 5.166% and 2.199%, respectively. This small control supports bounded weight-artifact fidelity with no pooled perplexity regression. It does not establish improved general quality, generation behavior, long-context fidelity, live INT8/FP8 numeric equivalence or Hindsight suitability. Native generation was not part of the diagnostic. The completed comparison exited 0 without OOM.
Single-GPU performance and capacity
The full matrix used one RTX 3090, INT8 computation, FP8 E4M3 KV, a 10.5 GiB GPU KV pool, 16 GiB CPU KV LRU, eight scheduler slots and 8,192-token chunked prefill, with profiling disabled. There were two repetitions per context/concurrency pair. All 16 cold-prefill batches (40 requests) and 16 warmed-decode batches (40 measured requests) passed stream/output/cache/backend-count checks with no preemption/retraction. Cold requests reported zero cached tokens. Decode used independent one-token prefix warmups and exactly 1,024 forced output tokens rather than natural answer completion.
| Nominal input | Cold prefill C1 tok/s | Cold prefill C4 aggregate tok/s | Warm decode C1 tok/s | Warm decode C4 aggregate tok/s |
|---|---|---|---|---|
| 4,096 | 17,229 | 20,766 | 213.1 | 607.3 |
| 16,384 | 16,376 | 18,251 | 210.2 | 565.5 |
| 38,619 | 13,193 | 13,827 | 202.3 | 498.6 |
| 65,536 | 10,347 | 10,579 | 189.7 | 441.0 |
These are client measurements. Prefill divides summed prompt usage by launch-to-last-first-token time and includes queue/network/first-token work. Decode uses the union from earliest first token to latest last token, excluding first chunks, with staggered streams/queues counted. Prompts are slightly longer than nominal due to fresh identifiers and templates. The rates are not pure kernel timings or guarantees for natural thinking-enabled answers.
The final experimental candidate reduces the GPU KV pool to 10 GiB for additional memory margin. Only a representative 64K/C4 confirmation was repeated at that budget: 10,653 aggregate cold-input tok/s and 439.1 aggregate warm-output tok/s, with both admission audits passing. Its engine capacity was 1,070,880 KV tokens / 8.17 full-length request equivalents; that is architecture-dependent pool accounting, not eight-request 128K admission or semantic validation. Model allocation was 7.07 GiB and CUDA graphs 0.25 GiB. The limited 10 GiB confirmation sampled 22,639 MiB total device usage and 1,486 MiB NVML free; device usage includes a 1,198 MiB retrieval-service baseline and is not an instantaneous allocator maximum. The complete matrix belongs to 10.5 GiB, not 10 GiB, and neither is an exhaustively determined maximum. Higher throughput does not reverse the semantic failures.
Publication and Hindsight status
This experimental community release is explicitly authorized despite the documented limitations. The authorization permits distribution of the calibrated, audited Thinking-quant artifact; it is not a claim that the semantic quality gate passed. Structural checks, actual loading/dispatch, bounded native fidelity and the small thinking-on sanity set support evaluating this artifact for other workloads, with the limitations above retained.
Hindsight semantic admission remains unpassed. Thinking-off compliance and actual-retain fidelity remain unresolved; this model is not an approved Hindsight replacement. No Hindsight-service/BEAM evaluation has run, and that evaluation remains held for a later explicit start request. No private extraction inputs, reasoning/final responses, customer content or credentials are included. Native holdout and performance outcomes are scoped diagnostics, not broad model-accuracy scores or guarantees for untested use cases.
Structural export checks establish packed coverage, shapes, finite positive scales, BF16 tensor preservation and source template/generation configuration preservation. They do not establish semantic quality. The original JetBrains model's benchmark scores are not scores for this derivative. Do not infer quality or throughput at untested context lengths, concurrency levels, runtime versions, activation compute settings or KV dtypes. Calibration at 2,048 tokens does not validate the source architecture's 131,072-token context limit.
No Hindsight/BEAM result is implied by calibration, packing, export auditing or direct model tests. Only completed evaluations explicitly described above are evidence for this derivative.
License and attribution
The pinned upstream model card declares Apache 2.0. This derivative retains that license; see LICENSE and NOTICE. Changes consist of AWQ smoothing, INT4 weight packing and publication of the recipe, provenance and validation summary. Original model development and training are attributable to JetBrains. This community derivative is not an official JetBrains release.
- Downloads last month
- 5
Model tree for blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32
Base model
JetBrains/Mellum2-12B-A2.5B-Base