Instructions to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32") model = AutoModelForCausalLM.from_pretrained("blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32
- SGLang
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with Docker Model Runner:
docker model run hf.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32
Download serving/README.md from blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32: direct link, hf CLI and curl.
- Browser
- Download file 3.18 kB
-
https://huggingface.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32/resolve/main/serving/README.md
- Command line
-
hf download hf://blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32/serving/README.md
-
curl -L -o README.md https://huggingface.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32/resolve/main/serving/README.md
Serving compatibility launcher
serve_compat.py is project-owned serving glue, independent of the frozen
calibration sources. Its SHA-256 and the tested serving image/version appear in
../publication-manifest.json. The tested Ampere fork uses vLLM 0.29.0 and
Transformers 5.17.0, with local immutable image
sha256:68d4a7a5303358be89f9cc6e035fa4d719a5fac554d912e6b31d16998b35a268.
This is an identity receipt, not a public image download. Compatibility with
other installations has not been established.
That installation's MellumConfig inherits Qwen3MoeConfig, whose constructor
adds an unused outer scalar to the source's nested RoPE parameters. The
supported callable hf_overrides restores the exact checkpoint dictionary
before validation. Full-attention layers retain the source YaRN factor 16,
theta 500000, original context 8192, beta 32/1 and attention factor
1.2772588722239782; sliding-attention layers retain default RoPE with theta
500000. The tokenizer registry reads configuration independently, so the
launcher supplies a temporary tokenizer-only directory with original file
bytes and --tokenizer-mode hf. That directory remains alive until serving
ends. No checkpoint files or installed packages are changed. A CLI dictionary
override is applied too late and must not replace this callable launcher.
Reserve a GPU and enter the compatible serving environment, which must provide vLLM, Transformers, uvloop and the tested attention/quantization kernels. Download this repository at its immutable published revision as shown in the model card. This includes
serving/serve_compat.py.Run a config/tokenizer preflight from the downloaded repository root:
VLLM_MARLIN_INPUT_DTYPE=int8 python serving/serve_compat.py \ --model . --compat-config-only --quantization compressed-tensors \ --dtype bfloat16 --max-model-len 131072 --kv-cache-dtype autoExpect
status: passed, 28 decoder layers, the exact checkpoint RoPE dictionaries and source BOS/EOS IDs 0/28. This uses actual engine config and tokenizer registry construction; it does not load weights or prove model quality. Missing core tokenizer files, the wrong model/layer layout, or changed RoPE semantics must fail. The launcher owns tokenizer selection and overrides: omit--tokenizerand--hf-overrides.Use the model-card launch command for actual serving. Confirm model loading, selected attention/MoE kernels, tokenizer IDs, thinking control and the allocated KV capacity. Run semantic/schema/tool tests and the context lengths needed by the workload before drawing quality or performance conclusions. FP8 KV needs its own matched comparison; the artifact does not calibrate KV scales. Preserve failures rather than automatically restarting the process.
Actual preflight passed for both the pinned BF16 source and quantized configuration in the tested installation. The BF16 source also fully loaded with TP2 and completed a short generation. The derivative's completed quality/performance evidence belongs in the model card; these source-control and compatibility checks do not establish that evidence.