Text Generation
Transformers
Safetensors
English
mellum
mixture-of-experts
compressed-tensors
awq
int4
w4a16
experimental
conversational
Instructions to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32") model = AutoModelForCausalLM.from_pretrained("blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32
- SGLang
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32 with Docker Model Runner:
docker model run hf.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32
Download publication-manifest.json from blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32: direct link, hf CLI and curl.
- Browser
- Download file 4.06 kB
-
https://huggingface.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32/resolve/main/publication-manifest.json
- Command line
-
hf download hf://blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32/publication-manifest.json
-
curl -L -o publication-manifest.json https://huggingface.co/blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32/resolve/main/publication-manifest.json
4.06 kB
| { | |
| "repo_id": "blake-lucas/Mellum2.1-12B-A2.5B-Thinking-AWQ-W4A16-G32", | |
| "source_model": "JetBrains/Mellum2.1-12B-A2.5B-Thinking", | |
| "source_revision": "92ddae9fc7665e9f801d141d2e5a6b2caf2460c4", | |
| "public_quality_summary_sha256": "47d3163d413ba3b792484b3e459334cd90b7e4249ab746ad71bd5c134531f904", | |
| "serving_compatibility": { | |
| "source_sha256": { | |
| "serve_compat.py": "4a5dc869255d884e95e06601a75d0122d5d6490729d77260edede45b612765e9" | |
| }, | |
| "tested_image_id": "sha256:68d4a7a5303358be89f9cc6e035fa4d719a5fac554d912e6b31d16998b35a268", | |
| "vllm": "0.29.0", | |
| "transformers": "5.17.0" | |
| }, | |
| "files": [ | |
| { | |
| "path": "LICENSE", | |
| "bytes": 11358, | |
| "sha256": "cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30" | |
| }, | |
| { | |
| "path": "NOTICE", | |
| "bytes": 756, | |
| "sha256": "d625f668cf02087a8b468882c4b6cb629470bc134400e9b41a7f9829e86a1753" | |
| }, | |
| { | |
| "path": "README.md", | |
| "bytes": 20738, | |
| "sha256": "47684e6fb117d8472842fe7c791ea57ebfc169fd192035da164a7e38244dd703" | |
| }, | |
| { | |
| "path": "calibration-run-receipt.json", | |
| "bytes": 3126, | |
| "sha256": "ef79f7844a58ca86edea302f2b70b57b93112613d645893c9d49204a435dbfa4" | |
| }, | |
| { | |
| "path": "chat_template.jinja", | |
| "bytes": 4849, | |
| "sha256": "4593a34d52f3364ac13ce57f4bea5688924e8942da6df8201e411db17e729f48" | |
| }, | |
| { | |
| "path": "config.json", | |
| "bytes": 4198, | |
| "sha256": "478a461e755e7ca4a4fd5d36c62b73b397e45552bff7eec4fa74f9b51cb35228" | |
| }, | |
| { | |
| "path": "export-audit.json", | |
| "bytes": 593, | |
| "sha256": "71503d5c21e8d9533a4763595518895f532aa3ccde8c9d664f41a6c9ad3b16ce" | |
| }, | |
| { | |
| "path": "generation_config.json", | |
| "bytes": 111, | |
| "sha256": "a9dc104ca398a2ef376d3f2fbb03037d99950ddb2db0cd041b98398cf1fc8d7c" | |
| }, | |
| { | |
| "path": "model.safetensors", | |
| "bytes": 7493755456, | |
| "sha256": "81c89dc887629530cefcb1a8eb2362da4ca74df4988bcdf951a86efc0de84927" | |
| }, | |
| { | |
| "path": "quantization-provenance.json", | |
| "bytes": 98333, | |
| "sha256": "64c36c8cffa54719b5f90ce5d3935a852687ea22e89c3a00aca6ac74d6c187a3" | |
| }, | |
| { | |
| "path": "quantization-recipe.yaml", | |
| "bytes": 1506, | |
| "sha256": "a73469b793eb4af8eacaa3241cecb6b699d3ccc88374e575458cf35500f6ab3a" | |
| }, | |
| { | |
| "path": "reproduction/Dockerfile.quantize", | |
| "bytes": 789, | |
| "sha256": "a6f281f8c21f025bb6e825e4efed7ce982af03e573e5a9d1214918bc6a44ec13" | |
| }, | |
| { | |
| "path": "reproduction/README.md", | |
| "bytes": 3951, | |
| "sha256": "d3bbf4448685284068322e6a1ca2636ce7047478d727772b312ef28d2006906c" | |
| }, | |
| { | |
| "path": "reproduction/audit_export.py", | |
| "bytes": 9610, | |
| "sha256": "9c2476d66193a505f43f8982a07c053fdabfe01205af051f6d6f7c37085ac199" | |
| }, | |
| { | |
| "path": "reproduction/check_quantizer.py", | |
| "bytes": 7075, | |
| "sha256": "c8bbb692f1e82e4193df78710df3f7b6d617793d66c8dfa8f0ab4836d05572c2" | |
| }, | |
| { | |
| "path": "reproduction/quantize.py", | |
| "bytes": 11740, | |
| "sha256": "890aef9cf2ceb5c50d663a0014a51326ff960b2098fc30f10efdf2d130d0fc33" | |
| }, | |
| { | |
| "path": "reproduction/recipe.yaml", | |
| "bytes": 1506, | |
| "sha256": "a73469b793eb4af8eacaa3241cecb6b699d3ccc88374e575458cf35500f6ab3a" | |
| }, | |
| { | |
| "path": "reproduction/requirements.quantize.txt", | |
| "bytes": 183, | |
| "sha256": "3dddfbf7bf5a7cdd012a56fd80f47ae895db1468f654b196f7ca10dab2f75617" | |
| }, | |
| { | |
| "path": "serving/README.md", | |
| "bytes": 3182, | |
| "sha256": "6ce2f9d75f3f8012bbb86d27eb3d9df5a31370fb41f3b3c16769a36b8bc1e6fb" | |
| }, | |
| { | |
| "path": "serving/serve_compat.py", | |
| "bytes": 5515, | |
| "sha256": "4a5dc869255d884e95e06601a75d0122d5d6490729d77260edede45b612765e9" | |
| }, | |
| { | |
| "path": "tokenizer.json", | |
| "bytes": 7091587, | |
| "sha256": "58548a346eb073e5132bf7d8ad17dc6971bca36ade378ca4d2bfbc49bf60da2a" | |
| }, | |
| { | |
| "path": "tokenizer_config.json", | |
| "bytes": 316, | |
| "sha256": "53bf0454ac01bbf2590b15ba8a01e4edcf2100492014d001b549929d746292cf" | |
| } | |
| ] | |
| } | |