How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8")
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8")
model = AutoModelForCausalLM.from_pretrained("liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8", device_map="auto")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

Mellum2.1-12B-A2.5B-Thinking — FP8 (dynamic)

FP8 quantization of JetBrains/Mellum2.1-12B-A2.5B-Thinking, published by Liodon AI.

Quantized with llm-compressor using the FP8_DYNAMIC scheme: weights are cast to FP8 (E4M3) per-channel ahead of time, activations are quantized to FP8 dynamically per-token at inference time. No calibration dataset is needed for this scheme, so the quantized weights are numerically just a direct cast of the original — no calibration-set bias to worry about. lm_head is left unquantized (standard practice — negligible size, disproportionate quality impact if quantized).

Original size: 24.3 GB → Quantized: 12.6 GB.

Quick Start

vLLM

vllm serve liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8

Text Generation Inference (TGI)

docker run --gpus all -p 8080:80 ghcr.io/huggingface/text-generation-inference \
    --model-id liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8

SGLang

python -m sglang.launch_server --model-path liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8

FP8 execution requires an NVIDIA GPU with compute capability ≥ 8.9 (Ada/Hopper/Blackwell — RTX 40-series, L4/L40S, H100/H200, B100/B200/GB10). On older GPUs, vLLM/TGI will dequantize to run, which loses the speed/memory benefit.

Source

Citation

@misc{liodonai_mellum2_1_12b_a2_5b_thinking_fp8,
  title        = {Mellum2.1-12B-A2.5B-Thinking — FP8},
  author       = {{Liodon AI}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8}},
  note         = {FP8 (dynamic) quantization of JetBrains/Mellum2.1-12B-A2.5B-Thinking}
}

Quantized by Liodon AI

Downloads last month
2
Safetensors
Model size
12B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8