--- license: other base_model: JetBrains/Mellum2.1-12B-A2.5B-Thinking base_model_relation: quantized library_name: transformers pipeline_tag: text-generation tags: - fp8 - compressed-tensors - vllm - quantized quantized_by: liodon-ai --- # Mellum2.1-12B-A2.5B-Thinking — FP8 (dynamic) FP8 quantization of [JetBrains/Mellum2.1-12B-A2.5B-Thinking](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking), published by [Liodon AI](https://huggingface.co/liodon-ai). Quantized with [llm-compressor](https://github.com/vllm-project/llm-compressor) using the `FP8_DYNAMIC` scheme: weights are cast to FP8 (E4M3) per-channel ahead of time, activations are quantized to FP8 dynamically per-token at inference time. No calibration dataset is needed for this scheme, so the quantized weights are numerically just a direct cast of the original — no calibration-set bias to worry about. `lm_head` is left unquantized (standard practice — negligible size, disproportionate quality impact if quantized). Original size: 24.3 GB → Quantized: 12.6 GB. ## Quick Start **vLLM** ```bash vllm serve liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8 ``` **Text Generation Inference (TGI)** ```bash docker run --gpus all -p 8080:80 ghcr.io/huggingface/text-generation-inference \ --model-id liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8 ``` **SGLang** ```bash python -m sglang.launch_server --model-path liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8 ``` FP8 execution requires an NVIDIA GPU with compute capability ≥ 8.9 (Ada/Hopper/Blackwell — RTX 40-series, L4/L40S, H100/H200, B100/B200/GB10). On older GPUs, vLLM/TGI will dequantize to run, which loses the speed/memory benefit. ## Source - **Model**: [JetBrains/Mellum2.1-12B-A2.5B-Thinking](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking) - **License**: other ## Citation ```bibtex @misc{liodonai_mellum2_1_12b_a2_5b_thinking_fp8, title = {Mellum2.1-12B-A2.5B-Thinking — FP8}, author = {{Liodon AI}}, year = {2026}, howpublished = {\url{https://huggingface.co/liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8}}, note = {FP8 (dynamic) quantization of JetBrains/Mellum2.1-12B-A2.5B-Thinking} } ``` --- *Quantized by [Liodon AI](https://huggingface.co/liodon-ai)*