PY-AI-Dev commited on
Commit
fb00258
·
verified ·
1 Parent(s): e6539c8

Add FP8 (dynamic) quantization for Mellum2.1-12B-A2.5B-Thinking

Browse files
README.md ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ base_model: JetBrains/Mellum2.1-12B-A2.5B-Thinking
4
+ base_model_relation: quantized
5
+ library_name: transformers
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - fp8
9
+ - compressed-tensors
10
+ - vllm
11
+ - quantized
12
+ quantized_by: liodon-ai
13
+ ---
14
+
15
+ # Mellum2.1-12B-A2.5B-Thinking — FP8 (dynamic)
16
+
17
+ FP8 quantization of [JetBrains/Mellum2.1-12B-A2.5B-Thinking](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking), published by [Liodon AI](https://huggingface.co/liodon-ai).
18
+
19
+ Quantized with [llm-compressor](https://github.com/vllm-project/llm-compressor) using the
20
+ `FP8_DYNAMIC` scheme: weights are cast to FP8 (E4M3) per-channel ahead of time, activations are
21
+ quantized to FP8 dynamically per-token at inference time. No calibration dataset is needed for this
22
+ scheme, so the quantized weights are numerically just a direct cast of the original — no calibration-set
23
+ bias to worry about. `lm_head` is left unquantized (standard practice — negligible size, disproportionate
24
+ quality impact if quantized).
25
+
26
+ Original size: 24.3 GB → Quantized: 12.6 GB.
27
+
28
+ ## Quick Start
29
+
30
+ **vLLM**
31
+ ```bash
32
+ vllm serve liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8
33
+ ```
34
+
35
+ **Text Generation Inference (TGI)**
36
+ ```bash
37
+ docker run --gpus all -p 8080:80 ghcr.io/huggingface/text-generation-inference \
38
+ --model-id liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8
39
+ ```
40
+
41
+ **SGLang**
42
+ ```bash
43
+ python -m sglang.launch_server --model-path liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8
44
+ ```
45
+
46
+ FP8 execution requires an NVIDIA GPU with compute capability ≥ 8.9 (Ada/Hopper/Blackwell — RTX 40-series,
47
+ L4/L40S, H100/H200, B100/B200/GB10). On older GPUs, vLLM/TGI will dequantize to run, which loses the
48
+ speed/memory benefit.
49
+
50
+ ## Source
51
+
52
+ - **Model**: [JetBrains/Mellum2.1-12B-A2.5B-Thinking](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking)
53
+ - **License**: other
54
+
55
+ ## Citation
56
+
57
+ ```bibtex
58
+ @misc{liodonai_mellum2_1_12b_a2_5b_thinking_fp8,
59
+ title = {Mellum2.1-12B-A2.5B-Thinking — FP8},
60
+ author = {{Liodon AI}},
61
+ year = {2026},
62
+ howpublished = {\url{https://huggingface.co/liodon-ai/Mellum2.1-12B-A2.5B-Thinking-FP8}},
63
+ note = {FP8 (dynamic) quantization of JetBrains/Mellum2.1-12B-A2.5B-Thinking}
64
+ }
65
+ ```
66
+
67
+ ---
68
+ *Quantized by [Liodon AI](https://huggingface.co/liodon-ai)*
chat_template.jinja ADDED
@@ -0,0 +1,103 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- macro normalize_content(content) -%}
2
+ {%- if content is string -%}
3
+ {{- content -}}
4
+ {%- elif content is iterable and content is not mapping -%}
5
+ {%- set ns_c = namespace(text='') -%}
6
+ {%- for part in content -%}
7
+ {%- if part is mapping -%}
8
+ {%- if part.type == 'text' -%}
9
+ {%- set ns_c.text = ns_c.text + (part.text or '') -%}
10
+ {%- elif part.type == 'tool-result' and part.output is mapping and part.output.type == 'text' -%}
11
+ {%- set ns_c.text = ns_c.text + (part.output.value or '') -%}
12
+ {%- endif -%}
13
+ {%- endif -%}
14
+ {%- endfor -%}
15
+ {{- ns_c.text -}}
16
+ {%- endif -%}
17
+ {%- endmacro -%}
18
+ {%- if tools %}
19
+ {{- '<|im_start|>system\n' }}
20
+ {%- if messages[0].role == 'system' %}
21
+ {{- normalize_content(messages[0].content) + '\n\n' }}
22
+ {%- endif %}
23
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
24
+ {%- for tool in tools %}
25
+ {{- "\n" }}
26
+ {{- tool | tojson }}
27
+ {%- endfor %}
28
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
29
+ {%- else %}
30
+ {%- if messages[0].role == 'system' %}
31
+ {{- '<|im_start|>system\n' + normalize_content(messages[0].content) + '<|im_end|>\n' }}
32
+ {%- endif %}
33
+ {%- endif %}
34
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
35
+ {%- for message in messages[::-1] %}
36
+ {%- set index = (messages|length - 1) - loop.index0 %}
37
+ {%- set msg_content = normalize_content(message.content) %}
38
+ {%- if ns.multi_step_tool and message.role == "user" and not(msg_content.startswith('<tool_response>') and msg_content.endswith('</tool_response>')) %}
39
+ {%- set ns.multi_step_tool = false %}
40
+ {%- set ns.last_query_index = index %}
41
+ {%- endif %}
42
+ {%- endfor %}
43
+ {%- for message in messages %}
44
+ {%- set content = normalize_content(message.content) %}
45
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
47
+ {%- elif message.role == "assistant" %}
48
+ {%- set reasoning_content = '' %}
49
+ {%- if message.reasoning_content is string %}
50
+ {%- set reasoning_content = message.reasoning_content %}
51
+ {%- else %}
52
+ {%- if '</think>' in content %}
53
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
54
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
55
+ {%- endif %}
56
+ {%- endif %}
57
+ {%- if loop.index0 > ns.last_query_index %}
58
+ {%- if reasoning_content %}
59
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
60
+ {%- else %}
61
+ {{- '<|im_start|>' + message.role + '\n' + content }}
62
+ {%- endif %}
63
+ {%- else %}
64
+ {{- '<|im_start|>' + message.role + '\n' + content }}
65
+ {%- endif %}
66
+ {%- if message.tool_calls %}
67
+ {%- for tool_call in message.tool_calls %}
68
+ {%- if (loop.first and content) or (not loop.first) %}
69
+ {{- '\n' }}
70
+ {%- endif %}
71
+ {%- if tool_call.function %}
72
+ {%- set tool_call = tool_call.function %}
73
+ {%- endif %}
74
+ {{- '<tool_call>\n{"name": "' }}
75
+ {{- tool_call.name }}
76
+ {{- '", "arguments": ' }}
77
+ {%- if tool_call.arguments is string %}
78
+ {{- tool_call.arguments }}
79
+ {%- else %}
80
+ {{- tool_call.arguments | tojson }}
81
+ {%- endif %}
82
+ {{- '}\n</tool_call>' }}
83
+ {%- endfor %}
84
+ {%- endif %}
85
+ {{- '<|im_end|>\n' }}
86
+ {%- elif message.role == "tool" %}
87
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
88
+ {{- '<|im_start|>user' }}
89
+ {%- endif %}
90
+ {{- '\n<tool_response>\n' }}
91
+ {{- content }}
92
+ {{- '\n</tool_response>' }}
93
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
94
+ {{- '<|im_end|>\n' }}
95
+ {%- endif %}
96
+ {%- endif %}
97
+ {%- endfor %}
98
+ {%- if add_generation_prompt %}
99
+ {{- '<|im_start|>assistant\n' }}
100
+ {%- if enable_thinking is defined and enable_thinking is false %}
101
+ {{- '<think>\n\n</think>\n\n' }}
102
+ {%- endif %}
103
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,188 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "MellumForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 0,
8
+ "dtype": "bfloat16",
9
+ "eos_token_id": 28,
10
+ "head_dim": 128,
11
+ "hidden_act": "silu",
12
+ "hidden_size": 2304,
13
+ "initializer_range": 0.02,
14
+ "intermediate_size": 7168,
15
+ "layer_types": [
16
+ "sliding_attention",
17
+ "sliding_attention",
18
+ "sliding_attention",
19
+ "full_attention",
20
+ "sliding_attention",
21
+ "sliding_attention",
22
+ "sliding_attention",
23
+ "full_attention",
24
+ "sliding_attention",
25
+ "sliding_attention",
26
+ "sliding_attention",
27
+ "full_attention",
28
+ "sliding_attention",
29
+ "sliding_attention",
30
+ "sliding_attention",
31
+ "full_attention",
32
+ "sliding_attention",
33
+ "sliding_attention",
34
+ "sliding_attention",
35
+ "full_attention",
36
+ "sliding_attention",
37
+ "sliding_attention",
38
+ "sliding_attention",
39
+ "full_attention",
40
+ "sliding_attention",
41
+ "sliding_attention",
42
+ "sliding_attention",
43
+ "full_attention"
44
+ ],
45
+ "max_position_embeddings": 131072,
46
+ "max_window_layers": 0,
47
+ "mlp_layer_types": [
48
+ "sparse",
49
+ "sparse",
50
+ "sparse",
51
+ "sparse",
52
+ "sparse",
53
+ "sparse",
54
+ "sparse",
55
+ "sparse",
56
+ "sparse",
57
+ "sparse",
58
+ "sparse",
59
+ "sparse",
60
+ "sparse",
61
+ "sparse",
62
+ "sparse",
63
+ "sparse",
64
+ "sparse",
65
+ "sparse",
66
+ "sparse",
67
+ "sparse",
68
+ "sparse",
69
+ "sparse",
70
+ "sparse",
71
+ "sparse",
72
+ "sparse",
73
+ "sparse",
74
+ "sparse",
75
+ "sparse"
76
+ ],
77
+ "model_type": "mellum",
78
+ "moe_intermediate_size": 896,
79
+ "norm_topk_prob": true,
80
+ "num_attention_heads": 32,
81
+ "num_experts_per_tok": 8,
82
+ "num_hidden_layers": 28,
83
+ "num_key_value_heads": 4,
84
+ "num_local_experts": 64,
85
+ "output_router_logits": false,
86
+ "pad_token_id": null,
87
+ "quantization_config": {
88
+ "config_groups": {
89
+ "group_0": {
90
+ "format": "float-quantized",
91
+ "input_activations": {
92
+ "actorder": null,
93
+ "block_structure": null,
94
+ "dynamic": true,
95
+ "group_size": null,
96
+ "num_bits": 8,
97
+ "observer": null,
98
+ "observer_kwargs": {},
99
+ "scale_dtype": null,
100
+ "strategy": "token",
101
+ "symmetric": true,
102
+ "type": "float",
103
+ "zp_dtype": null
104
+ },
105
+ "output_activations": null,
106
+ "targets": [
107
+ "Linear"
108
+ ],
109
+ "weights": {
110
+ "actorder": null,
111
+ "block_structure": null,
112
+ "dynamic": false,
113
+ "group_size": null,
114
+ "num_bits": 8,
115
+ "observer": "memoryless_minmax",
116
+ "observer_kwargs": {},
117
+ "scale_dtype": null,
118
+ "strategy": "channel",
119
+ "symmetric": true,
120
+ "type": "float",
121
+ "zp_dtype": null
122
+ }
123
+ }
124
+ },
125
+ "format": "float-quantized",
126
+ "global_compression_ratio": null,
127
+ "ignore": [
128
+ "model.layers.0.mlp.gate",
129
+ "model.layers.1.mlp.gate",
130
+ "model.layers.2.mlp.gate",
131
+ "model.layers.3.mlp.gate",
132
+ "model.layers.4.mlp.gate",
133
+ "model.layers.5.mlp.gate",
134
+ "model.layers.6.mlp.gate",
135
+ "model.layers.7.mlp.gate",
136
+ "model.layers.8.mlp.gate",
137
+ "model.layers.9.mlp.gate",
138
+ "model.layers.10.mlp.gate",
139
+ "model.layers.11.mlp.gate",
140
+ "model.layers.12.mlp.gate",
141
+ "model.layers.13.mlp.gate",
142
+ "model.layers.14.mlp.gate",
143
+ "model.layers.15.mlp.gate",
144
+ "model.layers.16.mlp.gate",
145
+ "model.layers.17.mlp.gate",
146
+ "model.layers.18.mlp.gate",
147
+ "model.layers.19.mlp.gate",
148
+ "model.layers.20.mlp.gate",
149
+ "model.layers.21.mlp.gate",
150
+ "model.layers.22.mlp.gate",
151
+ "model.layers.23.mlp.gate",
152
+ "model.layers.24.mlp.gate",
153
+ "model.layers.25.mlp.gate",
154
+ "model.layers.26.mlp.gate",
155
+ "model.layers.27.mlp.gate",
156
+ "lm_head"
157
+ ],
158
+ "kv_cache_scheme": null,
159
+ "quant_method": "compressed-tensors",
160
+ "quantization_status": "compressed",
161
+ "sparsity_config": {},
162
+ "transform_config": {},
163
+ "version": "0.18.0"
164
+ },
165
+ "rms_norm_eps": 1e-06,
166
+ "rope_parameters": {
167
+ "full_attention": {
168
+ "attention_factor": 1.2772588722239782,
169
+ "beta_fast": 32.0,
170
+ "beta_slow": 1.0,
171
+ "factor": 16.0,
172
+ "original_max_position_embeddings": 8192,
173
+ "rope_theta": 500000.0,
174
+ "rope_type": "yarn"
175
+ },
176
+ "sliding_attention": {
177
+ "rope_theta": 500000.0,
178
+ "rope_type": "default"
179
+ }
180
+ },
181
+ "router_aux_loss_coef": 0.001,
182
+ "sliding_window": 1024,
183
+ "tie_word_embeddings": false,
184
+ "transformers_version": "5.14.1",
185
+ "use_cache": true,
186
+ "use_sliding_window": true,
187
+ "vocab_size": 98304
188
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 0,
4
+ "eos_token_id": 28,
5
+ "transformers_version": "5.14.1"
6
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:097077a620a34c291af51acbebe48e8af0422e156d1f4b6bdd1f624d640309e3
3
+ size 12623670464
recipe.yaml ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ default_stage:
2
+ default_modifiers:
3
+ QuantizationModifier:
4
+ targets: [Linear]
5
+ ignore: [lm_head]
6
+ scheme: FP8_DYNAMIC
7
+ bypass_divisibility_checks: false
8
+ requires_calibration_data: false
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "bos_token": "<|endoftext|>",
4
+ "clean_up_tokenization_spaces": false,
5
+ "eos_token": "<|im_end|>",
6
+ "is_local": false,
7
+ "local_files_only": false,
8
+ "model_max_length": 131072,
9
+ "pad_token": "<|endoftext|>",
10
+ "tokenizer_class": "TokenizersBackend",
11
+ "unk_token": "<|endoftext|>"
12
+ }