File size: 6,443 Bytes
e9a6277
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
---
license: other
license_name: lfm1.0
license_link: LICENSE
language:
- en
pipeline_tag: text-generation
tags:
- liquid
- edge
- lfm2
- lfm2.5
- moe
- mixture-of-experts
- onnx
- onnxruntime
base_model:
- LiquidAI/LFM2.5-8B-A1B
---

<div align="center">
  <img 
    src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" 
    alt="Liquid AI" 
    style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;"
  />
  <div style="display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;">
    <a href="https://playground.liquid.ai/"><strong>Try LFM</strong></a> β€’ 
    <a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> β€’ 
    <a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> β€’ 
    <a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a>
  </div>
</div>

# LFM2.5-8B-A1B-ONNX

ONNX export of [LFM2.5-8B-A1B](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B) for cross-platform inference.

LFM2.5-8B-A1B is a Mixture of Experts model with 8B total parameters and about 1B active parameters per token. It uses 32 experts with 4 experts activated per token, combining the efficiency of sparse models with the quality of larger dense models.

## Recommended Variants

| Precision | Size | Use Case |
|-----------|------|----------|
| Q4F16 | ~4.7GB | Recommended (Q4 MoE + FP16 dense) |
| FP16 | ~15.8GB | Higher quality |
| Q4 | ~5.2GB | Smallest size |
| Q8 | ~30.4GB | Highest-fidelity quantized variant |

Note: This model is too large for WebGPU browser inference.

## Validation

This export was validated against the local PyTorch reference for `LiquidAI/LFM2.5-8B-A1B`.

- FP32 padded-batch parity passed for both left and right padding, with cosine similarity `1.0000` and top-5 overlap `5/5` at the last valid token for each row.
- Q4 decoder and coherence checks passed the repository thresholds. Average coherence similarity: `0.7144`.
- Q4F16 was runtime-validated on `CPUExecutionProvider` and matched the same decoder/coherence thresholds as Q4. Average coherence similarity: `0.7145`.
- Q8 decoder and coherence checks passed, and stayed very close to the PyTorch reference. Average coherence similarity: `0.9975`.

## Model Files

```
onnx/
β”œβ”€β”€ model.onnx              # FP32 model graph
β”œβ”€β”€ model.onnx_data*        # FP32 weights
β”œβ”€β”€ model_fp16.onnx         # FP16 model graph
β”œβ”€β”€ model_fp16.onnx_data*   # FP16 weights
β”œβ”€β”€ model_q4.onnx           # Q4 model graph
β”œβ”€β”€ model_q4.onnx_data*     # Q4 weights
β”œβ”€β”€ model_q4f16.onnx        # Q4 MoE experts + FP16 dense (recommended)
β”œβ”€β”€ model_q4f16.onnx_data*  # Q4F16 weights
β”œβ”€β”€ model_q8.onnx           # Q8 model graph
└── model_q8.onnx_data*     # Q8 weights

* Large models split weights across multiple files:
  model.onnx_data, model.onnx_data_1, model.onnx_data_2, etc.
  All data files must be in the same directory as the .onnx file.
```

## Python

### Installation

```bash
pip install onnxruntime transformers numpy huggingface_hub
# or with GPU support:
pip install onnxruntime-gpu transformers numpy huggingface_hub
```

### Inference

```python
from huggingface_hub import snapshot_download
from transformers import AutoConfig, AutoTokenizer
import numpy as np
import onnxruntime

# 1. Load config, tokenizer, and model
model_id = "LiquidAI/LFM2.5-8B-A1B-ONNX"
config = AutoConfig.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
eos_token_id = config.eos_token_id

filename = "model_q4f16.onnx"  # Options: model.onnx, model_fp16.onnx, model_q4.onnx, model_q4f16.onnx, model_q8.onnx
model_path = snapshot_download(repo_id=model_id, allow_patterns=f"onnx/{filename}*")
session = onnxruntime.InferenceSession(f"{model_path}/onnx/{filename}")
input_names = {inp.name for inp in session.get_inputs()}

# 2. Prepare inputs
prompt = "What is C. elegans?"
messages = [{"role": "user", "content": prompt}]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="np",
)
input_ids = inputs["input_ids"]
attention_mask = inputs["attention_mask"]
batch_size = input_ids.shape[0]

past_cache_values = {}
for inp in session.get_inputs():
    name = inp.name
    shape = inp.shape
    dtype = np.float32 if inp.type == "tensor(float)" else np.float16
    if name.startswith("past_key_values"):
        past_cache_values[name] = np.zeros([batch_size, shape[1], 0, shape[3]], dtype=dtype)
    elif name.startswith("past_conv"):
        past_cache_values[name] = np.zeros([batch_size, shape[1], shape[2]], dtype=dtype)

position_ids = np.arange(input_ids.shape[1], dtype=np.int64).reshape(1, -1)

# 3. Generation loop
max_new_tokens = 256
generated_tokens = np.array([[]], dtype=np.int64)
cur_len = input_ids.shape[1]
for i in range(max_new_tokens):
    if i == 0:
        ids = input_ids
        pos = position_ids
    else:
        ids = generated_tokens[:, -1:]
        pos = np.array([[cur_len - 1]], dtype=np.int64)

    feed = {
        "input_ids": ids,
        "attention_mask": attention_mask,
        **past_cache_values,
    }
    if "position_ids" in input_names:
        feed["position_ids"] = pos

    outputs = session.run(None, feed)
    logits = outputs[0]
    next_token = logits[:, -1].argmax(-1, keepdims=True)

    generated_tokens = (
        next_token if generated_tokens.shape[1] == 0
        else np.concatenate([generated_tokens, next_token], axis=-1)
    )
    attention_mask = np.concatenate(
        [attention_mask, np.ones_like(next_token, dtype=np.int64)],
        axis=-1,
    )

    output_names = [out.name for out in session.get_outputs()]
    cache_outputs = {
        name: value
        for name, value in zip(output_names[1:], outputs[1:])
    }
    for key in past_cache_values:
        present_key = key.replace("past_key_values", "present").replace("past_conv", "present_conv")
        past_cache_values[key] = cache_outputs[present_key]

    cur_len += 1
    if np.isin(next_token, eos_token_id).any():
        break

    print(tokenizer.decode(next_token[0]), end="", flush=True)
print()

# 4. Output result
print(tokenizer.batch_decode(generated_tokens, skip_special_tokens=True)[0])
```

## License

This model is released under the [LFM 1.0 License](LICENSE).