FastFlowLM commited on
Commit
7d0657b
·
verified ·
1 Parent(s): 67f9fad

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +35 -15
README.md CHANGED
@@ -201,21 +201,41 @@ LFM2.5-1.2B-Thinking offers extremely fast inference speed on CPUs with a low me
201
 
202
  In addition, we are partnering with AMD, Qualcomm, Nexa AI, and FastFlowLM to bring the LFM2.5 family to NPUs. These optimized models are available through our partners, enabling highly efficient on-device inference.
203
 
204
- We report the following numbers with 1K prefill and 100 decode tokens:
205
-
206
- | Device | Inference | Framework | Model | Prefill (tok/s) | Decode (tok/s) | Memory |
207
- | ---------------------------------------------------- | --------- | ---------------- | -------------------- | --------------- | -------------- | ------ |
208
- | AMD Ryzen AI 395+ | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 1487 | 60 | 1700MB |
209
- | AMD Ryzen AI 9 HX 370 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 1487 | 57 | 1700MB |
210
- | AMD Ryzen AI 7 HX 350 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 1431 | 63 | 1700MB |
211
- | AMD Ryzen AI 5 HX 340 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking | 1431 | 63 | 1700MB |
212
- | AMD Ryzen AI 9 HX 370 | GPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking | 2975 | 116 | 856MB |
213
- | Qualcomm Snapdragon® X Elite | NPU | NexaML | LFM2.5-1.2B-Thinking | 2591 | 63 | 0.9GB |
214
- | Qualcomm Snapdragon® Gen4 (ROG Phone9 Pro) | NPU | NexaML | LFM2.5-1.2B-Thinking | 4391 | 82 | 0.9GB |
215
- | Qualcomm Dragonwing IQ9 (IQ-9075) (IoT) | NPU | NexaML | LFM2.5-1.2B-Thinking | 2143 | 53 | 0.9 GB |
216
- | Qualcomm Snapdragon® Gen4 (Samsung Galaxy S25 Ultra) | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking | 335 | 70 | 719MB |
217
-
218
- LFM2.5-1.2B-Thinking works well with long contexts. On AMD NPU, it delivers **~59 tok/s at 4K** context and **~52 tok/s at 16K** during decoding, and reaches a **prefill speed of ~2,226 tok/s with a 4K-token prompt**. Check the detailed benchmark results [here](https://fastflowlm.com/docs/benchmarks/lfm2_results/).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
219
 
220
  These capabilities unlock new deployment scenarios across various devices, including vehicles, mobile devices, laptops, IoT devices, and embedded systems.
221
 
 
201
 
202
  In addition, we are partnering with AMD, Qualcomm, Nexa AI, and FastFlowLM to bring the LFM2.5 family to NPUs. These optimized models are available through our partners, enabling highly efficient on-device inference.
203
 
204
+ #### **Prefill Performance**
205
+
206
+ We report prefill throughput evaluated over a range of prompt lengths.
207
+
208
+ | Platform / Device | Inference | Framework | Model | 1K Prefill (tok/s) | 4K Prefill (tok/s) | 16K Prefill (tok/s) | Memory |
209
+ |----------------------------------------------------|-----------|------------------|-------------------|-------------------:|-------------------:|--------------------:| -------:|
210
+ | AMD Ryzen AI 395+ | NPU | FastFlowLM | LFM2.5-1.2B-Thinking| 1,487 | 2,226 | 1,670 | 1.7 GB |
211
+ | AMD Ryzen AI 9 HX 370 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking| 1,487 | 2,226 |1,670 | 1.7 GB |
212
+ | AMD Ryzen AI 7 HX 350 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking| 1,431 | 2,032 |1,519 | 1.7 GB |
213
+ | AMD Ryzen™ AI 5 HX 340 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking| 1,431 | 2,032 |1,519 | 1.7 GB |
214
+ | AMD Ryzen™ AI 9 HX 370 | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking| 2,975 | N/A | N/A | 856 MB |
215
+ | Qualcomm Snapdragon® X Elite | NPU | NexaML | LFM2.5-1.2B-Thinking| 2,591 | N/A | N/A | 0.9 GB |
216
+ | Qualcomm Snapdragon® Gen4 (ROG Phone 9 Pro) | NPU | NexaML | LFM2.5-1.2B-Thinking| 4,391 | N/A | N/A | 0.9 GB |
217
+ | Qualcomm Dragonwing IQ9 (IQ-9075, IoT) | NPU | NexaML | LFM2.5-1.2B-Thinking| 2,143 | N/A | N/A | 0.9 GB |
218
+ | Qualcomm Snapdragon® Gen4 (Galaxy S25 Ultra) | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking| 335 | N/A | N/A | 719 MB |
219
+
220
+ #### **Decode Performance**
221
+
222
+ The reported results correspond to decoding 100 tokens at different context lengths.
223
+
224
+ | Platform / Device | Inference | Framework | Model | Decode @1K (tok/s) | Decode @4K (tok/s) | Decode @16K (tok/s) | Memory |
225
+ |----------------------------------------------------|-----------|------------------|-------------------|-------------------:|-------------------:|--------------------:| -------:|
226
+ | AMD Ryzen™ AI 395+ | NPU | FastFlowLM | LFM2.5-1.2B-Thinking| 60 | 54 | 49 | 1.7 GB |
227
+ | AMD Ryzen™ AI 9 HX 370 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking| 57 | 54 |49 | 1.7 GB |
228
+ | AMD Ryzen™ AI 7 HX 350 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking| 63 | 59 |52 | 1.7 GB |
229
+ | AMD Ryzen™ AI 5 HX 340 | NPU | FastFlowLM | LFM2.5-1.2B-Thinking| 63 | 59 |52 | 1.7 GB |
230
+ | AMD Ryzen™ AI 9 HX 370 | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking| 116 | N/A | N/A | 856 MB |
231
+ | Qualcomm Snapdragon® X Elite | NPU | NexaML | LFM2.5-1.2B-Thinking| 63 | N/A | N/A | 0.9 GB |
232
+ | Qualcomm Snapdragon® Gen4 (ROG Phone 9 Pro) | NPU | NexaML | LFM2.5-1.2B-Thinking| 82 | N/A | N/A | 0.9 GB |
233
+ | Qualcomm Dragonwing IQ9 (IQ-9075, IoT) | NPU | NexaML | LFM2.5-1.2B-Thinking | 53 | N/A | N/A | 0.9 GB |
234
+ | Qualcomm Snapdragon® Gen4 (Galaxy S25 Ultra) | CPU | llama.cpp (Q4_0) | LFM2.5-1.2B-Thinking| 70 | N/A | N/A | 719 MB |
235
+
236
+ **LFM2.5-1.2B-Thinking excels at long-context inference.**
237
+ On AMD NPUs with FastFlowLM, decoding throughput sustains ~46 tok/s even at the full 32K context, indicating robust long-context scalability.
238
+ See detailed benchmark results (up to full context length) [here](https://fastflowlm.com/docs/benchmarks/lfm2_results/).
239
 
240
  These capabilities unlock new deployment scenarios across various devices, including vehicles, mobile devices, laptops, IoT devices, and embedded systems.
241