lfm2-5-8b-a1b-mxfp4-mlx

MLX quantization of LiquidAI/LFM2.5-8B-A1B for Apple Silicon.

Variant: Block float MX FP4
Disk size: 4308 MB
Quantized by: sahilchachra

Benchmark results

Evaluated on Apple M5 Pro with MLX. Model loaded once; performance and quality measured in a single pass.

Performance

This model FP16 baseline
Decode tok/s (steady-state) 205.03 77.32
Prefill tok/s (steady-state) 450.86 227.43
Decode tok/s (avg, long traces) 196.26 76.34
Peak memory (GB) 4.893 17.142
Disk size (MB) 4308 16169

Warmed, short-prompt, chat-templated, thinking disabled. Represents steady-state decode for typical chat use; long thinking traces will be slower due to KV-cache growth.

Quality

Benchmark This model FP16 baseline n
MATH-500 (math reasoning) 56.7% (answered 23/30) 70.0% (answered 25/30) 30
IFEval (instruction following) 77.3% 79.5% 44
HumanEval (code, pass@1) 73.3% 73.3% 30

MATH-500 per-level accuracy

Level This model FP16 baseline
level 1 83.3% 83.3%
level 2 83.3% 100.0%
level 3 50.0% 50.0%
level 4 66.7% 66.7%
level 5 0.0% 50.0%

Context scaling (decode tok/s)

Context length Decode tok/s
~128 tokens 206.1
~256 tokens 205.4
~512 tokens 205.6
~1024 tokens 204.5

Usage

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("sahilchachra/lfm2-5-8b-a1b-mxfp4-mlx")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True)

All variants in this collection

Model Variant
sahilchachra/lfm2-5-8b-a1b-mxfp4-mlx Block float MX FP4 ← this model
sahilchachra/lfm2-5-8b-a1b-optiq-5bpw-mlx OptiQ mixed-precision (target 5.0 bpw)

Notes

  • Requires Apple Silicon (M1 or later) with MLX
  • Benchmarks run on Apple M5 Pro, 24 GB unified memory
  • License: see LiquidAI/LFM2.5-8B-A1B for the original model's license

Original model

See LiquidAI/LFM2.5-8B-A1B for full model details and intended use.

Downloads last month
17
Safetensors
Model size
8B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sahilchachra/lfm2-5-8b-a1b-mxfp4-mlx

Quantized
(99)
this model