randmaru's picture
Update README.md
4ee144d verified
|
Raw History Blame
2.58 kB
metadata
license: apache-2.0
base_model:
  - Qwen/Qwen3.8-27B
library_name: mlx
tags:
  - text-generation-inference
pipeline_tag: text-generation

randmaru/Qwen3.8-27B-mlx-mxfp4-text-generation-only

This is an MXFP4 MLX quantization of Qwen/Qwen3.8-27BQwen3.8-27B for Apple Silicon inference.

Text-generation-only: all 333 vision weights in BF16 were removed.


MXFP4 vs 4Bit Quantization Comparison

Parameter MXFP4 4Bit
Quantization format 4‑bit floating point with microscaling, group 32, shared exponent E8M0 4‑bit integer (INT4/NF4)
Tensor types U8, U32, BF16 BF16, U32
Parameter size (safetensors) ~14.3 GB ~15.13 GB
Total storage (all files) ~14.32 GB ~15.15 GB
Hardware support Most efficient on GPUs with microscaling / FP8 tensor core support Broad support, but often requires specialized INT4 kernels
Apple Silicon compatibility Designed with hardware microscaling support in Apple Neural Engine / GPU Works, but without specialized Neural Engine optimization
Inference speed Higher on compatible hardware: FP path, lower dequantization overhead, higher throughput Kernel‑dependent; usually lower or comparable at similar quality
Quality Better preserves dynamic range, less degradation on outliers Higher risk of accuracy loss on outliers at the same bitrate

Key takeaways:

  • Parameter size: The MXFP4 version has a smaller safetensors footprint: ~14.3 GB vs ~15.13 GB for the 4-bit version.
  • Total storage: MXFP4 also occupies less total disk space: ~14.32 GB vs ~15.15 GB when summing all repository files.
  • Vision weights removed: Both repositories are text-generation-only. All 333 vision weights in BF16 were removed to reduce size and focus inference on text generation.
  • Performance: MXFP4 typically delivers higher inference throughput on hardware with microscaling/FP8 support, especially on Apple Silicon. The shared exponent per group of 32 elements reduces dequantization overhead and enables use of floating‑point tensor cores.
  • Apple Silicon optimization: MXFP4 is designed with Apple Neural Engine / GPU microscaling support in mind, making it the recommended choice for MacBook.
  • Quality: MXFP4 preserves dynamic range better, so generation quality can be higher at the same compression level.

Actual speed depends on the backend, GPU, batch size, and quantization implementation.