--- library_name: mlx pipeline_tag: image-text-to-text base_model: QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 tags: - mlx - nvfp4 - fp4 - quantization - quantization-aware-training - quasar - qwen3_5 --- # Qwen3.8-27B-QUASAR-NVFP4-mlx MLX conversion of [`QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4`](https://huggingface.co/QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4), a 4-bit NVFP4 (W4A4) quantization-aware-trained build of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B) where **all 496 linear layers** (self-attention, gated delta-net, and MLP) are NVFP4. Converted from the source `compressed-tensors` `nvfp4-pack-quantized` checkpoint to the MLX-native `nvfp4` layout for the omlx / mlx-vlm runtime on Apple silicon. - ~19.15 GiB: 4 weight shards (18.36 GiB) + `mtp.safetensors` (MTP draft head, 15 tensors, 810 MiB). - 1695 tensors: 496 packed NVFP4 weights (`uint32`) + 496 E4M3 scales (`uint8`) + 703 BF16 (embeddings, lm_head, vision tower, norms, conv1d, A_log, dt_bias) + the 15-tensor MTP head. ## Conversion notes MLX's `nvfp4` kernel is **single-level** and does not carry the per-tensor global scale, so the two-level source scaling is folded into the per-group E4M3 scales: the E2M1 codes are kept bit-exact and each per-group scale is stored as `E4M3(decode(weight_scale) / weight_global_scale)`. This single re-rounding of the (much smaller) scale tensor is the only precision change versus the source; the packed 4-bit weights themselves are byte-identical. MTP norm weights carry the MLX `+1.0` RMSNorm shift; the MTP head is a separate `mtp.safetensors` side file. Produced with `convert_vllm_nvfp4_to_mlx.py`. ## Quality benchmarks Perplexity and KL divergence on the WikiText-2 raw test split (297,053 tokens, 2048-token windows, fresh KV cache per window), measured against the bf16 reference (`Qwen3.8-27B-mlx`). | Model | Size (GiB) | PPL (lower = better) | ΔPPL vs bf16 | KLD vs bf16 (nats/token, lower = better) | |---|---|---|---|---| | `Qwen3.8-27B-mlx` (bf16 reference) | 51.75 | 6.935 | — | — | | [`Qwen3.8-27B-mlx-nvfp4-S`](https://huggingface.co/airagrp/Qwen3.8-27B-mlx-nvfp4-S) | 19.15 | 7.024 | +0.088 | 0.0573 | | [`Qwen3.8-27B-QAT-NVFP4`](https://huggingface.co/airagrp/Qwen3.8-27B-QUASAR-NVFP4-mlx) | 19.15 | 7.298 | +0.362 | 0.0753 | | [`Qwen3.8-27B-mlx-nvfp4-M`](https://huggingface.co/airagrp/Qwen3.8-27B-mlx-nvfp4-M) | 21.57 | 6.983 | +0.048 | 0.0490 | | [`Qwen3.8-27B-mlx-nvfp4-L`](https://huggingface.co/airagrp/Qwen3.8-27B-mlx-nvfp4-L) | 22.31 | 6.987 | +0.051 | 0.0407 | | [`Qwen3.8-27B-mlx-nvfp4-XL`](https://huggingface.co/airagrp/Qwen3.8-27B-mlx-nvfp4-XL) | 28.84 | 7.024 | +0.089 | 0.0382 | | [`Qwen3.8-27B-mlx-mxfp8-M`](https://huggingface.co/airagrp/Qwen3.8-27B-mlx-mxfp8-M) | 29.78 | 6.893 | -0.042 | 0.0068 | | [`Qwen3.8-27B-mlx-mxfp8-L`](https://huggingface.co/airagrp/Qwen3.8-27B-mlx-mxfp8-L) | 34.79 | 6.937 | +0.002 | 0.0046 | | [`Qwen3.8-27B-mlx-mxfp8-XL`](https://huggingface.co/airagrp/Qwen3.8-27B-mlx-mxfp8-XL) | 36.31 | 6.937 | +0.002 | 0.0037 | - **KLD** = mean per-token `D_KL(p_bf16 ‖ p_model)` over the full vocabulary — how much the quantized model's next-token distribution drifts from bf16. - Data: `wikitext-2-raw-v1` test split, SHA-256 `c5b5caea5bd655cb…`; tokenizer: `Qwen3.8-27B-mlx`; mlx-vlm 0.6.17. - Benchmarked 2026-09-03 with `benchmark_ppl_kld.py` (2048-token streaming windows, all positions except the first of each window scored). ### Plots ![Mean KLD vs quant size](benchmarks/kld-vs-size.png) ![PPL vs quant size](benchmarks/ppl-vs-size.png)