Qwen3.8-27B-QUASAR-NVFP4-mlx

MLX conversion of QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4, a 4-bit NVFP4 (W4A4) quantization-aware-trained build of Qwen/Qwen3.8-27B where all 496 linear layers (self-attention, gated delta-net, and MLP) are NVFP4. Converted from the source compressed-tensors nvfp4-pack-quantized checkpoint to the MLX-native nvfp4 layout for the omlx / mlx-vlm runtime on Apple silicon.

  • ~19.15 GiB: 4 weight shards (18.36 GiB) + mtp.safetensors (MTP draft head, 15 tensors, 810 MiB).
  • 1695 tensors: 496 packed NVFP4 weights (uint32) + 496 E4M3 scales (uint8)
    • 703 BF16 (embeddings, lm_head, vision tower, norms, conv1d, A_log, dt_bias)
    • the 15-tensor MTP head.

Conversion notes

MLX's nvfp4 kernel is single-level and does not carry the per-tensor global scale, so the two-level source scaling is folded into the per-group E4M3 scales: the E2M1 codes are kept bit-exact and each per-group scale is stored as E4M3(decode(weight_scale) / weight_global_scale). This single re-rounding of the (much smaller) scale tensor is the only precision change versus the source; the packed 4-bit weights themselves are byte-identical.

MTP norm weights carry the MLX +1.0 RMSNorm shift; the MTP head is a separate mtp.safetensors side file.

Produced with convert_vllm_nvfp4_to_mlx.py.

Quality benchmarks

Perplexity and KL divergence on the WikiText-2 raw test split (297,053 tokens, 2048-token windows, fresh KV cache per window), measured against the bf16 reference (Qwen3.8-27B-mlx).

Model Size (GiB) PPL (lower = better) ΔPPL vs bf16 KLD vs bf16 (nats/token, lower = better)
Qwen3.8-27B-mlx (bf16 reference) 51.75 6.935
Qwen3.8-27B-mlx-nvfp4-S 19.15 7.024 +0.088 0.0573
Qwen3.8-27B-QAT-NVFP4 19.15 7.298 +0.362 0.0753
Qwen3.8-27B-mlx-nvfp4-M 21.57 6.983 +0.048 0.0490
Qwen3.8-27B-mlx-nvfp4-L 22.31 6.987 +0.051 0.0407
Qwen3.8-27B-mlx-nvfp4-XL 28.84 7.024 +0.089 0.0382
Qwen3.8-27B-mlx-mxfp8-M 29.78 6.893 -0.042 0.0068
Qwen3.8-27B-mlx-mxfp8-L 34.79 6.937 +0.002 0.0046
Qwen3.8-27B-mlx-mxfp8-XL 36.31 6.937 +0.002 0.0037
  • KLD = mean per-token D_KL(p_bf16 ‖ p_model) over the full vocabulary — how much the quantized model's next-token distribution drifts from bf16.
  • Data: wikitext-2-raw-v1 test split, SHA-256 c5b5caea5bd655cb…; tokenizer: Qwen3.8-27B-mlx; mlx-vlm 0.6.17.
  • Benchmarked 2026-09-03 with benchmark_ppl_kld.py (2048-token streaming windows, all positions except the first of each window scored).

Plots

Mean KLD vs quant size

PPL vs quant size

Downloads last month
5
Safetensors
Model size
28B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for airagrp/Qwen3.8-27B-QAT-NVFP4

Base model

Qwen/Qwen3.8-27B
Quantized
(4)
this model