Qwen3.6-35B-A3B-MXFP4-GGUF

MXFP4 dense quant of ggml-org/Qwen3.6-35B-A3B-GGUF, from llama.cpp PR timlikesai/llama.cpp#14 ("ggml: mxfp4 quantization, KV cache and Blackwell MMA"). New dense ftype LLAMA_FTYPE_MOSTLY_MXFP4 (=42).

Quantized with the UOS e8m0 block scale (MXAttention paper, arXiv:2607.24377): Q_max = 7.25 for E2M1, all 2D tensors MXFP4, token embeddings and output tensor at Q8_0 until MXFP8 lands.

Comparisons measured vs: Q4_K_M, MXFP4_MOE. Full data in the PR.

Downloads last month
320
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for timtimtimtimtim/Qwen3.6-35B-A3B-MXFP4-GGUF

Quantized
(1)
this model

Paper for timtimtimtimtim/Qwen3.6-35B-A3B-MXFP4-GGUF