Qwen3.8-2.4T-A95B-NVFP4

NVFP4 quantization of Qwen/Qwen3.8-2.4T-A95B. Routed experts in NVFP4 (group size 16): attention, linear attention, shared experts, router gates, embeddings, and lm_head keep their original precision. Text-only, 262144 native context extended to 1M using RoPE at serving time.

Downloads last month
321
Safetensors
Model size
1.3T params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for modal-labs/Qwen3.8-2.4T-A95B-NVFP4

Quantized
(34)
this model
Finetunes
1 model