--- library_name: transformers base_model: Qwen/Qwen2.5-VL-3B-Instruct tags: - qwen2_5_vl - quantized - fp4 - mxfp4 - w4a4 - compressed-tensors - llm-compressor - vllm license: apache-2.0 language: - en - zh pipeline_tag: image-text-to-text --- # Qwen2.5-VL-3B-Instruct-MXFP4-W4A4-LM-Only This is an **MXFP4 W4A4 quantized** version of [Qwen/Qwen2.5-VL-3B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct), created using [llm-compressor](https://github.com/vllm-project/llm-compressor). **Only the LLM decoder is quantized.** The Vision Transformer (ViT) encoder remains in BF16 precision. ## Model Summary | Property | Value | |----------|-------| | **Base Model** | [Qwen/Qwen2.5-VL-3B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct) | | **Quantization** | MXFP4 W4A4 (4-bit float weights, 4-bit float activations) | | **Quantization Scope** | LLM decoder only (ViT encoder in BF16) | | **Strategy** | Per-tensor-group (group_size=32), symmetric (minmax observer) | | **Format** | `compressed-tensors` (`mxfp4-pack-quantized`) | | **Model Size** | ~3.3 GB (1 shard) | | **Ignored Layers** | `lm_head`, all `model.visual.*` layers | | **Tool** | [llm-compressor](https://github.com/vllm-project/llm-compressor) | | **Supported Runtime** | [vLLM](https://github.com/vllm-project/vllm) (with `compressed-tensors`) | ## Quantization Details - **Weights**: FP4 E2M1 (4-bit float), per-tensor-group (group_size=32) with E8M0 (uint8 exponent) block scales, static quantization - **Activations**: FP4 E2M1 (4-bit float), per-tensor-group (group_size=32), dynamic quantization - **Ignored**: `lm_head` (kept in BF16) and all ViT encoder layers (`model.visual.*`) - **Calibration**: 512 samples from CNN/DailyMail, max_seq_length=2048 - **Standard**: [OCP MX Specification](https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf) ## Quantization Recipe ```yaml quant_stage: quant_modifiers: QuantizationModifier: ignore: ["lm_head", "re:model.visual.*"] scheme: "MXFP4" targets: ["Linear"] ``` ## Hardware Requirements MXFP4 inference requires **NVIDIA Blackwell (SM120+)** GPUs with CUDA 12.8+ for native CUTLASS MXFP4 GEMM support. ## Usage ### With vLLM ```bash export VLLM_ATTENTION_BACKEND=TORCH_SDPA vllm serve JongYeop/Qwen2.5-VL-3B-Instruct-MXFP4-W4A4-LM-Only \ --trust-remote-code \ --max-model-len 4096 \ --enforce-eager ``` ### With Transformers ```python from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor from qwen_vl_utils import process_vision_info model = Qwen2_5_VLForConditionalGeneration.from_pretrained( "JongYeop/Qwen2.5-VL-3B-Instruct-MXFP4-W4A4-LM-Only", torch_dtype="auto", device_map="auto", ) processor = AutoProcessor.from_pretrained("Qwen/Qwen2.5-VL-3B-Instruct") messages = [{"role": "user", "content": [ {"type": "image", "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg"}, {"type": "text", "text": "Describe this image in detail."}, ]}] text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) image_inputs, video_inputs = process_vision_info(messages) inputs = processor(text=[text], images=image_inputs, videos=video_inputs, return_tensors="pt").to(model.device) output = model.generate(**inputs, max_new_tokens=256) result = processor.batch_decode(output[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True) print(result[0]) ``` ## Model Architecture - **Architecture**: Qwen2_5_VLForConditionalGeneration - **LLM Hidden Size**: 2048 - **LLM Intermediate Size**: 11008 - **LLM Layers**: 36 - **LLM Attention Heads**: 16 - **LLM KV Heads**: 2 (GQA) - **Vision Encoder Depth**: 32 - **Vision Hidden Size**: 1280 - **Vocabulary Size**: 151936 ## Differences from NVFP4 | Feature | MXFP4 | NVFP4 | |---------|-------|-------| | Scale Format | E8M0 (uint8 exponent) | E4M3 + FP32 global scale | | Group Size | 32 | 16 | | Standard | OCP MX Specification | NVIDIA proprietary | | Hardware | SM120+ (Blackwell) | SM89+ (Ada/Hopper/Blackwell) | ## Acknowledgments - Base model by [Qwen Team](https://huggingface.co/Qwen) - Quantization powered by [llm-compressor](https://github.com/vllm-project/llm-compressor) and [compressed-tensors](https://github.com/neuralmagic/compressed-tensors)