hsmin92's picture
Add AWQ W4A16 g128 checkpoint, model card, and quantization metadata
ff32ca7 verified
Raw
History Blame Contribute Delete
541 Bytes
InternVL3.5-4B-HF AWQ W4A16 (group size 128)
This model is a quantized derivative of:
OpenGVLab/InternVL3_5-4B-HF
https://huggingface.co/OpenGVLab/InternVL3_5-4B-HF
The upstream project is licensed under the Apache License 2.0.
The quantized checkpoint preserves the upstream model architecture and files,
with the language decoder Linear weights compressed to asymmetric INT4
(group size 128) using the AWQ algorithm via llm-compressor.
The vision tower, multimodal projector, input embeddings, and lm_head are
left unquantized in BF16.