Ovis Omni Embedding 3B โ€” experimental GPTQ W2/W4/W8

Experimental calibration-based GPTQ-derived variant of ATH-MaaS/Ovis-Omni-Embedding-3B.

Thinker layers 0โ€“13 use 2-bit weights, layers 14โ€“26 use 4-bit weights, and layer 27 uses 8-bit weights. The multimodal towers and non-linear/special modules remain in the original precision. Group scales were recomputed from the original BF16 weights after the compressor export produced invalid scales.

Compatibility warning: this W2 mixed artifact is published for experimentation, but vLLM 0.30.0 rejects its W2 packed shape in the Marlin compressed-tensors loader. Do not treat it as a vLLM-servable checkpoint. Use the separately published W4/W4/W8 vLLM-tested artifact for vLLM serving.

Downloads last month
17
Safetensors
Model size
2.94k params
Tensor type
BF16
ยท
I32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w2-w4-w8

Quantized
(2)
this model