Instructions to use pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w2-w4-w8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w2-w4-w8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-audio", model="pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w2-w4-w8")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w2-w4-w8") model = AutoModelForMultimodalLM.from_pretrained("pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w2-w4-w8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Ovis Omni Embedding 3B โ experimental GPTQ W2/W4/W8
Experimental calibration-based GPTQ-derived variant of ATH-MaaS/Ovis-Omni-Embedding-3B.
Thinker layers 0โ13 use 2-bit weights, layers 14โ26 use 4-bit weights, and layer 27 uses 8-bit weights. The multimodal towers and non-linear/special modules remain in the original precision. Group scales were recomputed from the original BF16 weights after the compressor export produced invalid scales.
Compatibility warning: this W2 mixed artifact is published for experimentation, but vLLM 0.30.0 rejects its W2 packed shape in the Marlin compressed-tensors loader. Do not treat it as a vLLM-servable checkpoint. Use the separately published W4/W4/W8 vLLM-tested artifact for vLLM serving.
- Downloads last month
- 17
Model tree for pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w2-w4-w8
Base model
ATH-MaaS/Ovis-Omni-Embedding-3B