Instructions to use pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w4-w4-w8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w4-w4-w8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-audio", model="pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w4-w4-w8")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w4-w4-w8") model = AutoModelForMultimodalLM.from_pretrained("pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w4-w4-w8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Ovis Omni Embedding 3B โ GPTQ W4/W4/W8 (vLLM-tested)
Experimental calibration-based weight-only GPTQ variant of ATH-MaaS/Ovis-Omni-Embedding-3B.
The Thinker transformer uses 4-bit weights for layers 0โ26 and 8-bit weights for layer 27. Audio, vision, talker, token2wav, embeddings, and norms remain in the original precision. Calibration used eight short text samples with LLM Compressor 0.14.0 and compressed-tensors 0.19.0.
This checkpoint includes a packaging repair: the compressor export contained invalid group scales, so scales were recomputed from the original BF16 weights per group before the vLLM test. Treat this as an experimental GPTQ-derived checkpoint and benchmark retrieval quality before production use.
vLLM test
Tested with vLLM 0.30.0 on an 8-GiB RTX 3080 Laptop GPU:
vllm serve . --runner pooling --convert embed \
--pooler-config '{"dimensions":1024}' \
--max-model-len 256 --gpu-memory-utilization 0.75 --enforce-eager
The server reached readiness and /v1/embeddings returned finite, normalized
2048-dimensional vectors. The Docker image for this artifact requests 1024
dimensions by default.
Full-context image
The companion image
profchan/kogwistar-llm-wiki-embedding:vllm-gptq-w4-w4-w8-fullcontext
uses vLLM 0.30.0 with --max-model-len 32768, FP8 KV cache, one sequence,
chunked prefill, and 2 GiB CPU offload. It was tested on the 8-GiB RTX 3080
Laptop GPU with a tokenizer-counted 32,701-token request. The default 1024
pooler returned a finite normalized vector (norm=0.999999997). The native
2048-dimensional output was also tested successfully. CPU offload requires
adequate host RAM and is slower than the constrained 256-token configuration.
Independent screening and fidelity status
The full-context vLLM image passed a tokenizer-counted 32,701-token embedding request on the RTX 3080 Laptop GPU. This proves serving/context capacity, not embedding fidelity. Cosine drift, nearest-neighbor agreement, and held-out Recall@K must be reported against the persisted BF16 reference before this experimental GPTQ derivative is treated as a quality-preserving replacement.
The completed 60-case BF16 comparison measured mean cosine -0.06758287, mean drift 1.06758287, nearest-neighbor agreement 23.33%, Recall@1 23.33%, Recall@5 63.33%, Recall@10 100%, and nDCG 0.53568221. The 32K serving result therefore does not establish retrieval fidelity for this experimental checkpoint.
- Downloads last month
- 44
Model tree for pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w4-w4-w8
Base model
ATH-MaaS/Ovis-Omni-Embedding-3B