Ovis Omni Embedding 3B โ€” GPTQ W4/W4/W8 (vLLM-tested)

Experimental calibration-based weight-only GPTQ variant of ATH-MaaS/Ovis-Omni-Embedding-3B.

The Thinker transformer uses 4-bit weights for layers 0โ€“26 and 8-bit weights for layer 27. Audio, vision, talker, token2wav, embeddings, and norms remain in the original precision. Calibration used eight short text samples with LLM Compressor 0.14.0 and compressed-tensors 0.19.0.

This checkpoint includes a packaging repair: the compressor export contained invalid group scales, so scales were recomputed from the original BF16 weights per group before the vLLM test. Treat this as an experimental GPTQ-derived checkpoint and benchmark retrieval quality before production use.

vLLM test

Tested with vLLM 0.30.0 on an 8-GiB RTX 3080 Laptop GPU:

vllm serve . --runner pooling --convert embed \
  --pooler-config '{"dimensions":1024}' \
  --max-model-len 256 --gpu-memory-utilization 0.75 --enforce-eager

The server reached readiness and /v1/embeddings returned finite, normalized 2048-dimensional vectors. The Docker image for this artifact requests 1024 dimensions by default.

Full-context image

The companion image profchan/kogwistar-llm-wiki-embedding:vllm-gptq-w4-w4-w8-fullcontext uses vLLM 0.30.0 with --max-model-len 32768, FP8 KV cache, one sequence, chunked prefill, and 2 GiB CPU offload. It was tested on the 8-GiB RTX 3080 Laptop GPU with a tokenizer-counted 32,701-token request. The default 1024 pooler returned a finite normalized vector (norm=0.999999997). The native 2048-dimensional output was also tested successfully. CPU offload requires adequate host RAM and is slower than the constrained 256-token configuration.

Independent screening and fidelity status

The full-context vLLM image passed a tokenizer-counted 32,701-token embedding request on the RTX 3080 Laptop GPU. This proves serving/context capacity, not embedding fidelity. Cosine drift, nearest-neighbor agreement, and held-out Recall@K must be reported against the persisted BF16 reference before this experimental GPTQ derivative is treated as a quality-preserving replacement.

The completed 60-case BF16 comparison measured mean cosine -0.06758287, mean drift 1.06758287, nearest-neighbor agreement 23.33%, Recall@1 23.33%, Recall@5 63.33%, Recall@10 100%, and nDCG 0.53568221. The 32K serving result therefore does not establish retrieval fidelity for this experimental checkpoint.

Downloads last month
44
Safetensors
Model size
2.94k params
Tensor type
BF16
ยท
I32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for pt810/Ovis-Omni-Embedding-3B-gptq-mixed-w4-w4-w8

Quantized
(2)
this model