You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Otzaria Qwen3 Embedding GGUF

GGUF export of the merged Otzaria Qwen3 embedding model.

Source merged model

EMD123/otzaria-qwen3-embedding

llama.cpp version used

b9933

Files

  • otzaria-qwen3-embedding-0.6b-merged-F16.gguf — 1.12 GB
  • otzaria-qwen3-embedding-0.6b-merged-Q4_K_M.gguf — 0.37 GB

Recommendation

For embedding quality, start with:

  • F16 for maximum quality.
  • Q8_0 for smaller size with usually limited quality loss.

Use Q4_K_M only after checking retrieval metrics.

Important limitation

This is a retrieval model. It retrieves relevant texts; it does not provide halachic rulings by itself.

Downloads last month
-
GGUF
Model size
0.6B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EMD123/otzaria-qwen3-embedding-GGUF

Quantized
(1)
this model