INT8 (W8A16, Marlin) Quantisierungen der Qwen3-Embedding-Modelle - Text (4B/8B) und multimodal (VL-8B). Fuer vLLM/CUDA.