How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "zecanard/Qwopus3.6-27B-v2-MLX-3bit-mixed_3_6"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "zecanard/Qwopus3.6-27B-v2-MLX-3bit-mixed_3_6",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker
docker model run hf.co/zecanard/Qwopus3.6-27B-v2-MLX-3bit-mixed_3_6
Quick Links

🦆 zecanard/Qwopus3.6-27B-v2-MLX-3bit-mixed_3_6

This model was converted to MLX from Jackrong/Qwopus3.6-27B-v2 using mlx-vlm version 0.6.3. Please refer to the original model card for more details.

🌟 Quality

Mixed-precision quantized vision language model with an effective 4.308 bits per weight. Combines the size and speed benefits of a 3-bit quant with higher precision where it matters most.

mlx_vlm.convert --quantize --q-group-size 32 --quant-predicate mixed_3_6

🛠️ Customizations

This quant is aware of the current date, and also enables thinking (if available). You may disable this behavior by deleting the following line from the chat template, or changing true to false:

{%- set enable_thinking = true %}

A fix is also included for a thinking-related performance issue in Qwen 3.6.

🖥️ Use with mlx

pip install -U mlx-vlm
mlx_vlm.generate --model zecanard/Qwopus3.6-27B-v2-MLX-3bit-mixed_3_6 --max-tokens 100 --temperature 0 --prompt "Describe this image." --image <path_to_image>
Downloads last month
10
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zecanard/Qwopus3.6-27B-v2-MLX-3bit-mixed_3_6

Adapter
(18)
this model

Datasets used to train zecanard/Qwopus3.6-27B-v2-MLX-3bit-mixed_3_6