Instructions to use QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ") model = AutoModelForMultimodalLM.from_pretrained("QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ
- SGLang
How to use QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ with Docker Model Runner:
docker model run hf.co/QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ
/no_think does not work
hi, I notice with this model the /no_think that works on Qwen/Qwen3-32B-AWQ does not work as intended.
also I notice it does not start it's reply with .
I am curious if is it a problem with the base model or is it a problem only with this quantized version?
{
"messages": [
{
"content": "/no_think What is the capital of Switzerland?",
"role": "user"
}
],
"model": "QuantTrio/Qwen3-VL-30B-A3B-Instruct-AWQ"
}
reply
{
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Okay, the user is asking for the capital of Switzerland. Hmm, this seems straightforward, but let me double-check because sometimes people confuse it with other cities like Zurich or Geneva. I remember that Bern is the capital, but I should verify to be absolutely sure. \n\nThe user didn't provide any extra context, so they probably just want a quick factual answer. No signs of confusion or deeper needs here—just a simple geography question. Maybe they're studying, traveling, or just curious. \n\nI'll confirm: Switzerland's capital is Bern, not Zurich (which is a major city but not the capital) or Geneva (which hosts the UN). Yep, Bern is correct. The answer should be short and clear since the user didn't ask for details. \n\nWait—should I mention why it's not Zurich? Nah, they didn't ask for explanation, so just the answer. But if they follow up, I can elaborate on the \"why\" later. \n\nFinal answer: Bern. Simple and precise.\n</think>\n\nThe capital of Switzerland is **Bern**. \n\nWhile Zurich and Geneva are major Swiss cities (Zurich is the largest city and Geneva is known for international organizations), **Bern** is the official federal capital, where the Swiss government and parliament are located. \n\nThis is a common point of confusion, but Bern has been the capital since 1848. 🇨🇭",
"refusal": null,
"annotations": null,
"audio": null,
"function_call": null,
"tool_calls": [],
"reasoning_content": null
},
"logprobs": null,
"finish_reason": "stop",
"stop_reason": null,
"token_ids": null
}
],
}
This 30B-VL model is an Instruct model, not a Thinking one.
Please note that starting from the Qwen3 2507 release, the Instruct and Thinking models were split into two separate variants.
Since then, the no_think flag in prompts is no longer effective.
For this model in particular, regardless of whether your prompt includes think or no_think,
the output will not contain tags.
If you want to use the Thinking mode, please refer to
👉 QuantTrio/Qwen3-VL-30B-A3B-Thinking-AWQ