Instructions to use Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Local Apps Settings

How to use Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8 with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Use Docker

docker model run hf.co/Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8

SGLang

How to use Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8 with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'

Docker Model Runner
How to use Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8 with Docker Model Runner:
```
docker model run hf.co/Qwen/Qwen2-VL-7B-Instruct-GPTQ-Int8
```

shuai bai commited on Sep 21, 2024

Commit

0768122

verified ·

1 Parent(s): 3d152a7

Update README.md

Browse files

Files changed (1) hide show

README.md +4 -3

README.md CHANGED Viewed

@@ -503,9 +503,10 @@ These limitations serve as ongoing directions for model optimization and improve
 If you find our work helpful, feel free to give us a cite.
 ```
-@article{Qwen2-VL,
-  title={Qwen2-VL},
-  author={Qwen team},
   year={2024}
 }

 If you find our work helpful, feel free to give us a cite.
 ```
+@article{Qwen2VL,
+  title={Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution},
+  author={Wang, Peng and Bai, Shuai and Tan, Sinan and Wang, Shijie and Fan, Zhihao and Bai, Jinze and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Fan, Yang and Dang, Kai and Du, Mengfei and Ren, Xuancheng and Men, Rui and Liu, Dayiheng and Zhou, Chang and Zhou, Jingren and Lin, Junyang},
+  journal={arXiv preprint arXiv:2409.12191},
   year={2024}
 }