Image-Text-to-Text
Transformers
Safetensors
qwen2_5_vl
hyperclick
gui-grounding
multimodal
confidence-calibration
reinforcement-learning
conversational
text-generation-inference
Instructions to use SeerRay-Lab/Qwen2.5-VL-7B-HyperClick with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SeerRay-Lab/Qwen2.5-VL-7B-HyperClick with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="SeerRay-Lab/Qwen2.5-VL-7B-HyperClick") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("SeerRay-Lab/Qwen2.5-VL-7B-HyperClick") model = AutoModelForMultimodalLM.from_pretrained("SeerRay-Lab/Qwen2.5-VL-7B-HyperClick", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SeerRay-Lab/Qwen2.5-VL-7B-HyperClick with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SeerRay-Lab/Qwen2.5-VL-7B-HyperClick" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/Qwen2.5-VL-7B-HyperClick", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/SeerRay-Lab/Qwen2.5-VL-7B-HyperClick
- SGLang
How to use SeerRay-Lab/Qwen2.5-VL-7B-HyperClick with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SeerRay-Lab/Qwen2.5-VL-7B-HyperClick" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/Qwen2.5-VL-7B-HyperClick", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SeerRay-Lab/Qwen2.5-VL-7B-HyperClick" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/Qwen2.5-VL-7B-HyperClick", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use SeerRay-Lab/Qwen2.5-VL-7B-HyperClick with Docker Model Runner:
docker model run hf.co/SeerRay-Lab/Qwen2.5-VL-7B-HyperClick
Add HyperClick-7B model card with usage and evaluation
Browse files
README.md
CHANGED
|
@@ -1,3 +1,133 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen2.5-VL-7B-Instruct
|
| 4 |
+
base_model_relation: finetune
|
| 5 |
+
library_name: transformers
|
| 6 |
+
pipeline_tag: image-text-to-text
|
| 7 |
+
tags:
|
| 8 |
+
- hyperclick
|
| 9 |
+
- gui-grounding
|
| 10 |
+
- multimodal
|
| 11 |
+
- confidence-calibration
|
| 12 |
+
- reinforcement-learning
|
| 13 |
+
- arxiv:2510.27266
|
| 14 |
---
|
| 15 |
+
|
| 16 |
+
# HyperClick-7B
|
| 17 |
+
|
| 18 |
+
[Paper](https://arxiv.org/pdf/2510.27266) 路 [Code](https://github.com/xiaomi-research/hyperclick) 路 [3B model](https://huggingface.co/SeerRay-Lab/Qwen2.5-VL-3B-HyperClick) 路 [7B model](https://huggingface.co/SeerRay-Lab/Qwen2.5-VL-7B-HyperClick)
|
| 19 |
+
|
| 20 |
+
## Overview
|
| 21 |
+
|
| 22 |
+
HyperClick grounds natural-language instructions in GUI screenshots and predicts a click point together with an explicit confidence score. This repository contains the **7B checkpoint**, based on [Qwen2.5-VL-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct), with full model weights in Safetensors format.
|
| 23 |
+
|
| 24 |
+
The training framework combines supervised fine-tuning with reinforcement fine-tuning. Its rewards check output format, grounding correctness, and confidence alignment using a truncated Gaussian spatial target and the Brier score. See the [training code](https://github.com/xiaomi-research/hyperclick/blob/main/src/open-r1-multimodal/src/open_r1/hyperclick.py) for details.
|
| 25 |
+
|
| 26 |
+

|
| 27 |
+
|
| 28 |
+
## Reported results
|
| 29 |
+
|
| 30 |
+
Grounding accuracy (%) from the [project evaluation table](https://github.com/xiaomi-research/hyperclick#evaluation). These are the project's reported results, not a new evaluation of the uploaded files.
|
| 31 |
+
|
| 32 |
+
| Model | ScreenSpot | ScreenSpot-v2 | ScreenSpot-Pro | MMBench-GUI | UI-I2E-Bench | CAGUI | UI-Vision |
|
| 33 |
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
| 34 |
+
| HyperClick-3B | 88.5 | 90.6 | 41.3 | 71.4 | 71.8 | 81.0 | 19.6 |
|
| 35 |
+
| HyperClick-7B | 91.5 | 93.7 | 48.2 | 79.6 | 76.5 | 82.9 | 25.7 |
|
| 36 |
+
|
| 37 |
+
## Quick start
|
| 38 |
+
|
| 39 |
+
The example below uses Transformers on a CUDA GPU. Install PyTorch for your CUDA environment, then install the inference dependencies:
|
| 40 |
+
|
| 41 |
+
```bash
|
| 42 |
+
pip install "transformers==4.49.0" "accelerate==1.10.0" "qwen-vl-utils==0.0.11" pillow
|
| 43 |
+
```
|
| 44 |
+
|
| 45 |
+
Replace `screenshot.png` and the instruction with your own input. The prompt follows the HyperClick training template.
|
| 46 |
+
|
| 47 |
+
```python
|
| 48 |
+
import torch
|
| 49 |
+
from PIL import Image
|
| 50 |
+
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
|
| 51 |
+
from qwen_vl_utils import process_vision_info
|
| 52 |
+
|
| 53 |
+
model_id = "SeerRay-Lab/Qwen2.5-VL-7B-HyperClick"
|
| 54 |
+
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
|
| 55 |
+
model_id, torch_dtype=torch.bfloat16, device_map="auto"
|
| 56 |
+
).eval()
|
| 57 |
+
processor = AutoProcessor.from_pretrained(
|
| 58 |
+
model_id, min_pixels=3136, max_pixels=4390400, use_fast=False
|
| 59 |
+
)
|
| 60 |
+
|
| 61 |
+
screenshot = Image.open("screenshot.png").convert("RGB")
|
| 62 |
+
instruction = "Click the search button"
|
| 63 |
+
messages = [dict(role="user", content=[
|
| 64 |
+
dict(type="image", image=screenshot, min_pixels=3136, max_pixels=4390400),
|
| 65 |
+
dict(type="text", text=(
|
| 66 |
+
f'Point to the element related to the instruction "{instruction}" '
|
| 67 |
+
'on the screenshot with your confidence.'
|
| 68 |
+
)),
|
| 69 |
+
])]
|
| 70 |
+
prompt = processor.apply_chat_template(
|
| 71 |
+
messages, tokenize=False, add_generation_prompt=True
|
| 72 |
+
)
|
| 73 |
+
images, _ = process_vision_info(messages)
|
| 74 |
+
inputs = processor(text=[prompt], images=images, return_tensors="pt").to(model.device)
|
| 75 |
+
with torch.inference_mode():
|
| 76 |
+
generated = model.generate(**inputs, max_new_tokens=128, do_sample=False)
|
| 77 |
+
answer = processor.batch_decode(
|
| 78 |
+
generated[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
|
| 79 |
+
)[0]
|
| 80 |
+
print(answer)
|
| 81 |
+
|
| 82 |
+
# Dimensions of the image coordinate space used by the model.
|
| 83 |
+
patch_size = processor.image_processor.patch_size
|
| 84 |
+
_, grid_h, grid_w = inputs.image_grid_thw[0].tolist()
|
| 85 |
+
input_width, input_height = grid_w * patch_size, grid_h * patch_size
|
| 86 |
+
print("Model image size:", input_width, input_height)
|
| 87 |
+
```
|
| 88 |
+
|
| 89 |
+
### Output format and coordinates
|
| 90 |
+
|
| 91 |
+
The expected output format is:
|
| 92 |
+
|
| 93 |
+
```text
|
| 94 |
+
<point>[x,y]</point><confidence>conf</confidence>
|
| 95 |
+
```
|
| 96 |
+
|
| 97 |
+
`x` and `y` are pixel coordinates in the processed image; `conf` is a confidence estimate between 0 and 1. To map a predicted point back to the original screenshot, use:
|
| 98 |
+
|
| 99 |
+
```python
|
| 100 |
+
# x and y are parsed from the model response.
|
| 101 |
+
x_original = x * screenshot.width / input_width
|
| 102 |
+
y_original = y * screenshot.height / input_height
|
| 103 |
+
```
|
| 104 |
+
|
| 105 |
+
Image resizing affects the coordinate system. Validate the response format and point bounds before using a prediction. Confidence is a learned estimate and can be incorrect, especially for unfamiliar interfaces or ambiguous instructions.
|
| 106 |
+
|
| 107 |
+
## Training
|
| 108 |
+
|
| 109 |
+
The [GitHub repository](https://github.com/xiaomi-research/hyperclick) provides training setup, example annotation formats, and the reinforcement fine-tuning entry points:
|
| 110 |
+
|
| 111 |
+
```bash
|
| 112 |
+
bash src/open-r1-multimodal/run_hyperclick_7b.sh
|
| 113 |
+
```
|
| 114 |
+
|
| 115 |
+
## License
|
| 116 |
+
|
| 117 |
+
This repository retains its Apache-2.0 license designation. See the [base model license](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct/blob/main/LICENSE).
|
| 118 |
+
|
| 119 |
+
## Citation
|
| 120 |
+
|
| 121 |
+
The latest arXiv version is titled *Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning*.
|
| 122 |
+
|
| 123 |
+
```bibtex
|
| 124 |
+
@misc{zhang2025hyperclick,
|
| 125 |
+
title={Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning},
|
| 126 |
+
author={Shaojie Zhang and Pei Fu and Ruoceng Zhang and Jiahui Yang and Anan Du and Xiuwen Xi and Shaokang Wang and Ying Huang and Bin Qin and Zhenbo Luo and Jian Luan},
|
| 127 |
+
year={2025},
|
| 128 |
+
eprint={2510.27266},
|
| 129 |
+
archivePrefix={arXiv},
|
| 130 |
+
primaryClass={cs.CV},
|
| 131 |
+
url={https://arxiv.org/abs/2510.27266}
|
| 132 |
+
}
|
| 133 |
+
```
|