--- license: apache-2.0 language: - en pipeline_tag: video-text-to-text base_model: - Qwen/Qwen3-4B-Instruct-2507 - google/siglip2-so400m-patch16-384 datasets: - nyu-visionx/Cambrian-Alignment - nyu-visionx/Cambrian-10M - nyu-visionx/Cambrian-S-3M - lmms-lab/LLaVA-Video-178K tags: - gaze-attention - multimodal - vision-language - cambrian --- # Gaze Attention, Qwen3-4B, video Video model of the paper **Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs** (Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun; NAVER AI Lab, KAIST). - Paper: https://arxiv.org/abs/2605.13080 - Code: https://github.com/junha1125/Gaze-Attention - Project page: https://june-page.github.io/gaze-attention Multimodal LLMs attend to all visual tokens at every generation step. With Gaze Attention, the cached keys of the visual tokens are grouped into spatial regions, and every text query attends only to the TopK regions that match it best, selected per layer, per head and per token. A few learnable context tokens per image or frame keep the global view of the scene. The KV cache stays complete, so the attended regions change from one generated token to the next. ## Model | | | |---|---| | Language model | Qwen3-4B-Instruct-2507 | | Vision encoder | SigLIP2-So400m/16, 384px (finetuned) | | Projector | 2-layer MLP with a 3 x 3 convolution (stride 2): 4x fewer visual tokens | | Visual tokens | 144 per frame | | Gaze regions | 4 regions of 6 x 6 tokens per frame | | Attended visual KV entries per query | 10% of the regions of all frames + 4 context tokens per frame (64 frames: 1,192 of 9,216 visual tokens) | | Context tokens | 4 per frame | | Parameters | 4.5B, stored in fp16 | ## Usage The model is loaded with the `cambrian` package of the [Gaze Attention repository](https://github.com/junha1125/Gaze-Attention): ```bash git clone https://github.com/junha1125/Gaze-Attention cd Gaze-Attention pip install torch==2.2.2 torchvision==0.17.2 pip install -e . ``` ```python import copy import torch from cambrian.constants import DEFAULT_IMAGE_TOKEN, IMAGE_TOKEN_INDEX from cambrian.conversation import conv_templates from cambrian.mm_utils import tokenizer_image_token from cambrian.model.builder import load_pretrained_model tokenizer, model, image_processor, _, load_video = load_pretrained_model( "junha1125/gaze-attention-qwen3-4b-video", None, "cambrian_qwen3", device_map="cuda:0", return_video_loader=True ) model.eval() frames, frame_idx, meta = load_video("assets/jobs.mp4", video_fps=1, frames_upbound=64) # 1 FPS, at most 64 frames video = image_processor.preprocess(frames, return_tensors="pt")["pixel_values"].to(dtype=torch.float16, device="cuda:0") conv = copy.deepcopy(conv_templates["qwen_3"]) conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nDescribe what happens in this video.") conv.append_message(conv.roles[1], None) input_ids = tokenizer_image_token(conv.get_prompt(), tokenizer, IMAGE_TOKEN_INDEX, return_tensors="pt").unsqueeze(0).to("cuda:0") output_ids = model.generate(input_ids, images=[video], modalities=["video"], do_sample=False, max_new_tokens=256) print(tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0]) ``` or `python demo/video_inference.py --model-path junha1125/gaze-attention-qwen3-4b-video --video assets/jobs.mp4 --max-frames 64`. The model was trained with inputs of up to 12,288 tokens (about 80 frames). For more frames load it with `overwrite_config={"tokenizer_model_max_length": 24576}`. `gaze_topk_ratio` (0.1 in this checkpoint) is the fraction of gaze regions that a query attends to. It can be changed when the model is loaded: `load_pretrained_model(..., overwrite_config={"gaze_topk_ratio": 0.25})`. The generation config of Qwen3 samples by default; pass `do_sample=False` for reproducible outputs. ## Training | Stage | Trained parts | Data | Learning rate | Batch size | |---|---|---|---|---:| | 1-1 | projector | Cambrian-Alignment (1.88M samples) | 1e-3 | 512 | | 1-2 | projector, context tokens | Cambrian-Alignment | 1e-3 | 512 | | 2 | all (images, TopK ratio 0.6, then 0.1) | two 17% subsets of Cambrian-7M (1.17M samples each) | 1e-5 (vision encoder 2e-6) | 256 | | 3 | all (videos, 1 FPS, at most 64 frames) | Cambrian-S-3M subset: 776K video samples + 235K text-only samples | 1e-5 (vision encoder 2e-6) | 256 | Stage 3 uses the progressive TopK schedule (0.9 to 0.1 during the first 60% of the training steps). Recipe: `scripts/train/video_384` of the repository. ## License Apache License 2.0. The model is built on [Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) and [SigLIP2](https://huggingface.co/google/siglip2-so400m-patch16-384) (both Apache 2.0). The training datasets have their own licenses. ## Citation ```bibtex @inproceedings{song2026gazeattention, title={Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal {LLM}s}, author={Junha Song and Byeongho Heo and Geonmo Gu and Jaegul Choo and Dongyoon Han and Sangdoo Yun}, booktitle={Third Conference on Language Modeling (CoLM)}, year={2026} } ```