Gaze Attention, Qwen3-4B, video

Video model of the paper Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs (Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun; NAVER AI Lab, KAIST).

Multimodal LLMs attend to all visual tokens at every generation step. With Gaze Attention, the cached keys of the visual tokens are grouped into spatial regions, and every text query attends only to the TopK regions that match it best, selected per layer, per head and per token. A few learnable context tokens per image or frame keep the global view of the scene. The KV cache stays complete, so the attended regions change from one generated token to the next.

Model

Language model Qwen3-4B-Instruct-2507
Vision encoder SigLIP2-So400m/16, 384px (finetuned)
Projector 2-layer MLP with a 3 x 3 convolution (stride 2): 4x fewer visual tokens
Visual tokens 144 per frame
Gaze regions 4 regions of 6 x 6 tokens per frame
Attended visual KV entries per query 10% of the regions of all frames + 4 context tokens per frame (64 frames: 1,192 of 9,216 visual tokens)
Context tokens 4 per frame
Parameters 4.5B, stored in fp16

Usage

The model is loaded with the cambrian package of the Gaze Attention repository:

git clone https://github.com/junha1125/Gaze-Attention
cd Gaze-Attention
pip install torch==2.2.2 torchvision==0.17.2
pip install -e .
import copy
import torch

from cambrian.constants import DEFAULT_IMAGE_TOKEN, IMAGE_TOKEN_INDEX
from cambrian.conversation import conv_templates
from cambrian.mm_utils import tokenizer_image_token
from cambrian.model.builder import load_pretrained_model

tokenizer, model, image_processor, _, load_video = load_pretrained_model(
    "junha1125/gaze-attention-qwen3-4b-video", None, "cambrian_qwen3", device_map="cuda:0", return_video_loader=True
)
model.eval()

frames, frame_idx, meta = load_video("assets/jobs.mp4", video_fps=1, frames_upbound=64)   # 1 FPS, at most 64 frames
video = image_processor.preprocess(frames, return_tensors="pt")["pixel_values"].to(dtype=torch.float16, device="cuda:0")

conv = copy.deepcopy(conv_templates["qwen_3"])
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nDescribe what happens in this video.")
conv.append_message(conv.roles[1], None)
input_ids = tokenizer_image_token(conv.get_prompt(), tokenizer, IMAGE_TOKEN_INDEX, return_tensors="pt").unsqueeze(0).to("cuda:0")

output_ids = model.generate(input_ids, images=[video], modalities=["video"], do_sample=False, max_new_tokens=256)
print(tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0])

or python demo/video_inference.py --model-path junha1125/gaze-attention-qwen3-4b-video --video assets/jobs.mp4 --max-frames 64.

The model was trained with inputs of up to 12,288 tokens (about 80 frames). For more frames load it with overwrite_config={"tokenizer_model_max_length": 24576}.

gaze_topk_ratio (0.1 in this checkpoint) is the fraction of gaze regions that a query attends to. It can be changed when the model is loaded: load_pretrained_model(..., overwrite_config={"gaze_topk_ratio": 0.25}). The generation config of Qwen3 samples by default; pass do_sample=False for reproducible outputs.

Training

Stage Trained parts Data Learning rate Batch size
1-1 projector Cambrian-Alignment (1.88M samples) 1e-3 512
1-2 projector, context tokens Cambrian-Alignment 1e-3 512
2 all (images, TopK ratio 0.6, then 0.1) two 17% subsets of Cambrian-7M (1.17M samples each) 1e-5 (vision encoder 2e-6) 256
3 all (videos, 1 FPS, at most 64 frames) Cambrian-S-3M subset: 776K video samples + 235K text-only samples 1e-5 (vision encoder 2e-6) 256

Stage 3 uses the progressive TopK schedule (0.9 to 0.1 during the first 60% of the training steps). Recipe: scripts/train/video_384 of the repository.

License

Apache License 2.0. The model is built on Qwen3-4B-Instruct-2507 and SigLIP2 (both Apache 2.0). The training datasets have their own licenses.

Citation

@inproceedings{song2026gazeattention,
  title={Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal {LLM}s},
  author={Junha Song and Byeongho Heo and Geonmo Gu and Jaegul Choo and Dongyoon Han and Sangdoo Yun},
  booktitle={Third Conference on Language Modeling (CoLM)},
  year={2026}
}
Downloads last month
30
Safetensors
Model size
5B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for junha1125/gaze-attention-qwen3-4b-video

Finetuned
(2363)
this model

Datasets used to train junha1125/gaze-attention-qwen3-4b-video

Paper for junha1125/gaze-attention-qwen3-4b-video