--- license: other license_name: tongyi-qianwen license_link: https://huggingface.co/Qwen/Qwen-VL-Chat/blob/main/LICENSE base_model: Qwen/Qwen-VL-Chat library_name: peft pipeline_tag: image-text-to-text tags: - graph-learning - multimodal-graphs - vision-language-model - qwen-vl - lora - node-classification - link-prediction ---
# OMG-VLM ### One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models **Accepted at EMNLP 2026 · Main Conference** Jiayi Yang† · Yifang Chen† · Yuanfu Sun · Jiajin Liu · Qiaoyu Tan
† Equal contribution [![Paper](https://img.shields.io/badge/arXiv-2607.19128-B31B1B?style=flat-square&logo=arxiv&logoColor=white)](https://arxiv.org/abs/2607.19128) [![Code](https://img.shields.io/badge/GitHub-Code-181717?style=flat-square&logo=github)](https://github.com/Jo-eyang/OMG-VLM) ![Qwen-VL](https://img.shields.io/badge/Backbone-Qwen--VL--Chat-6B5BFF?style=flat-square) Overview of the OMG-VLM framework

One shared vision-language backbone for text-attributed, image-attributed, and multimodal graphs.

This repository holds the fine-tuned **OMG-VLM** weights: a checkpoint with 32 visual queries and 8 text context tokens, for use with the code at [github.com/Jo-eyang/OMG-VLM](https://github.com/Jo-eyang/OMG-VLM). ## Overview **OMG-VLM** is a unified vision-language framework for attributed graph learning under heterogeneous modality schemas. It builds on **Qwen-VL** and adds graph-aware adapters that inject neighborhood signals into the VLM-native embedding space for both image and text attributes. | Item | Description | | --- | --- | | Task family | Attributed graph learning with image, text, or mixed node attributes | | Backbone | Qwen-VL-Chat | | Graph modules | Graph-aware visual adapter and target-aware textual aggregation | | This checkpoint | LoRA adapter + both graph modules | - **Graph-aware visual adapter:** center image features attend to compressed neighbor visual features in the Qwen-VL embedding space. - **Target-aware textual aggregation:** target node text retrieves and compresses neighbor text into learnable `` context tokens. ## Files | File | Contents | |---|---| | `adapter_model.safetensors`, `adapter_config.json` | LoRA adapter for the Qwen-VL-Chat language model (r=64, α=16; `attn.c_attn`, `attn.c_proj`, `w1`, `w2`) | | `visual_adapter.pth` | Graph-aware visual adapter (`transformer.visual_adapter`) | | `textual_aggregation.pth` | Target-aware textual aggregation (`transformer.textual_aggregation`) | The base Qwen-VL-Chat weights are **not** included. Download them from [Qwen/Qwen-VL-Chat](https://huggingface.co/Qwen/Qwen-VL-Chat). ## Configuration These values must match at inference time: | Setting | Value | |---|---| | Visual adapter layers / queries / heads | 1 / 32 / 32 | | Visual compressor | on | | Textual aggregation heads / pool layers / MLP ratio | 16 / 1 / 4.0 | | Textual aggregation context tokens | 8 | | Max neighbors | 10 | ## Usage This checkpoint is loaded with the evaluation script from the OMG-VLM repository. ```bash # 1. Code git clone https://github.com/Jo-eyang/OMG-VLM.git cd OMG-VLM pip install -r requirements.txt # 2. Base weights go into the repo's Qwen_VL_Chat/ folder, next to the OMG-VLM model code. # Download only the weight shards so the repo's config and code are kept. hf download Qwen/Qwen-VL-Chat --local-dir Qwen_VL_Chat \ --include "pytorch_model-*.bin" "pytorch_model.bin.index.json" # 3. This checkpoint hf download oofwite/OMG-VLM --local-dir checkpoints/omg_vlm # 4. Evaluate (the repo root must be on PYTHONPATH so the model code can import omg_vlm) export PYTHONPATH=$(pwd):$PYTHONPATH python evaluate_omg_vlm.py \ --model_name_or_path Qwen_VL_Chat \ --adapter_path checkpoints/omg_vlm \ --data_path /path/to/test.json \ --neighbor_data_path /path/to/neighbors.json \ --text_info_path /path/to/text_info.json \ --output_path outputs/predictions.jsonl \ --max_neighbors 10 \ --visual_adapter_num_layers 1 --visual_adapter_num_queries 32 --visual_adapter_num_heads 32 \ --textual_aggregation_num_heads 16 --textual_aggregation_pool_layers 1 \ --textual_aggregation_pool_mlp_ratio 4.0 --textual_aggregation_context_tokens 8 ``` For the conversation, neighbor, and text-attribute file formats, and for training your own checkpoint, see the [Data Format](https://github.com/Jo-eyang/OMG-VLM#data-format) and [Training](https://github.com/Jo-eyang/OMG-VLM#training) sections of the GitHub README. **Scoring note:** `evaluate_omg_vlm.py` reports strict string exact match, so for example `yes` against a target of `yes.` counts as wrong. For yes/no or label tasks, normalize case and punctuation before scoring. Tested with `torch 2.7.1`, `transformers 4.37.2`, and `peft 0.10.0`. ## Training - **Base model:** Qwen-VL-Chat; the vision encoder is frozen. - **Data:** a mixture of node-classification and link-prediction tasks over image, text, and multimodal graphs: - Amazon Arts (node classification); - arXiv (node classification); - RedditS (link prediction, image-only); - Movies (link prediction). - **Optimization:** 3 epochs; learning rate 1e-5 with cosine schedule and 1% warmup; weight decay 0.1; effective batch size 8; max sequence length 2048. ## License These weights are derived from Qwen-VL-Chat and are subject to the [Tongyi Qianwen License Agreement](https://huggingface.co/Qwen/Qwen-VL-Chat/blob/main/LICENSE). ## Citation ```bibtex @inproceedings{yang2026one, title={One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models}, author={Yang, Jiayi and Chen, Yifang and Sun, Yuanfu and Liu, Jiajin and Tan, Qiaoyu}, booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing}, year={2026} } ``` ## Acknowledgment This work builds on Qwen-VL and widely used open-source tooling from the VLM/LLM ecosystem. We thank the authors and maintainers of Qwen-VL, Hugging Face Transformers, PEFT, FastChat, DeepSpeed, and related projects.