---
license: other
license_name: tongyi-qianwen
license_link: https://huggingface.co/Qwen/Qwen-VL-Chat/blob/main/LICENSE
base_model: Qwen/Qwen-VL-Chat
library_name: peft
pipeline_tag: image-text-to-text
tags:
- graph-learning
- multimodal-graphs
- vision-language-model
- qwen-vl
- lora
- node-classification
- link-prediction
---
# OMG-VLM
### One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models
**Accepted at EMNLP 2026 · Main Conference**
Jiayi Yang
† · Yifang Chen
† · Yuanfu Sun · Jiajin Liu · Qiaoyu Tan
† Equal contribution
[](https://arxiv.org/abs/2607.19128)
[](https://github.com/Jo-eyang/OMG-VLM)

One shared vision-language backbone for text-attributed, image-attributed, and multimodal graphs.
This repository holds the fine-tuned **OMG-VLM** weights: a checkpoint with 32 visual queries and 8 text context tokens, for use with the code at [github.com/Jo-eyang/OMG-VLM](https://github.com/Jo-eyang/OMG-VLM).
## Overview
**OMG-VLM** is a unified vision-language framework for attributed graph learning under heterogeneous modality schemas. It builds on **Qwen-VL** and adds graph-aware adapters that inject neighborhood signals into the VLM-native embedding space for both image and text attributes.
| Item | Description |
| --- | --- |
| Task family | Attributed graph learning with image, text, or mixed node attributes |
| Backbone | Qwen-VL-Chat |
| Graph modules | Graph-aware visual adapter and target-aware textual aggregation |
| This checkpoint | LoRA adapter + both graph modules |
- **Graph-aware visual adapter:** center image features attend to compressed neighbor visual features in the Qwen-VL embedding space.
- **Target-aware textual aggregation:** target node text retrieves and compresses neighbor text into learnable `