MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval
Abstract
A new benchmark and steering method improve retrieval coverage of diverse perspectives across domains and modalities for open-ended queries.
Information retrieval (IR) increasingly targets open-ended queries that admit diverse perspectives. Existing IR benchmarks, however, focus primarily on closed-ended queries, while even open-ended benchmarks largely consist of queries whose supporting documents span a single subject domain and modality. We introduce Multi^3IR, a benchmark that evaluates how well retrievers cover the multifaceted perspectives of open-ended queries across diverse domains and modalities. It comprises 104.9K Stack Exchange queries, each annotated with perspective descriptions that capture the query's implicit viewpoints. We further propose SPIN, a parameter- and label-efficient method that learns noise vectors to steer embeddings toward diverse yet meaningful semantic directions. Experiments show that existing multimodal retrievers suffer from single-perspective bias, while SPIN substantially improves perspective coverage on Multi^3IR and generalizes well to unseen open-ended IR benchmarks. The dataset and experimental code are available at https://github.com/seokwon99/Multi3IR.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation (2026)
- KAMR: Grounding Generation via Knowledge-Aligned Multi-hop Retrieval (2026)
- Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering (2026)
- Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability (2026)
- Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval (2026)
- UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering (2026)
- Diverse-Intent Multi-Turn Fashion Image Retrieval (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.30949 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper