Arrow Research search
Back to ICML

ICML 2025

Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAG

Conference Paper Accept (oral) Artificial Intelligence ยท Machine Learning

Abstract

High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs). To drive progress beyond the limits of heuristic methods, this paper advances HR perception capabilities of MLLMs by harnessing cutting-edge long-context techniques such as retrieval-augmented generation (RAG). Towards this end, this paper presents the first study exploring the use of RAG to address HR perception challenges. Specifically, we propose Retrieval-Augmented Perception (RAP), a training-free framework that retrieves and fuses relevant image crops while preserving spatial context using the proposed Spatial-Awareness Layout. To accommodate different tasks, the proposed Retrieved-Exploration Search (RE-Search) dynamically selects the optimal number of crops based on model confidence and retrieval scores. Experimental results on HR benchmarks demonstrate the significant effectiveness of RAP, with LLaVA-v1. 5-13B achieving a 43% improvement on $V^*$ Bench and 19% on HR-Bench. Code is available at https: //github. com/DreamMr/RAP.

Authors

Keywords

  • Multimodal Large Language Models
  • High-resolution Image Perception

Context

Venue
International Conference on Machine Learning
Archive span
1993-2025
Indexed papers
16471
Paper id
821802019949060805
v2026.09.13