EAAI Journal 2026 Journal Article
A novel multi-modal attentional collaborative learning framework with semantic enhancement for audio–visual question answering
- Jie Yang
- Miao Ma
- Peng Wang
- Yutong Li
- Zhao Pei
- Chao Yao
- Longjiang Guo
The Audio–Visual Question Answering (AVQA) task aims to extract audio and visual cues from videos for answering the questions. The popular two-stage method, such as Progressive Spatio-Temporal Perception Network (PSTP-Net), first locates key segments in the audio–visual scene based on the question and then identifies the most relevant audio–visual regions. While this reduces cue redundancy, it overlooks the complementary role of rich cues, which is crucial for a comprehensive understanding of audio–visual content. In this paper, we propose a novel framework to start from the question itself, guide the entire multi-modal collaborative learning process, and conduct audio–visual question answering. This method includes a semantically enhanced strategy using Multi-modal Large Language Models (MLLMs) applied as an engineering solution, and a multi-modal attentional collaborative learning process, which is the core algorithmic innovation. Extensive experiments on the Music Audio–Visual Question Answering dataset (MUSIC-AVQA) and Music Audio–Visual Question Answering dataset version 2 (MUSIC-AVQA v2) demonstrate the effectiveness of our method. Compared to the PSTP-Net, our method reduces the number of training parameters by 61. 23% and Floating-point Operations (FLOPs) by 60. 83%, while achieving 2. 61 percentage-point improvement in accuracy. This indicates that our method effectively captures and aligns rich audio–visual cues, significantly enhancing reasoning efficiency. Our code will be publicly available soon.