EAAI Journal 2026 Journal Article
Dual-level information interactive learning model for text-image person Re-identification
- Jia Sun
- Yanfeng Li
- Houjin Chen
- Luyifu Chen
- Minjun Wang
Text-image person re-identification (TI-ReID), which retrieves corresponding person images via textual descriptions, stands out as a prominent research area within the field of object tracking. The core challenge of TI-ReID lies in the significant discrepancies between the text and image modalities, making it difficult to effectively associate positive sample pairs. Existing studies only focus on the relationships between samples at the instance-level, and overly emphasizing the uniqueness of positive pairs, resulting in insufficient learnable content for the model. In this study, building upon the improvement of instance-level work, we introduce a class-level learning component and put forward a novel Dual-level Information Interactive Learning (DIIL) model. The aim is to jointly learn the inter-modal correlation relationships from both the class-level and the instance-level. Specifically, DIIL consists of two principal components: (1) a class-level teacher guidance (CTG) module that constructs two sample embedding banks at the class-level to provide more comprehensive guidance for instance samples. (2) an instance-level information blending (IIB) module that establishes the bidirectional correlation between text and image from the two perspectives of mask prediction and information blending, thus fully narrowing the gap between the features of the two modalities. We conduct sufficient experiments on three public datasets, and the experimental results demonstrate that the DIIL model achieves state-of-the-art results, especially in terms of the mean average precision (mAP).