EAAI Journal 2026 Journal Article
FEAN: A Fragments Embedding and Aligning Network for image-text matching
- Xianlun Tang
- Lin Jiang
- Yu Xia
- Wuquan Deng
- Bo Tang
- Tianyu Xiang
- Hao Zhu
Most recent image-text matching methods use fragments alignment structures to achieve cross modal interaction. Although they compensate for the lack of cross modal interaction caused by overall embedding, the accuracy of image-text matching is compromised by the cross modal semantic gap and redundant alignment. To address these issues, we propose a Fragments Embedding and Aligning Network (FEAN) to achieve accurate image-text matching, which focuses on the representation of global and local image-text features and their similarity. Specifically, we embed local image and textual features into fragments to obtain global image and textual features representations and global similarity scores, and use cross modal weight calculation to obtain fragments aligned local image and textual features and local similarity matrices, in order to mitigate the cross modal semantic gap. To avoid redundant alignment in image and textual fragments, a Similarity Pooling (SP) strategy is proposed to aggregate global similarity scores and local similarity matrices. In addition, to compensate for the missing contextual semantic information on the image region features, we add the Position Weight Feature Reinforcement (PWFR) module before embedding alignment to achieve consistency with the text semantic information. Extensive experiments on two publicly available datasets for image-text matching, Flickr30K and MSCOCO, have shown that our fragments embedding and aligning network achieve the best R@1 and RSUM values in bidirectional matching with optimal accuracy.