EAAI Journal 2026 Journal Article
Leveraging multi-strategy labels for long-tailed classification
- Wu Zeng
- Mei Li
Most datasets in the real world predominantly exhibit long-tailed distribution characteristics. This inter-class imbalance phenomenon in datasets causes trained models to focus excessively on majority-class samples. Meanwhile, the models exhibit significantly poorer detection performance for minority-class samples. Although existing data augmentation methods can generate new samples that incorporate minority-class pixel regions. This approach improves the model’s recognition performance for minority-class samples. However, these methods typically assign labels to new samples solely based on the proportional area of pixel coverage between samples. This practice of ignoring semantic correlations among new samples often results in generated labels that substantially deviate from actual sample content, thereby providing erroneous supervisory signals that ultimately misguide the model training process. To address this issue, this paper proposes a multi-strategy mixed label with semantic awareness for recognition (MMLSAR) method. Specifically, we first construct a semantic predictor connected to the backbone network to more accurately assess the semantic composition of augmented images. Secondly, we propose an adaptive semantic adjustment module that enables dynamic allocation of label values. This allocation is achieved by integrating geometric features of cropped regions from augmented samples with semantic information from sample content. Finally, we design an augmented-sample label prediction auxiliary task. This task provides superior supervisory signals to the model by improving the prediction accuracy of augmented sample labels. We validated the effectiveness of our method on multiple datasets. In engineering applications, this technology can provide technical support to address class imbalance issues in real-world scenarios.