EAAI Journal 2026 Journal Article
A conditional diffusion vision transformer model via data augmentation for few-shot fault diagnosis
- Beijia Zhao
- Dongsheng Yang
- Jiayue Sun
- Yanhong Luo
- Zhong Luo
- Xin Wang
The scarcity of labeled training data degrades the performance of accurate fault diagnosis models, highlighting the critical need for research in few-shot fault diagnosis (FSFD). Despite being a predominant FSFD solution, current data augmentation-based methods still suffer from distribution mismatch between generated and real data, as well as insufficient hierarchical diversity, particularly in fault severities. To overcome these limitations, the conditional diffusion vision transformer model (CDViT) is proposed for FSFD. CDViT leverages a dual-constrained denoising diffusion probabilistic model to accurately model the underlying distribution of real fault data. Subsequently, a fault severity attention module is designed to effectively extract fault severity features by capturing local–global hierarchical characteristics. Additionally, a fault refinement classifier is employed to better capture fault severity, improving diagnostic performance. The effectiveness of CDViT has been validated through multiple comparative experiments conducted on three datasets, demonstrating superior performance compared to mainstream methods in FSFD.