EAAI Journal 2026 Journal Article
A zero-shot anomaly detection network with patch-augmented prompts and test-time adaptation
- Hegui Zhu
- Chunmeng Zhao
- Yuan Yuan
- Meiling Liu
In recent years, pre-trained visual-language models, such as Contrastive Language-Image Pre-training (CLIP), have demonstrated strong generalization capabilities in zero-shot anomaly detection tasks. Due to the scarcity and diversity of anomaly samples, it is impossible to collect enough anomaly samples in many application scenarios, especially in privacy-sensitive scenarios. Moreover, most zero-shot anomaly detection methods with single and fixed text prompts face significant challenges in accurately detecting and localizing anomalies across domains. To address these challenges, this study proposes a zero-shot anomaly detection network, Patch-augmented Prompts and Test-time Adaptation (PPTA-CLIP), which incorporates four key modules to achieve robust anomaly detection performance. First, the modified visual encoder extracts multi-level visual features for subsequent tasks. Second, the patch-augmented text prompts module enhances learnable text prompts by incorporating local visual patch information, thus improving the model’s ability to detect subtle and rare anomaly patterns. Next, the modified text encoder module extracts enhanced text features. Finally, the test-time feature adaptation module refines the learned global features for test inputs, further improving detection performance. After being trained on an auxiliary dataset, the model can be directly used for testing without requiring retraining on other datasets. Extensive experimental results on 14 real-world datasets from industrial and medical fields show that PPTA-CLIP achieves state-of-the-art performance compared to other competitive zero-shot anomaly detection methods.