IS Journal 2026 Journal Article
Evaluating the Adversarial Robustness of Vision–Language Models for Facial Expression Recognition
- Hui Kuurila-Zhang
- Haoyu Chen
- Guoying Zhao
Facial expression recognition (FER) using vision–language models (VLMs) shows strong performance, but their robustness under adversarial conditions is underexplored. Noting that most VLMs rely on CLIP-style vision encoders vulnerable to gradient-based perturbations, we study how attacks on a CLIP encoder affect downstream recognition. We test zero-shot classifiers (CLIP, EVA-CLIP, Exp-CLIP) and generative VLMs (BLIP2, LLaVA, Qwen3-VL) performing FER via visual question answering. Using a CLIP surrogate for gradient-based attacks, we evaluate on AffectNet, RAF-DB, and FERPlus. Key findings include: 1) zero-shot classifiers are highly fragile, 2) attacks transfer only between models sharing the same vision encoder, and 3) generative VLMs are more robust than CLIP variants despite task-agnostic training. These results identify the vision encoder as the main bottleneck and highlight the need for robustness-focused design in future FER systems.