Arrow Research search
Back to ICML

ICML 2025

When and How Does CLIP Enable Domain and Compositional Generalization?

Conference Paper Accept (spotlight poster) Artificial Intelligence ยท Machine Learning

Abstract

The remarkable generalization performance of contrastive vision-language models like CLIP is often attributed to the diversity of their training distributions. However, key questions remain unanswered: Can CLIP generalize to an entirely unseen domain when trained on a diverse mixture of domains (domain generalization)? Can it generalize to unseen classes within partially seen domains (compositional generalization)? What factors affect such generalization? To answer these questions, we trained CLIP models on systematically constructed training distributions with controlled domain diversity and object class exposure. Our experiments show that domain diversity is essential for both domain and compositional generalization, yet compositional generalization can be surprisingly weaker than domain generalization when the training distribution contains a suboptimal subset of the test domain. Through data-centric and mechanistic analyses, we find that successful generalization requires the learning of sufficiently shared representations in intermediate layers and circuits.

Authors

Keywords

  • CLIP
  • Compositional Generalization
  • Domain Generalization
  • Out-of-Distribution Robustness
  • OOD generalization

Context

Venue
International Conference on Machine Learning
Archive span
1993-2025
Indexed papers
16471
Paper id
728383163170201747
v2026.09.13