Arrow Research search
Back to AAAI

AAAI 2024

Multimodal Ensembling for Zero-Shot Image Classification

Short Paper AAAI Undergraduate Consortium Artificial Intelligence

Abstract

Artificial intelligence has made significant progress in image classification, an essential task for machine perception to achieve human-level image understanding. Despite recent advances in vision-language fields, multimodal image classification is still challenging, particularly for the following two reasons. First, models with low capacity often suffer from underfitting and thus underperform on fine-grained image classification. Second, it is important to ensure high-quality data with rich cross-modal representations of each class, which is often difficult to generate. Here, we utilize ensemble learning to reduce the impact of these issues on pre-trained models. We aim to create a meta-model that combines the predictions of multiple open-vocabulary multimodal models trained on different data to create more robust and accurate predictions. By utilizing ensemble learning and multimodal machine learning, we will achieve higher prediction accuracies without any additional training or fine-tuning, meaning that this method is completely zero-shot.

Authors

Keywords

  • Fine-Grained Image Classification
  • Image Classification
  • machine learning
  • Machine Perception
  • Multimodal Machine Learning

Context

Venue
AAAI Conference on Artificial Intelligence
Archive span
1980-2026
Indexed papers
28718
Paper id
863065175350991583
v2026.09.13