Arrow Research search
Back to IJCAI

IJCAI 2020

Embodied Multimodal Multitask Learning

Conference Paper Machine Learning Artificial Intelligence

Abstract

Visually-grounded embodied language learning models have recently shown to be effective at learning multiple multimodal tasks such as following navigational instructions and answering questions. In this paper, we address two key limitations of these models, (a) the inability to transfer the grounded knowledge across different tasks and (b) the inability to transfer to new words and concepts not seen during training using only a few examples. We propose a multitask model which facilitates knowledge transfer across tasks by disentangling the knowledge of words and visual attributes in the intermediate representations. We create scenarios and datasets to quantify cross-task knowledge transfer and show that the proposed model outperforms a range of baselines in simulated 3D environments. We also show that this disentanglement of representations makes our model modular and interpretable which allows for transfer to instructions containing new concepts.

Authors

Keywords

  • Machine Learning Applications: Applications of Reinforcement Learning
  • Machine Learning: Deep Reinforcement Learning
  • Machine Learning: Transfer, Adaptation, Multi-task Learning

Context

Venue
International Joint Conference on Artificial Intelligence
Archive span
1969-2025
Indexed papers
14525
Paper id
451377765487299488
v2026.09.13