Arrow Research search
Back to IROS

IROS 2024

Multimodal Evolutionary Encoder for Continuous Vision-Language Navigation

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Can multimodal encoder evolve when facing increasingly tough circumstances? Our work investigates this possibility in the context of continuous vision-language navigation (continuous VLN), which aims to navigate robots under linguistic supervision and visual feedback. We propose a multimodal evolutionary encoder (MEE) comprising a unified multimodal encoder architecture and an evolutionary pre-training strategy. The unified multimodal encoder unifies rich modalities, including depth and sub-instruction, to enhance the solid understanding of environments and tasks. It also effectively utilizes monocular observation, reducing the reliance on panoramic vision. The evolutionary pre-training strategy exposes the encoder to increasingly unfamiliar data domains and difficult objectives. The multi-stage adaption helps the encoder establish robust intra- and inter-modality connections and improve its generalization to unfamiliar environments. To achieve such evolution, we collect a large-scale multi-stage dataset with specialized objectives, addressing the absence of suitable continuous VLN pre-training. Evaluation on VLN-CE demonstrates the superiority of MEE over other direct action-predicting methods. Furthermore, we deploy MEE in real scenes using self-developed service robots, showcasing its effectiveness and potential for real-world applications. Our code and dataset are available at https://github.com/RavenKiller/MEE.

Authors

Keywords

  • Visualization
  • Costs
  • Codes
  • Navigation
  • Service robots
  • Linguistics
  • Feature extraction
  • Solids
  • Decoding
  • Intelligent robots
  • Vision-language Navigation
  • Visual Feedback
  • Monocular
  • Evolutionary Strategy
  • Real Scenes
  • Understanding Tasks
  • Stage 2
  • Feature Space
  • Discrete Set
  • Depth Information
  • Hierarchical Method
  • Successive Stages
  • Reconstruction Loss
  • Quality Of Representations
  • Attention Layer
  • t-SNE Plot
  • Imitation Learning
  • Attention Block
  • Pre-training Stage
  • Navigation Performance
  • Partial Alignment
  • Navigation Path
  • Navigation Model
  • Alignment Loss
  • Hierarchical Groups
  • Limited Visibility
  • Feature Representation
  • Ablation

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
192063557273671159
v2026.09.13