Arrow Research search
Back to IROS

IROS 2020

MAPPER: Multi-Agent Path Planning with Evolutionary Reinforcement Learning in Mixed Dynamic Environments

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Multi-agent navigation in dynamic environments is of great industrial value when deploying a large scale fleet of robot to real-world applications. This paper proposes a decentralized partially observable multi-agent path planning with evolutionary reinforcement learning (MAPPER) method to learn an effective local planning policy in mixed dynamic environments. Reinforcement learning-based methods usually suffer performance degradation on long-horizon tasks with goal-conditioned sparse rewards, so we decompose the long-range navigation task into many easier sub-tasks under the guidance of a global planner, which increases agents' performance in large environments. Moreover, most existing multi-agent planning approaches assume either perfect information of the surrounding environment or homogeneity of nearby dynamic agents, which may not hold in practice. Our approach models dynamic obstacles' behavior with an image-based representation and trains a policy in mixed dynamic environments without homogeneity assumption. To ensure multi-agent training stability and performance, we propose an evolutionary training approach that can be easily scaled to large and complex environments. Experiments show that MAPPER is able to achieve higher success rates and more stable performance when exposed to a large number of non-cooperative dynamic obstacles compared with traditional reaction-based planner LRA* and the state-of-the-art learning-based method.

Authors

Keywords

  • Training
  • Learning systems
  • Navigation
  • Reinforcement learning
  • Path planning
  • Planning
  • Task analysis
  • Dynamic Environment
  • Evolutionary Reinforcement Learning
  • Complex Environment
  • Real-world Applications
  • Local Policy
  • Learning-based Methods
  • Policy Planning
  • Global Plan
  • Reinforcement Learning Methods
  • Dynamic Obstacles
  • Neural Network
  • Learning Rate
  • Value Function
  • Deep Neural Network
  • Evolutionary Algorithms
  • Sensor Data
  • Training Procedure
  • Easy Task
  • Number Of Agents
  • Reference Path
  • Markov Decision Process
  • Planning Problem
  • Dynamic Objects
  • Central Method
  • Planning Methods
  • Navigation Problem
  • Local Observations
  • Large-scale Environments
  • Stochastic Policy

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
60219538379780997
v2026.09.13