Arrow Research search
Back to ICRA

ICRA 2021

Distributed Heuristic Multi-Agent Path Finding with Communication

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

Multi-Agent Path Finding (MAPF) is essential to large-scale robotic systems. Recent methods have applied reinforcement learning (RL) to learn decentralized polices in partially observable environments. A fundamental challenge of obtaining collision-free policy is that agents need to learn co-operation to handle congested situations. This paper combines communication with deep Q-learning to provide a novel learning based method for MAPF, where agents achieve cooperation via graph convolution. To guide RL algorithm on long-horizon goal-oriented tasks, we embed the potential choices of shortest paths from single source as heuristic guidance instead of using a specific path as in most existing works. Our method treats each agent independently and trains the model from a single agent’s perspective. The final trained policy is applied to each agent for decentralized execution. The whole system is distributed during training and is trained under a curriculum learning strategy. Empirical evaluation in obstacle-rich environment indicates the high success rate with low average step of our method.

Authors

Keywords

  • Training
  • Learning systems
  • Convolution
  • Law enforcement
  • Scalability
  • Robot kinematics
  • Heuristic algorithms
  • Pathfinding
  • Multi-Agent Path Finding
  • Single Agent
  • Shortest Path
  • Graph Convolution
  • Specific Path
  • Average Step
  • Deep Q-learning
  • Path Choice
  • Time Step
  • Convolutional Layers
  • Actual Values
  • Attention Mechanism
  • Finite Set
  • Joint Action
  • Number Of Agents
  • Binary Matrix
  • Multi-agent Systems
  • Reward Function
  • Deep Reinforcement Learning
  • Deep Q-network
  • Multi-agent Reinforcement Learning
  • Position Of Agent
  • Action-value Function
  • Communication Latency
  • Adjacent Vertices
  • Attention Heads
  • Policy Gradient Method
  • Total Return
  • Local Observations

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
292871982943232847
v2026.09.13