Arrow Research search
Back to IROS

IROS 2021

Centralizing State-Values in Dueling Networks for Multi-Robot Reinforcement Learning Mapless Navigation

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

We study the problem of multi-robot mapless navigation in the popular Centralized Training and Decentralized Execution (CTDE) paradigm. This problem is challenging when each robot considers its path without explicitly sharing observations with other robots and can lead to non-stationary issues in Deep Reinforcement Learning (DRL). The typical CTDE algorithm factorizes the joint action-value function into individual ones, to favor cooperation and achieve decentralized execution. Such factorization involves constraints (e. g. , monotonicity) that limit the emergence of novel behaviors in an individual as each agent is trained starting from a joint action-value. In contrast, we propose a novel architecture for CTDE that uses a centralized state-value network to compute a joint state-value, which is used to inject global state information in the value-based updates of the agents. Consequently, each model computes its gradient update for the weights, considering the overall state of the environment. Our idea follows the insights of Dueling Networks as a separate estimation of the joint state-value has both the advantage of improving sample efficiency, while providing each robot information whether the global state is (or is not) valuable. Experiments in a robotic navigation task with 2 4, and 8 robots, confirm the superior performance of our approach over prior CTDE methods (e. g. , VDN, QMIX).

Authors

Keywords

  • Training
  • Navigation
  • Computational modeling
  • Estimation
  • Reinforcement learning
  • Computer architecture
  • Task analysis
  • Mapless Navigation
  • Factorization
  • Global Status
  • Global Information
  • Deep Reinforcement Learning
  • Separate Estimates
  • Navigation Task
  • Robot Navigation
  • Robotic Tasks
  • Action-value Function
  • Poor Performance
  • Minimum Distance
  • Target Location
  • Linear Velocity
  • Training Environment
  • Target Network
  • Independent Learning
  • Average Path Length
  • Process Of Agents
  • Deep Reinforcement Learning Algorithm
  • Multi-agent Reinforcement Learning
  • Double Deep Q-network
  • Separate Streams
  • Value-based Approach
  • Robot Operating System
  • Average Reward
  • Motivating Example
  • Deep Q-network
  • Minutes Of Training

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
822459214520422373
v2026.09.13