Arrow Research search
Back to IROS

IROS 2016

D++: Structural credit assignment in tightly coupled multiagent domains

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Autonomous multi-robot teams can be used in complex coordinated exploration tasks to improve exploration performance in terms of both speed and effectiveness. However, use of multi-robot systems presents additional challenges. Specifically, in domains where the robots' actions are tightly coupled, coordinating multiple robots to achieve cooperative behavior at the group level is difficult. In this paper, we demonstrate that reward shaping can greatly benefit learning in multi-robot exploration tasks. We propose a novel reward framework based on the idea of counterfactuals to tackle the coordination problem in tightly coupled domains. We show that the proposed algorithm provides superior performance (166% performance improvement and a quadruple convergence speed up) compared to policies learned using either the global reward or the difference reward [1].

Authors

Keywords

  • Robot kinematics
  • Robot sensing systems
  • Neural networks
  • Environmental monitoring
  • Multi-robot systems
  • Training
  • Superior Performance
  • Multi-agent Systems
  • Policy Learning
  • Exploration Task
  • Neural Network
  • Learning Algorithms
  • Evaluation Of Function
  • Evolutionary Algorithms
  • Fitness Function
  • Global Rate
  • Feed-forward Network
  • Number Of Agents
  • Multiple Agents
  • Reward Function
  • Learning Agent
  • Reward Signal
  • Simultaneous Observation
  • Swarm Robotics
  • Neural Network Control
  • Tight Coordination
  • Individual Robots

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
362581154570737935
v2026.09.13