Arrow Research search
Back to IROS

IROS 2025

Bridging the Reality Gap: Communication-Aware Task Allocation with Multi-Objective Asynchronous Policy Learning

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Distributed task allocation in the UAV swarm is sensitive to excessive communication overhead and frequent transmissions. Combining reinforcement learning and task allocation demonstrates great potential in enhancing algorithm performance and optimizing communication. However, existing studies rely on ideal communication assumptions and the nonphysical environment, making training and validation impractical in applying networked swarms. This paper proposes the Communication-Aware Task Allocation, which aims to train a gating mechanism policy to coordinate the transmission timing, improving robustness and timelessness of the task allocation. First, the policy learning problem is formalized as a POMDP, for which the channel access and other features are designed for observations, actions are inter-agent adaptive gating mechanisms, and the shared reward reflects global task conflicts. Second, to address the asynchronous learning under the CTDE, an asynchronous experience collection and splicing method is proposed to align trajectories. Then, the MOCPPO is proposed, which combines a primal-dual operator with proximal policy optimization, updating the optimal Lagrange multiplier and strategy parameters to simultaneously minimize task conflicts and communication overhead. Finally, sim-to-real experiments are conducted in the HIL environment, and results illustrate the best trade-off optimization of the proposed method over all state-of-the-art approaches.

Authors

Keywords

  • Training
  • Splicing
  • Scalability
  • Bandwidth
  • Autonomous aerial vehicles
  • Robustness
  • Trajectory
  • Timing
  • Resource management
  • Optimization
  • Policy Learning
  • Task Allocation
  • Asynchronous Learning
  • Lagrange Multiplier
  • Unmanned Aerial Vehicles
  • Optimal Policy
  • Gating Mechanism
  • Communication Overhead
  • Channel Access
  • Task Conflict
  • Proximal Policy Optimization
  • Time Step
  • Optimization Algorithm
  • Global Status
  • Communication Strategies
  • Kullback-Leibler
  • Multi-objective Optimization
  • Time Slot
  • Training Center
  • Reward Function
  • Multi-agent Reinforcement Learning
  • Virtual Nodes
  • State Transition Function
  • Convergence Time
  • Asynchronous Condition
  • Lagrangian Method
  • Self-organizing Network
  • Real-world Environments
  • Critic Network
  • Markov Decision Process

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
673508713878086151
v2026.09.13