Arrow Research search
Back to IROS

IROS 2023

Energy Constrained Multi-Agent Reinforcement Learning for Coverage Path Planning

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

For multi-agent area coverage path planning problem, existing researches regard it as a combination of Traveling Salesman Problem (TSP) and Coverage Path Planning (CPP). However, these approaches have disadvantages of poor observation ability in online phase and high computational cost in offline phase, making it difficult to be applied to energy-constrained Unmanned Aerial Vehicles (UAVs) and adjust strategy dynamically. In this paper, we decompose the task into two sub-problems: multi-agent path planning and sub-region CPP. We model the multi-agent path planning problem as a Collective Markov Decision Process (C-MDP), and design an Energy Constrained Multi-Agent Reinforcement Learning (ECMARL) algorithm based on the centralized training and distributed execution concept. Taking into account energy constraint of UAVs, the UAV propulsion power model is established to measure the energy consumption of UAVs, and load balancing strategy is applied to dynamically allocate target areas for each UAV. If the UAV is under energy-depleted situation, ECMARL can adjust the mission strategy in real time according to environmental information and energy storage conditions of other UAVs. When UAVs reach each sub-region of interest, Back-an-Forth Paths (BFPs) are adopted to solve CPP problem, which can ensure full coverage, optimality and complexity of the sub-problem. Comprehensive theoretical analysis and experiments demonstrate that ECMARL is superior to the traditional offline TSP-CPP strategy in terms of solution quality and computational time, and can effectively deal with the energy-constrained UAVs.

Authors

Keywords

  • Training
  • Power measurement
  • Reinforcement learning
  • Traveling salesman problems
  • Propulsion
  • Markov processes
  • Autonomous aerial vehicles
  • Path Planning
  • Multi-agent Reinforcement Learning
  • Coverage Path
  • Coverage Path Planning
  • Energy Consumption
  • Computation Time
  • Target Area
  • Unmanned Aerial Vehicles
  • Markov Decision Process
  • Reinforcement Learning Algorithm
  • Energy Constraints
  • Traveling Salesman Problem
  • Online Phase
  • Deep Learning
  • Field Of View
  • Time Step
  • State Space
  • Environment Interactions
  • Constant Speed
  • Number Of Agents
  • Deep Reinforcement Learning
  • Deep Q-learning
  • Q-learning
  • State-value Function
  • Polygon Area
  • Multi-agent Systems
  • Unmanned Aerial Vehicle Position
  • Convex Polygon
  • Flight Path
  • Grid Map

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
398149328223709435
v2026.09.13