Arrow Research search
Back to IROS

IROS 2025

Robust Reinforcement Learning based on Momentum Adversarial Training

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

Reinforcement learning (RL) is a fundamental and pivotal algorithm in the advancement of autonomous intelligence, including Embodied Intelligence and Physical Intelligence. The performance of RL directly influences the quality and efficiency of a robot’s decision-making and execution during interactions with its environment. Moreover, the robustness of RL remains a critical challenge that needs to be addressed. A promising approach to enhancing robustness is adversarial reinforcement learning. However, the existing methods primarily focus on perturbations in the state space, while perturbations in the action space have been relatively underexplored. The action space in RL is as crucial as the state space in autonomous intelligence. Furthermore, action-space perturbations provide a more comprehensive evaluation of RL robustness. Therefore, it is necessary and valuable to investigate RL robustness under action-space perturbations for the development of autonomous intelligence. To this end, we propose an adversarial learning framework that employs momentum-based gradient descent to model perturbations in the action space, such as actuator disturbances. Furthermore, we introduce an improved optimization method that integrates historical gradient information into conventional Stochastic Gradient Descent (SGD). This approach enhances training stability and improves perturbation efficiency. The proposed method is evaluated through simulations in the MuJoCo environment and UAV control experiments in GymFC, demonstrating significant improvements in robustness and adaptability under action-space perturbations. Additionally, real-world UAV flight tests are conducted to further validate the effectiveness of the proposed framework. The results confirm that the Sim-to-Real transfer is successful, providing empirical evidence for the applicability of our method in real-world scenarios. This study shows that enhancing RL robustness through action-space perturbations is feasible and effective. More importantly, our findings contribute to the future development of autonomous intelligence, particularly in improving its resilience to uncertainties and dynamic environments.

Authors

Keywords

  • Training
  • Adaptation models
  • Uncertainty
  • Perturbation methods
  • Stochastic processes
  • Autonomous aerial vehicles
  • Robustness
  • Stability analysis
  • Resilience
  • Quadrotors
  • Adversarial Training
  • Gradient Descent
  • State Space
  • Comprehensive Evaluation
  • Dynamic Environment
  • Stochastic Gradient Descent
  • Generative Adversarial Networks
  • Robust Improvement
  • Real-world Test
  • Flight Test
  • Real-world Applications
  • Local Optimum
  • Saddle Point
  • Proportional-integral-derivative
  • Reward Function
  • Robot Control
  • Deep Reinforcement Learning
  • Markov Decision Process
  • Policy Learning
  • Gradient-based Methods
  • Adversarial Perturbations
  • Projected Gradient Descent
  • Pitch Axis
  • Proximal Policy Optimization
  • Reinforcement Learning Policy
  • Noisy Conditions
  • Roll Axis
  • Flight Control
  • Yaw Axis
  • Momentum Term

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
431601085767698273
v2026.09.13