Arrow Research search
Back to AAMAS

AAMAS 2026

Accelerating Action-Robust Deep Deterministic Policy Gradient via Parallel Optimization

Conference Paper Extended Abstracts Autonomous Agents and Multiagent Systems

Abstract

The Noisy Action Robust MDP (NR-MDP) framework is effective for training robust policies but suffers from significant computational inefficiency due to the iterative, alternating training of a protagonist and an adversary. While reducing the frequency of adversary updates can shorten training time, it typically results in a degradation of robust performance. To address this trade-off, we propose Efficient Action Robust DDPG (EAR-DDPG), which eliminates the need for explicit adversarial policy learning by generating adversarial actions on-the-fly. We show that the approximately optimal adversarial action lies at a vertex of the hypercube. This insight allows us to replace conventional iterative optimization with instantaneous parallel computation over these vertex candidates.

Authors

Keywords

  • Robust reinforcement learning
  • Adversarial learning
  • Parallel computing

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
767552196116744254
v2026.09.13