Arrow Research search
Back to EWRL

EWRL 2023

Acceleration in Policy Optimization

Workshop Paper EWRL16 Artificial Intelligence · Machine Learning · Reinforcement Learning

Abstract

We work towards a unifying paradigm for accelerating policy optimization methods in reinforcement learning (RL) through predictive and adaptive directions of (functional) policy ascent. Leveraging the connection between policy iteration and policy gradient methods, we view policy optimization algorithms as iteratively solving a sequence of surrogate objectives, local lower bounds on the original objective. We define optimism as predictive modelling of the future behavior of a policy, and hindsight adaptation as taking immediate and anticipatory corrective actions to mitigate accumulating errors from overshooting predictions or delayed responses to change. We use this shared lens to jointly express other well-known algorithms, including model-based policy improvement based on forward search, and optimistic meta-learning algorithms. We show connections with Anderson acceleration, Nesterov's accelerated gradient, extra-gradient methods, and linear extrapolation in the update rule. We analyze properties of the formulation, design an optimistic policy gradient algorithm, adaptive via meta-gradient learning, and empirically highlight several design choices pertaining to acceleration, in an illustrative task.

Authors

Keywords

  • acceleration
  • Actor-Critic
  • adaptivity
  • extragradient
  • inexact policy gradients
  • meta-gradients
  • meta-learning
  • momentum
  • optimism
  • policy optimization
  • Reinforcement Learning

Context

Venue
European Workshop on Reinforcement Learning
Archive span
2008-2025
Indexed papers
649
Paper id
165580686375247702
v2026.09.13