EWRL 2018
TD-Regularized Actor-Critic Methods
Abstract
Actor-critic methods can achieve incredible performance on difficult reinforcement-learning problems, but they are also prone to instability due to the interplay between the actor and critic during learning. To improve their stability, we propose a novel TD-regularized actorcritic method. Our method regularizes the learning objective of the actor by penalizing the temporal difference error of the critic. This improves stability by avoiding overconfident steps in the actor update when the critic is highly inaccurate. We show that our TD-regularization can be easily applied to existing actor-critic methods, e.g., deterministic policy gradient and trust-region policy optimization, with only a slight increase in computation. Evaluations on standard benchmarks show that our method improves stability and exhibits better performance and data-efficiency than its non-regularized counterparts.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- European Workshop on Reinforcement Learning
- Archive span
- 2008-2025
- Indexed papers
- 649
- Paper id
- 902628855155112216