EWRL 2018
Safely Exploring Policy Gradient
Abstract
Exploration is a fundamental aspect of reinforcement learning (RL) agents. In real-world applications, such as the control of industrial processes, exploratory behavior may have immediate costs that are only repaid in the long run. Selecting the right amount of exploration is a challenging task that can significantly improve overall performance. Safe RL algorithms focus on limiting the immediate costs. This conservative approach, however, carries the risk of sacrificing too much in terms of learning speed and exploration. Starting from the idea that a practical algorithm should be as safe as needed, but not more, we identify an interesting safety scenario and propose Safely-Exploring Policy Gradient (SEPG) to solve it. To do this, we generalize the existing bounds on performance improvement for Gaussian policies to the case of adaptive variance and propose policy updates that are both safe and exploratory. We evaluate our algorithm on simulated continuous control tasks.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- European Workshop on Reinforcement Learning
- Archive span
- 2008-2025
- Indexed papers
- 649
- Paper id
- 548487324086176410