Arrow Research search
Back to EWRL

EWRL 2018

Safely Exploring Policy Gradient

Workshop Paper Accepted Paper Artificial Intelligence · Machine Learning · Reinforcement Learning

Abstract

Exploration is a fundamental aspect of reinforcement learning (RL) agents. In real-world applications, such as the control of industrial processes, exploratory behavior may have immediate costs that are only repaid in the long run. Selecting the right amount of exploration is a challenging task that can significantly improve overall performance. Safe RL algorithms focus on limiting the immediate costs. This conservative approach, however, carries the risk of sacrificing too much in terms of learning speed and exploration. Starting from the idea that a practical algorithm should be as safe as needed, but not more, we identify an interesting safety scenario and propose Safely-Exploring Policy Gradient (SEPG) to solve it. To do this, we generalize the existing bounds on performance improvement for Gaussian policies to the case of adaptive variance and propose policy updates that are both safe and exploratory. We evaluate our algorithm on simulated continuous control tasks.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
European Workshop on Reinforcement Learning
Archive span
2008-2025
Indexed papers
649
Paper id
548487324086176410
v2026.09.13