PRL 2021
Scalable Risk-Sensitive Planning by Gradient Descent
Abstract
Planning provides a framework for optimizing sequential decisions in potentially complex environments. A recent advance in efficient planning in deterministic high-dimensional domains with continuous action spaces leverages backpropagation through a model of the environment to directly optimize the actions. However, this method does not take risk into account when optimizing decisions in highly stochastic environments. We address this problem by introducing RiskAware Planning using PyTorch (RAPTOR), a framework that handles risk in stochastic planning domains through an endto-end optimization of entropic utility. While we cannot directly formalize the distributionally-defined entropic utility in closed-form for end-to-end planning, in settings where all MDP stochasticity is defined through the location-scale family, we can reparameterize the objective and apply stochastic backpropagation. What is notable in this approach is that the entropic utility is defined based on sufficient statistics computed from forward sampled trajectories, but due to the nature of autodifferentiation, we can still backpropagate through the entropic utility and these sufficient statistics. The resulting sequence of actions, which we call the risk-sensitive straightline plan, provides a lower bound on the utility of the optimal policy and can be seen as a form of hindsight optimization. We evaluate RAPTOR on two highly stochastic domains, including nonlinear navigation and linear reservoir control, demonstrating the ability to manage risk in complex MDPs.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Bridging the Gap Between AI Planning and Reinforcement Learning
- Archive span
- 2020-2025
- Indexed papers
- 151
- Paper id
- 191253188333317540