Arrow Research search
Back to RLDM

RLDM 2013

From offline to online reinforcement learning using kernel-based stochastic factorization

Conference Abstract Accepted abstract Artificial Intelligence · Decision Making · Machine Learning · Reinforcement Learning

Abstract

(Joint work with André M. S. Barreto and Doina Precup.) Recent years have witnessed the emergence of several reinforcement-learning techniques that make it possi- ble to learn a decision policy from a batch of sample transitions. Among them, kernel-based reinforcement learning (KBRL) stands out for two reasons. First, unlike other approximation schemes, KBRL always con- verges to a unique solution. Second, KBRL is consistent in the statistical sense, meaning that adding more data improves the quality of the resulting policy and eventually leads to optimal performance. Despite its nice theoretical properties, KBRL has not been widely adopted by the reinforcement learning communi- ty. One possible explanation for this is that the size of the KBRL approximator grows with the number of sample transitions, which makes the approach impractical for large problems. In this work, we introduce a novel algorithm to improve the scalability of KBRL. We use a special de- composition of a transition matrix, called stochastic factorization, which allows us to fix the size of the approximator while at the same time incorporating all the information contained in the data. We apply this technique to compress the size of KBRL-derived models to a fixed dimension. This approach is not only advantageous because of the model-size reduction; it also allows a better bias-variance trade-off, by incor- porating more samples in the model estimate. The resulting algorithm, kernel-based stochastic factorization (KBSF), is much faster than KBRL, yet still converges to a unique solution. We derive a theoretical bound on the distance between KBRL’s solution and KBSF’s solution. We show that it is also possible to construct the KBSF solution in a fully incremental way, thus freeing the space complexity of the approach from its dependence on the number of sample transitions. The incremental version of KBSF (iKBSF) is able to pro- cess an arbitrary amount of data, which results in a model-based reinforcement learning algorithm that can be used to solve large continuous MDPs in on-line regimes. We present experiments on a variety of challenging RL domains, including the double and triple pole- balancing tasks, the Helicopter domain, the penthatlon event featured in the Reinforcement Learning Com- petition 2013, and a model of epileptic rat brains in which the goal is to learn a neurostimulation policy to suppress the occurrence of seizures.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Multidisciplinary Conference on Reinforcement Learning and Decision Making
Archive span
2013-2025
Indexed papers
1004
Paper id
170401914143647283
v2026.09.13