Arrow Research search
Back to RLDM

RLDM 2015

Approximate Linear Successor Representation

Conference Abstract Accepted abstract Artificial Intelligence · Decision Making · Machine Learning · Reinforcement Learning

Abstract

The dependency of the value function on the dynamics and a fixed rewards function makes the reuse of information difficult when domains share dynamics but differ in their reward functions. If instead of a value function, a successor representation is learned for some fixed dynamics, then any value function defined on any reward function can be computed efficiently. This setting can be particularly useful for reusing options in a hierarchical planning framework. Unfortunately, even linear parametrization of successor representation require a quadratic number of parameters with respect to the number of features and as many operations per temporal difference update step. We present a simple temporal difference-like algorithm for learning an approximate version of the successor representation with an amortized quadratic runtime with respect to the maximum rank of the approximation. Preliminary results indicate that this parameter can be much smaller than the number of features.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
Multidisciplinary Conference on Reinforcement Learning and Decision Making
Archive span
2013-2025
Indexed papers
1004
Paper id
998388822988013192
v2026.09.13