RLDM 2017
Time-adaptive temporal difference reinforcement learning
Abstract
Anticipating the timing of rewards is as crucial to adaptive behavior as predicting what those rewards will be. In the brain, reward learning is thought to depend on dopamine signals that convey a prediction error whenever reward predictions do not accord with reality. Research further suggests that neural circuits in the basal ganglia use these error signals to learn reward predictions. This well-articulated behavioral and neural story nevertheless leaves an important question unaddressed: how does the neural machinery in the basal ganglia support learning of exactly when rewards should be expected? Prominent temporal difference reinforcement learning models purport to explain how the timing of rewards is learned, but rely for this on simplifying assumptions that are not tenable in the biological circuits that support reward prediction and learning in the brain. To address this question, we introduce time-adaptive temporal differ- ence reinforcement learning (time-adaptive TDRL), in which both the value and duration of underlying task states are learned concurrently. This model builds on the theory of partially-observable semi-Markov deci- sion processes, and introduces a mechanism for learning the duration of hidden task states by tracking the elapsed time between observations. We use this model to reproduce a number of features of dopaminergic reward prediction error signals to manipulations in the timing of reward delivery in simple learning tasks and interpret these response patterns as the reflection of an inference process that unfolds in time over the true underlying state of the task. In this framework, expectations about the likely time of state transitions are used to gate prediction error signaling, thereby concentrating learning at the time of predicted changes in the underlying state of the task, making testable predictions about the neural computation of dopaminergic prediction errors during learning.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 827559709172629184