RLDM 2015
A learning mechanism for variability-sensitive reinforcement learning
Abstract
Variability in reward outcome is known to influence motivated behavior in humans and animals. While this sensitivity to so-called risk is a well-established behavioral phenomenon, the neural mechanisms that underlie its action are not well understood. We propose a model of reinforcement learning in the stria- tum that is sensitive to both the average of rewards and their variability, thereby outlining a putative neural mechanism for the influence of risk on learning and decision-making. Current theories of reinforcement learning in the basal ganglia propose a central role for dopamine in signaling errors in the prediction of reward, and hypothesize a central role for dopamine-mediated plasticity in the striatum in learning the as- sociation between states of the environment and the average future rewards they predict. We extend such a model of striatal reinforcement learning by introducing a parallel learning circuit that monitors ongoing dopaminergic prediction errors as a proxy for variability in reward outcomes around their mean. The spe- cific pattern of risk learnt from probabilistic rewards in the environment is dictated by nonlinearities in the response of the variability learning system and the step-size of its update rule. Coupling between the vari- ability learning system and the primary average reinforcement learning circuit allows learnt risk to affect the iterative update of state value, driving differentiation between states of equal expected future reward according to the weighting on their variability. This model demonstrates how parallel update systems tied to the same dopaminergically-mediated prediction error signal can interact locally in a neural circuit to pro- duce adaptive learning based on the experienced variability of rewards. We discuss the striatal cholinergic system as a putative neural substrate of the variability learning system and consider its potential role in the modulation of reinforcement learning in the striatum.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 1009404845291647772