RLDM 2013
“Identity prediction errors” and model-based learning
Abstract
It is known that humans and animals perform model-based reinforcement learning, in which decision-making uses a full model of the environment, including the transition probabilities between states. This is in contrast to model-free reinforcement learning, which relies on a “value” for each state or state- action pair. In model-free reinforcement learning, state values are thought to be learned using “value predic- tion errors” the difference between the expected value and the actual value observed. This is a computational- ly efficient means of learning, and, furthermore, these predictions errors famously seem to be represented by the activity of midbrain dopamine neurons. How are the full models, needed for model based reinforcement learning, learned? It has been hypothesized that there exist analogous “identity (state) prediction errors” in the brain, which are used to learn the transition probabilities between states. Identity prediction errors are elicited when a specific outcome (or state) is unexpected, regardless of whether the utility value of the outcome is different from what was expected. Here, we use fMRI to find signals corresponding to identity prediction errors in the brain. Based on prior work, we expected to find such signals in the orbitofrontal cortex.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 472309507001326906