RLDM 2017
Episodic Contributions to Model-Based Reinforcement Learning
Abstract
RL theories of human and animal behavior often assume that choice relies on incrementally learned running averages of previous events, either action values for model-free (MF) or one-step models for model-based (MB) accounts. However, a third suggestion, supported by recent findings, posits that in- dividual trajectories are also stored as separate episodic memories and can later be retrieved or sampled to guide choice. Such a third way” raises particular questions for classic arguments that animals use a cognitive map or model to plan actions in sequential tasks: Individual trajectories embody the same state-action-state relationships summarized in a world model and might be used to similar end. Conversely, their use might confound standard tests for model use. To investigate the contribution of memories for individual trials in sequential choice, we created a task that combines 2-step MDP dynamics, of the sort previously used to distinguish MB from MF, with single trial memory cues (unique objects) that also predict reward. This allowed us to investigate whether episodic information about a cued object’s previous reward influences MB or MF evaluation, and also how these effects trade off against incrementally learned estimates. 80 human subjects competed 200 trials online. In addition to significant signatures of traditional MF and MB strate- gies based on running averages, subjects displayed a significant capacity for MB planning using individually cued episodes. Furthermore, on trials that contained episodic cues (vs. those that didn’t), traditional (puta- tively incremental) MB planning was significantly reduced. This finding raises the possibility that previous interpretations of choices as reflecting running averages may instead reflect covert retrieval of individual episodes, which are replaced by explicitly cued episodes when these are provided.
Authors
Keywords
No keywords are indexed for this paper.
Context
- Venue
- Multidisciplinary Conference on Reinforcement Learning and Decision Making
- Archive span
- 2013-2025
- Indexed papers
- 1004
- Paper id
- 555270428175451915