RLDM Conference 2019 Conference Abstract
Episodic Memory Contributions to Model-Based Reinforcement Learning
- Oliver Vikbladh
- Daphna Shohamy
- Nathaniel D. Daw
RL theories of biological behavior often assume that choice relies on incrementally learned running averages of previous events, either action values for model-free (MF) or one-step models for model- based (MB) accounts. A third suggestion, supported by recent findings, posits that individual trajectories are also stored as episodic memories and can later be sampled to guide choice. This raises questions for classic arguments that animals use a world model to plan actions in sequential tasks: Individual trajectories embody the same state-action-state relationships summarized in a model and might be used to similar end. To investigate the contribution of episodic memories of trial-specific outcomes during sequential choice, we scanned 30 subjects with fMRI while they performed a task that combines 2-step MDP dynamics, of the sort previously used to distinguish MB from MF, with single trial memory cues (trial unique stimuli associated with a specific reward outcomes). This allowed us to investigate whether episodic memory about the cued stimulus influences choice, how reliance of episodic cues trades off against reliance on estimates which are usually understood to be learned incrementally, and test the hypothesis whether this episodic sampling process might specifically underpin putatively incremental MB (but not MF) learning. We find behavioral evidence for an episodic choice strategy. We also show behavioral competition between this episodic strategy and the putatively incremental MB (but not MF) strategies, suggesting some shared (and other results show, item category specific) substrate. Furthermore, we demonstrate fMRI evidence that incremental MB learning shares a common hippocampal substrate with the episodic strategy, for outcome encoding. Taken together, these data provide evidence for convert episodic retrieval, rather than incremental model learning, as a possible implementation of seemingly MB choice in the human brain.