Arrow Research search

Author name cluster

Nathaniel Daw

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

27 papers
1 author row

Possible papers

27

NeurIPS Conference 2023 Conference Paper

Would I have gotten that reward? Long-term credit assignment by counterfactual contribution analysis

  • Alexander Meulemans
  • Simon Schug
  • Seijin Kobayashi
  • Nathaniel Daw
  • Gregory Wayne

To make reinforcement learning more sample efficient, we need better credit assignment methods that measure an action’s influence on future rewards. Building upon Hindsight Credit Assignment (HCA), we introduce Counterfactual Contribution Analysis (COCOA), a new family of model-based credit assignment algorithms. Our algorithms achieve precise credit assignment by measuring the contribution of actions upon obtaining subsequent rewards, by quantifying a counterfactual query: ‘Would the agent still have reached this reward if it had taken another action? ’. We show that measuring contributions w. r. t. rewarding states, as is done in HCA, results in spurious estimates of contributions, causing HCA to degrade towards the high-variance REINFORCE estimator in many relevant environments. Instead, we measure contributions w. r. t. rewards or learned representations of the rewarding objects, resulting in gradient estimates with lower variance. We run experiments on a suite of problems specifically designed to evaluate long-term credit assignment capabilities. By using dynamic programming, we measure ground-truth policy gradients and show that the improved performance of our new model-based credit assignment methods is due to lower bias and variance compared to HCA and common baselines. Our results demonstrate how modeling action contributions towards rewarding outcomes can be leveraged for credit assignment, opening a new path towards sample-efficient reinforcement learning.

NeurIPS Conference 2022 Conference Paper

Using natural language and program abstractions to instill human inductive biases in machines

  • Sreejan Kumar
  • Carlos G. Correa
  • Ishita Dasgupta
  • Raja Marjieh
  • Michael Y Hu
  • Robert Hawkins
  • Jonathan D Cohen
  • Nathaniel Daw

Strong inductive biases give humans the ability to quickly learn to perform a variety of tasks. Although meta-learning is a method to endow neural networks with useful inductive biases, agents trained by meta-learning may sometimes acquire very different strategies from humans. We show that co-training these agents on predicting representations from natural language task descriptions and programs induced to generate such tasks guides them toward more human-like inductive biases. Human-generated language descriptions and program induction models that add new learned primitives both contain abstract concepts that can compress description length. Co-training on these representations result in more human-like behavior in downstream meta-reinforcement learning agents than less abstract controls (synthetic language descriptions, program induction without learned primitives), suggesting that the abstraction supported by these representations is key.

RLDM Conference 2019 Conference Abstract

Evidence for a cost of cognitive control effect on foraging behavior

  • Laura A Bustamante
  • Allison Burton
  • Nathaniel Daw
  • Jonathan Cohen

Objective: Evidence suggests exerting cognitive control carries an intrinsic cost and that indi- vidual differences in subjective costs may account for differences in everyday control allocation. Previous studies have demonstrated individual differences in the subjective effort associated with engaging control but are limited in that the choices are explicit and may introduce experimenter demand characteristics, or the choice period is separated from the realization of the cognitive effort. We sought to build on this literature using a novel method to quantify individual differences in the cost of cognitive control that addresses these limitations. Methods: We designed a method for quantifying control costs using a patch foraging task in which participants (N=20) had to complete a control-demanding task (N-Back) to travel between patches. We predicted that participants would over-exploit a patch, yielding diminishing rewards, when performance of the more demanding 3-Back task vs. a 1-Back task was required to travel. We applied the Marginal Value Theorem to quantify how costly participants treated the 3-Back task based on their shift of exit threshold. Results: Most participants treated control as costly and exited later in the 3-Back condition. Control costs may be separable from error avoidance as there was no reliable correlation with N-Back task performance. Conclusions: Our results demonstrate that along with time costs, cognitive control registers as a cost in a patch foraging environment. Advantages of this design include that control costs can be measured implicitly and cost is expressed directly in terms of reward (money). Additionally reward and cost learning are expe- riential, and control allocation is an immediate consequence of choice. This measure can be used to explore the extent to which control costs are experienced and utilized in decisions about control differently across individuals.

RLDM Conference 2019 Conference Abstract

Reconciling dopaminergic response heterogeneity with reward prediction er- ror models

  • Nathaniel Daw
  • Ilana Witten
  • Rachel Lee

The reward prediction error model of the midbrain dopamine system has been successful in part because the global, scalar error signal it describes seems well matched to the sweeping, diffuse projections of dopamine neurons and their apparently homogenous phasic responses. The model’s use of a single reward prediction error for both reward- and action-related learning also seems to explain an apparent lack of movement-related responses in the system, a surprising feature of early studies given the neuromodulator’s implication in movement disorders. However, we review recent evidence that now clearly demonstrates that the dopamine response is instead heterogeneous both from target area to target area and even from neuron to neuron. There are now also unambiguous reports of precisely the types of movement-related responses whose earlier apparent absence the model had seemed to explain. We revisit the role of the scalar error signal in temporal-difference learning, and lay out a pair of related computational proposals how the model might accommodate heterogeneous prediction errors. We aim for a maximally simple and generic account, making few assumptions beyond the standard model and changing only its mapping onto the circuitry. Our core insight is that in a realistic biological system the state input to the learning system is continuous and high-dimensional (unlike most previous models). If this input is represented with a distributed feature code, then this population code may be inherited by the prediction error signal for which it serves as both input and target. We show that these interactions between prediction error and a high dimensional state space can explain many of the seemingly anomalous features of the heterogeneous dopamine response.

RLDM Conference 2017 Conference Abstract

A rational model of prioritized experience replay

  • Marcelo Mattar
  • Nathaniel Daw

Psychologists have long argued that animals and humans can use maps or models to plan actions, a process often viewed in RL terms as model-based action evaluation when an uncertain choice is faced. However, recent experiments suggest that many flexible choice phenomena previously considered to support such just-in-time model usage actually depend on computations occurring much earlier, during offline rest or when parts of the world model are first encountered. Analogously, position representations in rodent hippocampus can run ahead of the animal at choice points, suggesting a substrate for forward evaluation. But these events are only one instance of a heterogeneous family of non-local sequences, which include both forward and backward replay, during both behavior and rest. Jointly, these data indicate that theories suggesting that organisms selectively deploy model-based forward evaluation must be generalized to explain offline computation. We propose a rational model that prioritizes individual Bellman backups in a DYNA setting according to their expected utility. This utility is the product of a need term, measuring the number of times the agent will visit a target state, with a gain term, defined as the relative return improvement obtained by the backup operation. The balance between these imperatives during exploration and rest produces a heterogeneous pattern of experience replay that resembles sequential place cell activity in the hippocampus. In particular, encountering a large prediction error such as an unexpected reward drives reverse replay to propagate its gain; at other times, need drives “forward sweeps” ahead of the current state. By viewing forward search as a special case of a more general prioritized evaluation scheme, the model qualitatively accounts for a variety of empirical findings in both the human and rodent literature and suggests that the content of memory access may reflect the rational investment of limited computational resources.

RLDM Conference 2017 Conference Abstract

Cognitive effort and the opportunity cost of time: a behavioral examination

  • Ross Otto
  • Nathaniel Daw

Principles of rational choice may govern an agent’s internal allocation of resources, as well as its overt actions. In many tasks, an organism’s behavior reflects a fundamental tradeoff between accuracy and cognitive effort. Although theorists have proposed that decisions to expend cognitive effort can be conceptualized as rational cost-benefit tradeoffs, few experiments directly test this claim. We sought to 1) quantify expenditure of cognitive effort, and 2) manipulate its costs and benefits to investigate how people solve this effort-accuracy tradeoff. We tested the hypothesis that expenditure of cognitive effort is sensitive to the opportunity cost of time, extending a popular theory about physical effort (Niv et al. 2007). This account holds that energetic costs of vigorous responding trade off against the opportunity cost of time spent by acting more slowly. Because in many settings this cost equals the overall average reward rate, the theory predicts speedier behavior in richer environments. We extend this framework from physical to cognitive effort using established tasks, for which 1) the cognitive effort demanded varies from trial to trial and 2) effort affects performance via measurable speed-accuracy tradeoffs. In one experiment, subjects completed a perceptual decision-making task, while available rewards fluctuated from trial to trial. Response speeds and accuracies tracked the recently experienced average reward rate: when the opportunity cost of time was high, subjects responded more quickly and less accurately. In a second experiment, subjects completed a Simon response conflict task. On incongruent trials—for which correct responses demand cognitive effort— we again observed reward-rate-dependent speeding and a reduction in accuracy. Last, in a task-switching experiment, the average reward rate engendered more errors on (effortful) task switches. Thus, across diverse task domains, expenditure of cognitive effort tracked the opportunity cost of time.

RLDM Conference 2017 Conference Abstract

Episodic Contributions to Model-Based Reinforcement Learning

  • Oliver Vikbladh
  • Nathaniel Daw

RL theories of human and animal behavior often assume that choice relies on incrementally learned running averages of previous events, either action values for model-free (MF) or one-step models for model-based (MB) accounts. However, a third suggestion, supported by recent findings, posits that in- dividual trajectories are also stored as separate episodic memories and can later be retrieved or sampled to guide choice. Such a third way” raises particular questions for classic arguments that animals use a cognitive map or model to plan actions in sequential tasks: Individual trajectories embody the same state-action-state relationships summarized in a world model and might be used to similar end. Conversely, their use might confound standard tests for model use. To investigate the contribution of memories for individual trials in sequential choice, we created a task that combines 2-step MDP dynamics, of the sort previously used to distinguish MB from MF, with single trial memory cues (unique objects) that also predict reward. This allowed us to investigate whether episodic information about a cued object’s previous reward influences MB or MF evaluation, and also how these effects trade off against incrementally learned estimates. 80 human subjects competed 200 trials online. In addition to significant signatures of traditional MF and MB strate- gies based on running averages, subjects displayed a significant capacity for MB planning using individually cued episodes. Furthermore, on trials that contained episodic cues (vs. those that didn’t), traditional (puta- tively incremental) MB planning was significantly reduced. This finding raises the possibility that previous interpretations of choices as reflecting running averages may instead reflect covert retrieval of individual episodes, which are replaced by explicitly cued episodes when these are provided.

RLDM Conference 2017 Conference Abstract

Excessive Deliberation in Social Anxiety

  • Elana Meer
  • Lindsay Hunter
  • Daw Lab
  • Nathaniel Daw

Recent work has argued that mental health disorders can be understood in terms of dysfunction in the brain’s RL mechanisms. Notably, by comparing humans’ psychiatric symptoms to their choices in RL tasks, researchers have linked the compulsive aspect of drugs of abuse, and other disorders such as OCD, to excessive use of automatic (model-free) over deliberative (model-based) action evaluation. There have also been suggestions, mostly theoretical, that a converse pathology of excess deliberation might be linked to other disorders, e. g. rumination in mood disorders. We investigated this hypothesis directly, by assessing how symptoms of social anxiety disorder (SAD) predict model-based (vs model-free) learning in a socially framed RL task. 489 participants from a general population sample (Amazon Mechanical Turk) completed the Liebowitz Social Anxiety Scale (LSAS), and played 80 rounds of a competitive economic game, the Patent Race, against a computerized opponent. SAD is an appealing test both because it is prevalent in the Turk population, and because the focus of the anxiety is pertinent to a socially framed task. Previous research has captured human choices and neural responses on such tasks with the Experience Weighted Attraction (EWA) model. EWA nests two RL strategies, learning action values by a weighted combination of model-free reward sampling, vs. model-based learning (marginalizing) of the opponent’s move distribution. Estimating the parameters of EWA that best fit subjects’ choices, we verified, in accord with our hypothesis, that self-reported social anxiety was selectively associated with increased use of model- based evaluation (P ¡. 01; 6% increase in MB learning per 1 SD increase in LSAS). Other model parameters were unaffected. These results ground the deleterious symptoms of SAD, such as overthinking and paralysis in social interactions, in well-characterized neuro-computational mechanisms, and offer a rare example of enhanced function in disease.

RLDM Conference 2017 Conference Abstract

Mechanisms of Overharvesting in Patch Foraging

  • Gary Kane
  • Aaron Bornstein
  • Amitai Shenhav
  • Robert Wilson
  • Nathaniel Daw
  • Jonathan Cohen

Serial stay-or-search decisions are ubiquitous across many domains, including decisions regard- ing employment, relationships, and foraging for resources or information. Studies of animal foraging, in which animals decide to harvest depleting rewards contained within a patch or to leave the patch in search of a new, full one, have revealed a consistent bias towards overharvesting, or staying in patches longer than is predicted by optimal foraging theory (the Marginal Value Theorem; MVT). Yet, the cognitive biases that lead to overharvesting are poorly understood. We attempt to determine the cognitive biases that underlie overharvesting in rats. We characterized rat foraging behavior in response to two basic manipulations in patch foraging tasks: travel time between reward sources and depletion rate of the source; and to two novel manipulations to the foraging environment: proportional changes to the size of rewards and length of delays, and placement of delays (pre- vs. post-reward). In response to the basic manipulations, rats qualitatively followed predictions of MVT, but stayed in patches longer than is predicted. In the latter two manipulations, rats deviated from predictions of MVT, exhibiting changes in behavior not predicted by MVT. We formally tested whether four separate cognitive biases — subjective costs, decreasing marginal utility for reward, discounting of future reward, and ignoring post-reward delays — could explain overharvesting in the former two manipulations and deviations from MVT in the latter two. All the biases tested explained overharvest- ing behavior in the former contexts, but only one bias — in which rats ignore post-reward delays — also explained deviations from MVT in the latter contexts. Our results reveal that multiple cognitive biases may contribute to overharvesting, but inaccurate estimation of post-reward delays provided the best explanation across all contexts.

RLDM Conference 2017 Conference Abstract

Offline Replay Supports Planning: fMRI Evidence from Reward Revaluation

  • Ida Momennejad
  • Ross Otto
  • Nathaniel Daw
  • Ken Norman

We offer fMRI evidence for the idea of planning as learning from replay. Learning to make ad- vantageous decisions in sequentially structured tasks, like mazes, requires integrating information acquired across multiple learning episodes. This is a challenge for many popular learning approaches that work fully “online”, adjusting representations that summarize ongoing experience. A proposed mechanism to support such challenging integration is to replay real or simulated experiences “offline”. Here we used a decision task called retrospective revaluation in which participants must integrate initial experience about a task with later experience about a change in its goals. We hypothesized that replaying past experience during inter- mittent rest periods helps ‘piece together’ trajectories that were not directly experienced, enabling the inte- gration of new relevant information to update previously learned policies. A key question for this account is how the brain prioritizes whether or which experiences to replay. Based on research in machine learning, we hypothesize that the brain should preferentially replay experiences ‘tagged’ with prediction errors, signaling increased uncertainty that may have consequences for other states and decisions. To test this, we acquired fMRI data as participants performed a sequential decision task with revaluation and control trials. We used multi-voxel pattern analysis (MVPA) to measure replay as classifier evidence for reactivation of past states. We report three main results, n=24. (a) MVPA evidence for replay during rest predicts revaluation during test. (b) Evidence for replay during rest is predicted by frontoparietal sensitivity to prediction errors during learning. (c) Brain’s memory and evaluation networks (hippocampus and medial prefrontal cortex) show higher activation during revaluation vs. control rest. These findings further our understanding of how the brain leverages offline mechanisms in planning and goal-directed behavior.

RLDM Conference 2017 Conference Abstract

The Hippocampus as a Common Neural Substrate for Spatial Navigation and Model-Based

  • Oliver Vikbladh
  • Nathaniel Daw

The hippocampus (HC) famously supports spatial memory. However, little direct evidence cor- roborates the commonly asserted hypothesis that these spatial functions extend to map- or model-based planning, in space or otherwise. We addressed the relationship between these functions by probing both goal-directed planning and spatial memory in patients with damage to the HC following anterior temporal lobectomy (ATL). We hypothesized that if both abilities indeed share a common HC substrate, they should both be attenuated following HC damage but covary with one another the when HC is intact. 19 unilateral ATL patients and 19 controls (matched for age and IQ) participated in the study. Subjects performed a se- quential decision-making task used to differentiate reliance on model-free (MF) and model-based (MB) RL strategies. Subjects also performed a spatial navigation task that distinguished hippocampal dependent place memory”, referenced to environmental boundaries, from striatal dependent “response memory”, referenced to a discrete landmark. As predicted, patients displayed attenuated place memory but not response memory. Patients were also biased away from MB and toward MF strategies on the sequential decision-making task. Comparing the two tasks, place memory performance was correlated with the use of MB planning in the control group. Importantly, no such correlation was found in the patient group. These results show a causal role of the HC in goal-directed MB planning and speak to multiple, potentially competing, decision systems in the human brain. Finally, the demonstration that place memory correlates with MB planning only in the control group suggests that the HC may serve as a common neural substrate for flexible behavior during spatial navigation and sequential decision-making, but that following damage to this structure, differential compensatory mechanisms emerge.

RLDM Conference 2013 Conference Abstract

Acute Stress Effects on Model-Based versus Model-Free Reinforcement Learning

  • Ross Otto
  • Candace Raio
  • Elizabeth Phelps
  • Nathaniel Daw

Contemporary accounts of reinforcement-learning (RL) posit the operation of separate, compet- ing valuation systems in the control of choice behavior: model-free RL, which learns action preferences in a manner in accord with the ”law of effect”, is contrasted with the more flexible model-based RL, which ex- plicitly represents environment structure in order to prospectively evaluate actions. On the basis of previous work demonstrating that working-memory (WM) resources are necessary for implementing model-based choice, and that acute stress deleteriously impacts executive functions including WM, we investigated how acute stress (and its concomitant physiological response) alters expression of model-based and model-free RL in a sequential choice task affording disentanglement of the two choice strategies. To induce a neuro- physiological stress response, we administered the Cold Pressor Test shortly before subjects completed the sequential choice task. In line with our predictions we find that cortisol response attenuates model-based, but not model-free contributions to choice behavior.

RLDM Conference 2013 Conference Abstract

Better things to do: opportunity cost may contribute to cognitive depletion effects

  • Y-Lan Boureau
  • Nathaniel Daw

Influential research suggests that exercising self-control depletes a limited cognitive resource. However, the identity of this putative resource remains unclear. We examined whether the behavior associ- ated with such cognitive depletion could instead be understood in terms of rational learning and choice. The key finding is that performing a control-demanding task, relative to a baseline, reduces performance (eg, by increasing quitting) on a subsequent control-demanding task. We hypothesized that rather than depleting a resource, the first task might affect later behavior in part by informing subjects about the average reward available in the environment. This is a measure of the opportunity cost of time, so that if it is higher, rational subjects should quit earlier or be more rushed on the second task (Charnov 1976; Niv et al. 2007). We examined this hypothesis in experiments in which subjects unscrambled anagrams after different ini- tial tasks. We first replicated the standard result: subjects unscrambled fewer words successfully (and quit earlier) following a demanding constrained writing task, compared to free writing. Consistent with an aver- age reward account, subjects in the demanding condition (and also controls who quit the anagrams earliest) reported higher engagement in the writing task. In a second experiment, we more explicitly manipulated the average reward by replacing the writing task with a simple slot-machine task for which different groups received two levels of monetary payoffs. Anagram performance tracked the reward level in the first phase, with higher rewards associated with earlier quitting. Thus, results from our experiments suggest that opportunity costs may indeed explain some behavioral patterns usually ascribed to the depletion of some resource.

RLDM Conference 2013 Conference Abstract

Dissociable effects of dopamine and serotonin on reversal learning

  • Hanneke den Ouden
  • Nathaniel Daw
  • Guillen Fernandez
  • Joris Elshout
  • Mark Rijpkema
  • Martine Hoogman
  • Barbara Franke
  • Roshan Cools

Serotonin and dopamine are speculated to subserve motivationally opponent functions, but this hypothesis has not been directly tested. We studied the role of these neurotransmitters in probabilistic re- versal learning in nearly 700 individuals as a function of two polymorphisms in the genes encoding the serotonin and dopamine transporters (5HTTLPR plus rs25531; DAT1 3’UTR VNTR). A double dissocia- tion was observed. The SERT polymorphism altered behavioral adaptation after losses (lose-shifting), with increased lose-shift associated with L’-L’ homozygosity, while leaving unaffected perseveration after rever- sal. In contrast, the DAT1 genotype affected the influence of prior choices on perseveration, while leaving lose-shifting unaltered. A model of reinforcement learning captured the dose-dependent effect of DAT1 genotype, such that an increasing number of 9R-alleles resulted in a stronger reliance on previous experi- ence, and therefore reluctance to update learned associations. These data provide direct evidence for doubly dissociable effects of serotonin and dopamine.

RLDM Conference 2013 Conference Abstract

Episodic memory interferes with reward learning and decreases striatal prediction errors

  • G Wimmer
  • Erin Kendall Braun
  • Nathaniel Daw
  • Daphna Shohamy

Learning from experience is central to adaptive decision making. Research on memory systems has demonstrated distinct cognitive and neural systems for learning stimulus-reward associations and for encoding episodes. In even simple experiences, however, these two types of learning often co-occur and may interact. Currently, it is unknown whether and how learning of stimulus-reward associations is influ- enced by memory for learning-related events. Here we sought to address this by examining how incremental reinforcement learning and reward-guided choices are influenced by episodic memory formation for the experience. During the experiment, participants made choices between two options (colored squares), each associated with a drifting probability of reward, with the goal to earn as much money as possible. Incidental, trial- unique object pictures, which were unrelated to the reward learning task, were overlaid on each option. The next day, participants were given a surprise memory test for these pictures. We found that choices were significantly influenced by recent reward experience. Participants also exhibited significant memory for the object pictures that were presented during learning, although the objects were unrelated to the reward learning task. This memory formation interacted with how reward guided choices: both across and within-participants, successful memory formation was associated with a decreased influence of recent reward experience on choice. Neurally, the striatal reward prediction error signal was decreased when memory was successfully formed, and this decrease was preceded by enhanced functional connectivity between the hippocampus and striatum. These results demonstrate a mechanism by which reward-guided choices can be influenced by multiple memory systems. Further, they provide insight into the interactions between neural systems for reward learning and episodic memory.

RLDM Conference 2013 Conference Abstract

How instructed knowledge shapes aversive learning

  • Lauren Atlas
  • Bradley Doll
  • Nathaniel Daw
  • Jian Li

In humans, expectations reflect prior experience and instructed knowledge. Most models of aver- sive learning make predictions about brain responses as a function of reinforcement alone. Recent studies of reward learning indicate that striatal learning is modulated when participants are instructed about stimulus contingencies. The aim of this study was to test whether instructed knowledge modulates associative fear learning. Participants performed a Pavlovian aversive learning paradigm. Two cues were presented: One (the CS+) was paired with shock on 30 % of trials, whereas the second (the CS-) was never paired with shock. Fol- lowing 20 trials, contingencies reversed. There were three reversals across the session. Participants were assigned to two groups: The Instructed Group was informed about contingencies prior to learning and upon each reversal, whereas the Feedback Group received no information. We analyzed skin conductance responses (SCRs) and brain responses to cues. Fear expression tracked con- tingency reversals (i. e. larger SCRs for current CS+ than CS-), and the Instructed Group showed stronger differential responses. The Instructed Group showed greater activation in right DLPFC, while the Feedback Group showed greater activation in bilateral striatum. We fit a quantitative model with a dynamic learning rate to SCRs to isolate the timecourse of learning in the Feedback Group, focusing on prediction error and associability. We then tested whether Instructions modulated the neural correlates of feedback-driven sig- nals. We observed group differences in bilateral ventral striatum, such that only the Feedback Group showed striatal prediction errors. These results reveal that instructed knowledge influences aversive learning. Instructions enhance fear acqui- sition and expression, and prediction errors are not observed when instructions are veridical. The DLPFC is likely to play a key role in maintaining instructions, which in turn modulate fear expression.

RLDM Conference 2013 Conference Abstract

Learned Myopic or Far-Sighted: Experience Shapes Human Temporal Horizon in Sequential

  • Hang Zhang
  • Hyoseok Kim
  • Nathaniel Daw
  • Laurence Maloney

We investigated how well people make sequential decisions to achieve the long-term goal. In video-game-like settings, a spaceship flew across a row of three mountains of increasing heights. Before each mountain, subjects could elevate the spaceship by either a constant and small height (CS) or a variant but on average larger height (VL) to avoid crashing. The goal was to survive beyond the last mountain. The optimal choice before a specific mountain depended on the heights of all future mountains. We tested whether subjects could learn the optimal policy or base their choices only on a short horizon, i. e. on the immediate mountain. Methods: We constructed two combinations of mountain heights, A and B, which differed in how early a short horizon would be penalized. For A, a short horizon would yield the optimal choice before the first mountain and not increase crash rate until the last mountain. In contrast, for B, a short horizon would increase crash rate as early as the second mountain. Each subject completed 4 blocks of 60 trials, in the block order of ABAB or BABA. Sixteen naı̈ve subjects were evenly assigned to the two groups. Results: The two groups differed in their learning trajectories. (1) The ABAB group achieved a higher probability of survival in the last (. 63) than in the first two blocks (. 50), but the BABA did not (both. 54). (2) The ABAB had a shorter horizon than the BABA: When VL was the optimal choice and involved long- term considerations, ABAB chose VL less than BABA did (53 % vs. 67 %). (3) The BABA appeared to be far-sighted: When CS was the optimal choice and reduced crash at the immediate mountain, BABA chose CS less than ABAB did (60 % vs. 81 %). Conclusion: Human individuals’ temporal horizon in a sequential- decision task depends on their initial experience with the task. People may learn to be myopic or far-sighted.

RLDM Conference 2013 Conference Abstract

Learning the value of time

  • Sara Constantino
  • Nathaniel Daw

Although single-shot decision tasks have contributed to our knowledge of value based decision making, they do not address the sequential dependence and temporal aspect of most choices. In this study, we investigate these decisions in a patch-foraging context, which requires subjects to consider future out- comes when deciding how to allocate time between harvesting a depleting resource and searching for a new one. The Marginal Value Theorem (MVT) states that in this class of tasks, the complex problem can be optimally summarized by a threshold rule on the average reward rate or opportunity cost of time. In two experiments, we varied the average reward rate and consistently found that subjects adjusted behavior in the direction predicted by the theory. The MVT threshold rule suggests a simple method for learning the quality of an environment (threshold) by averaging rewards in time and can be contrasted with the predominant incremental learning theory in neuroscience, the Temporal-Difference (TD) algorithm. We examined trial- by-trial decisions and found an effect of recently experienced reward sequences that was better explained by an MVT-based learning model than TD. In a subsequent study, we investigated the suggested role of tonic dopamine (DA) as a signal of average reward rate and opportunity cost of time by looking at the foraging decisions of Parkinson’s disease (PD) patients. In particular, we examined the effect of DA depletion (due to PD) and replacement (due to medication) on the subjective opportunity cost of time implied by the for- aging choices and found that it was consistent with a tonic DA signaled average reward rate. These studies unify work in ecology, economics and computational neuroscience that look at time-sensitive, sequential decisions. We expand upon previous work by suggesting that humans may, in certain contexts, implement a simple threshold-based decision rule on the average reward rate and that this quantity may be signaled by tonic DA.

RLDM Conference 2013 Conference Abstract

Model based and model free reinforcement learning in the brain

  • Nathaniel Daw

There is a long line of research in neuroscience suggesting that the brain’s dopamine systems implement model-free reinforcement learning by temporal-difference methods. However, this theory cannot explain a range of well-demonstrated animal and human behaviors in learning tasks. For this reason, we have more recently suggested that the brain also implements model-based reinforcement learning, as a relatively sepa- rate, competing behavioral control system. I discuss how (and why) these two approaches might be traded off in different circumstances, and how we can dissociate their unique contributions to choices and neural signals in humans learning Markov decision games in the laboratory. Finally, I consider how this theory formalizes a long line of fuzzier dual-systems theories in psychology, and may shed light on disorders of compulsion such as drug abuse. Peter Smittenaar*, Thomas FitzGerald, George Prichard, Vincenzo Romei, Nicholas Wright, Joern

RLDM Conference 2013 Conference Abstract

Neural correlates of forward planning in model-based reinforcement learning

  • Bradley Doll
  • Katherine Duncan
  • Dylan Simon
  • Daphna Shohamy
  • Nathaniel Daw

Incremental learning across species is well described by reinforcement learning (RL) algorithms. The bulk of such demonstrations correlate behavioral and biological signals with signatures of model-free RL. More recently, interest has grown in correlates of model-based RL which can exhibit more cognitive flexibility than model-free RL, though at greater computational cost. Using fMRI, we investigated the neural correlates of learning in a task that dissociates model-based from model-free RL, and permits distinct model- based strategies. Participants navigated to terminal task states from different starting states in search of monetary reward. Model-based behavior in the task may arise from the forward planning typical of these algorithms or from a computational shortcut whereby the representation of actions that produce the same outcomes are joined. The states in this task were represented by different classes of stimuli that activate unique regions of visual cortex. This task feature permitted us to assess RL strategies by decoding brain activations at different task states. Across the population, choice behavior showed evidence of both model-based and model-free learning. Pre- liminary fMRI results permitted closer investigation of the mechanisms by which these classes of learning algorithms are implemented. For each subject, we identified ROIs that showed preferential responses to the stimulus categories used to represent the different task states in an independent functional localizer. We then assessed the activity in these ROIs during the reward learning task start states. We looked for activation of states to be navigated to, as well as representational compression of equivalent start states. Activation related to the former correlated with model-based task behavior, consistent with forward planning in model-based RL.

RLDM Conference 2013 Conference Abstract

Reward-guided decisions are affected by episodic cues

  • Aaron Bornstein
  • Mel Khaw
  • Nathaniel Daw

In traditional reinforcement learning (RL) models, decisions for reward depend on a running average estimate of action values. A different way of approaching decisions is to estimate the value of actions online, at the time of choice, by drawing on discrete episodes of past experience with those actions. In this procedure, the episodes are called ”samples’, and the process is called ”decisions by sampling’ (Erev et al. , 2008b). It has been suggested that sampling models may provide a mechanistic explanation for many idiosyncratic choice behaviors that are not captured by RL (Stewart et al. , 2006; Erev et al. , 2008a). Here, building on the premise that these sampled experiences are encoded by the episodic memory system, we exploited a feature of episodic memories — the ability to bring to mind past contexts using associative cues — to privilege the sampling of particular trial episodes. We show that these episodic cues — from choices, on average, 40 trials past — have immediate and specific effect on subsequent choices. Quanti- tatively, the rewards experienced on these cued trials impact choices about as much as rewards obtained through direct experience just 2 trials in the past. These results are consistent with sampling models of de- cisions, and suggest that manipulations that alter retrieval of episodic memories — such as the cues used here — can also alter choices. The effect may have particular impact the study of choices in natural settings, as episodic information cued by various environmental factors may bias decisions in ways not captured by standard mechanisms.

NeurIPS Conference 2011 Conference Paper

Environmental statistics and the trade-off between model-based and TD learning in humans

  • Dylan Simon
  • Nathaniel Daw

There is much evidence that humans and other animals utilize a combination of model-based and model-free RL methods. Although it has been proposed that these systems may dominate according to their relative statistical efficiency in different circumstances, there is little specific evidence -- especially in humans -- as to the details of this trade-off. Accordingly, we examine the relative performance of different RL approaches under situations in which the statistics of reward are differentially noisy and volatile. Using theory and simulation, we show that model-free TD learning is relatively most disadvantaged in cases of high volatility and low noise. We present data from a decision-making experiment manipulating these parameters, showing that humans shift learning strategies in accord with these predictions. The statistical circumstances favoring model-based RL are also those that promote a high learning rate, which helps explain why, in psychology, the distinction between these strategies is traditionally conceived in terms of rule-based vs. incremental learning.

NeurIPS Conference 2007 Conference Paper

The rat as particle filter

  • Aaron Courville
  • Nathaniel Daw

Although theorists have interpreted classical conditioning as a laboratory model of Bayesian belief updating, a recent reanalysis showed that the key features that theoretical models capture about learning are artifacts of averaging over subjects. Rather than learning smoothly to asymptote (reflecting, according to Bayesian models, the gradual tradeoff from prior to posterior as data accumulate), subjects learn suddenly and their predictions fluctuate perpetually. We suggest that abrupt and unstable learning can be modeled by assuming subjects are conducting in- ference using sequential Monte Carlo sampling with a small number of samples — one, in our simulations. Ensemble behavior resembles exact Bayesian models since, as in particle filters, it averages over many samples. Further, the model is capable of exhibiting sophisticated behaviors like retrospective revaluation at the ensemble level, even given minimally sophisticated individuals that do not track uncertainty in their beliefs over trials.

NeurIPS Conference 2005 Conference Paper

How fast to work: Response vigor, motivation and tonic dopamine

  • Yael Niv
  • Nathaniel Daw
  • Peter Dayan

Reinforcement learning models have long promised to unify computa- tional, psychological and neural accounts of appetitively conditioned be- havior. However, the bulk of data on animal conditioning comes from free-operant experiments measuring how fast animals will work for rein- forcement. Existing reinforcement learning (RL) models are silent about these tasks, because they lack any notion of vigor. They thus fail to ad- dress the simple observation that hungrier animals will work harder for food, as well as stranger facts such as their sometimes greater produc- tivity even when working for irrelevant outcomes such as water. Here, we develop an RL framework for free-operant behavior, suggesting that subjects choose how vigorously to perform selected actions by optimally balancing the costs and benefits of quick responding. Motivational states such as hunger shift these factors, skewing the tradeoff. This accounts normatively for the effects of motivation on response rates, as well as many other classic findings. Finally, we suggest that tonic levels of dopamine may be involved in the computation linking motivational state to optimal responding, thereby explaining the complex vigor-related ef- fects of pharmacological manipulation of dopamine.

NeurIPS Conference 2004 Conference Paper

Similarity and Discrimination in Classical Conditioning: A Latent Variable Account

  • Aaron Courville
  • Nathaniel Daw
  • David Touretzky

We propose a probabilistic, generative account of configural learning phenomena in classical conditioning. Configural learning experiments probe how animals discriminate and generalize between patterns of si- multaneously presented stimuli (such as tones and lights) that are dif- ferentially predictive of reinforcement. Previous models of these issues have been successful more on a phenomenological than an explanatory level: they reproduce experimental findings but, lacking formal founda- tions, provide scant basis for understanding why animals behave as they do. We present a theory that clarifies seemingly arbitrary aspects of pre- vious models while also capturing a broader set of data. Key patterns of data, e. g. concerning animals' readiness to distinguish patterns with varying degrees of overlap, are shown to follow from statistical inference.

NeurIPS Conference 2003 Conference Paper

Model Uncertainty in Classical Conditioning

  • Aaron Courville
  • Geoffrey Gordon
  • David Touretzky
  • Nathaniel Daw

We develop a framework based on Bayesian model averaging to explain how animals cope with uncertainty about contingencies in classical con- ditioning experiments. Traditional accounts of conditioning fit parame- ters within a fixed generative model of reinforcer delivery; uncertainty over the model structure is not considered. We apply the theory to ex- plain the puzzling relationship between second-order conditioning and conditioned inhibition, two similar conditioning regimes that nonethe- less result in strongly divergent behavioral outcomes. According to the theory, second-order conditioning results when limited experience leads animals to prefer a simpler world model that produces spurious corre- lations; conditioned inhibition results when a more complex model is justified by additional experience.

NeurIPS Conference 2002 Conference Paper

Timing and Partial Observability in the Dopamine System

  • Nathaniel Daw
  • Aaron Courville
  • David Touretzky

According to a series of influential models, dopamine (DA) neurons sig- nal reward prediction error using a temporal-difference (TD) algorithm. We address a problem not convincingly solved in these accounts: how to maintain a representation of cues that predict delayed consequences. Our new model uses a TD rule grounded in partially observable semi-Markov processes, a formalism that captures two largely neglected features of DA experiments: hidden state and temporal variability. Previous models pre- dicted rewards using a tapped delay line representation of sensory inputs; we replace this with a more active process of inference about the under- lying state of the world. The DA system can then learn to map these inferred states to reward predictions using TD. The new model can ex- plain previously vexing data on the responses of DA neurons in the face of temporal variability. By combining statistical model-based learning with a physiologically grounded TD theory, it also brings into contact with physiology some insights about behavior that had previously been confined to more abstract psychological models.

v2026.09.13