Arrow Research search

Author name cluster

Yael Niv

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

29 papers
2 author rows

Possible papers

29

RLJ Journal 2024 Journal Article

States as goal-directed concepts: an epistemic approach to state-representation learning

  • Nadav Amir
  • Yael Niv
  • Angela J Langdon

Goals fundamentally shape how we experience the world. For example, when we are hungry, we tend to view objects in our environment according to whether or not they are edible (or tasty). Alternatively, when we are cold, we view the very same objects according to their ability to produce heat. Computational theories of learning in cognitive systems, such as reinforcement learning, use state-representations to describe how agents determine behaviorally-relevant features of their environment. However, these approaches typically assume ground-truth state representations that are known to the agent, and reward functions that need to be learned. Here we suggest an alternative approach in which state-representations are not assumed veridical, or even pre-defined, but rather emerge from the agent's goals through interaction with its environment. We illustrate this novel perspective using a rodent odor-guided choice task and discuss its potential role in developing a unified theory of experience based learning in natural and artificial agents.

RLC Conference 2024 Conference Paper

States as goal-directed concepts: an epistemic approach to state-representation learning

  • Nadav Amir
  • Yael Niv
  • Angela J Langdon

Goals fundamentally shape how we experience the world. For example, when we are hungry, we tend to view objects in our environment according to whether or not they are edible (or tasty). Alternatively, when we are cold, we view the very same objects according to their ability to produce heat. Computational theories of learning in cognitive systems, such as reinforcement learning, use state-representations to describe how agents determine behaviorally-relevant features of their environment. However, these approaches typically assume ground-truth state representations that are known to the agent, and reward functions that need to be learned. Here we suggest an alternative approach in which state-representations are not assumed veridical, or even pre-defined, but rather emerge from the agent's goals through interaction with its environment. We illustrate this novel perspective using a rodent odor-guided choice task and discuss its potential role in developing a unified theory of experience based learning in natural and artificial agents.

RLDM Conference 2019 Conference Abstract

Not smart enough: most rats fail to learn a parsimonious task representation

  • Mingyu Song
  • Angela Langdon
  • Yuji Takahashi
  • Yael Niv

As humans designing tasks for laboratory animals, we often assume (or presume) that animals will represent the task as we understand it. This assumption may be wrong. In the worst case, ignoring discrepancies between the way we and our experimental subjects represent an experimental task can lead to data analysis that is meaningless. On the positive side, exploring these discrepancies can shed light on how animals (and humans) learn implicitly, without instructions, task representations that are aligned with the true rules or structure of the environment. Here, we explore how rats represent a moderately complex odor-guided choice task in which different trial types share the same underlying reward structure. Acquiring this shared representation is not necessary for performing the task, but can help the animal learn faster and earn more rewards in a shorter period of time. By fitting rats’ choice behavior to reinforcement-learning models with different state representations, we found that most rats were not able to acquire this shared rep- resentation, but instead learned about different trial types separately. A small group of rats, however, showed partial knowledge of the correct, parsimonious task representation, which opens up interesting questions on individual differences and the mechanism of representation learning.

RLDM Conference 2019 Conference Abstract

Rumination Steals Attention from Potentially Reinforcing Cues

  • Peter F Hitchcock
  • Yael Niv
  • Evan Forman
  • Nina Rothstein
  • Chris R Sims

Rumination is the tendency to respond to negative emotion or to the onset of a bad mood by ana- lyzing oneself and one’s distress. Rumination is associated with poor mental health and with various kinds of undesirable behavior, although how precisely rumination disrupts healthy behavior is mysterious. Some clues come from research showing that rumination alters both reinforcement learning (RL) and attention al- location, and that attention and RL systems cooperate to carry out adaptive behavior. If rumination disrupts this cooperation, this could explain its link to diverse maladaptive behaviors and ultimately to mental health problems. Thus, we investigated how rumination alters the cooperative interaction of attention and RL by experimentally inducing rumination and measuring its effects on a task designed to assay this interaction. As predicted, rumination impaired performance on the task, however, in a way that is not easily captured by computational models. To understand the subtle patterns of interference, we used trial-wise analyses and found that rumination disrupts the calibration of response speed to trial difficulty. This lack of calibration is especially evident in trials that follow cues to ruminate (suggesting that the cues may set off rumination that spills over into the task period, then gradually wanes) and correspond to the part of the task that rumination most impairs. In ongoing work, we are striving to model formally the precise mechanisms that rumination disrupts.

RLDM Conference 2017 Conference Abstract

A positive feedback loop between dopamine and freezing opposes extinction of fear

  • Lili Cai
  • Ilana Witten
  • Yael Niv

Striatal dopamine generates positive reinforcement, a property which is well accepted to con- tribute to drug addiction. However, it has been unclear if and how the reinforcing nature of striatal dopamine affects behavioral responses to aversive stimuli, and how this may be relevant to related psychiatric disor- ders such as post-traumatic stress disorder (PTSD). Here we identify a maladaptive function for striatal dopamine in the extinction of a fearful memory: striatal dopamine activity and fear behavior are related to each other through a positive feedback loop. This positive feedback loop opposes the extinction of a fear- ful memory and supports individual variability in fear extinction across mice. Thus, this work suggests that dopamine-mediated positive feedback loops may be a general mechanism underlying not only addiction, but also PTSD, and likely numerous other neuropsychiatric disorders characterized by dopamine dysfunction.

RLDM Conference 2017 Conference Abstract

Assessing the Potential of Computational Modeling in Clinical Science

  • Peter Hitchcock
  • Yael Niv
  • Angela Radulescu
  • Chris Sims

There has been much recent interest in using reinforcement learning (RL) model parameters as outcome measures in clinical science. A prerequisite to developing an outcome measure that might co-vary with a clinical variable of interest (such as an experimental manipulation, intervention, or diagnostic status) is first showing that the measure is stable within the same subject, absent any change in the clinical variable. Yet researchers often neglect to establish test-retest reliability. This is especially a problem with behavioral measures derived from laboratory tasks, as these often have abysmal test-retest reliability. Computational models of behavior may offer a solution. Specifically, model-based analyses should yield measures with lower measurement error than simple summaries of raw behavior. Hence model-based measures should have higher test-retest reliability than behavioral measures. Here, we show, in two datasets, that a pair of RL model parameters derived from modeling a trial-and-error learning task indeed show much higher test-retest reliability than a pair of raw behavioral summaries from the same task. We also find that the reliabilities of the model parameters tend to improve with time on task, suggesting that parameter estimation improves with time. Our results attest to the potential of computational modeling in clinical science.

RLDM Conference 2017 Conference Abstract

Bias in neural representational similarity analysis and a Bayesian method for reducing bias

  • Ming Bo Cai
  • Nicholas W. Schuck
  • Michael J. Anderson
  • Jonathan W. Pillow
  • Yael Niv

Understanding how the human brain represents the state space of a task is crucial for under- standing the neural basis of model-based learning and decision making. One approach towards this goal is representational similarity analysis (RSA), which allows one to analyze the structure of the neural represen- tation of different states as a participant is undertaking an RL task. However, when the transition between different task states is not entirely counterbalanced, the standard approach of RSA is guaranteed to intro- duce bias in the representational structure. Here we first illustrate the severity of this bias and analytically derive the source of the bias: serial correlations in fMRI noise, together with overlapping of hemodynamic responses between cognitive events, introduce structured noise in the estimated neural patterns. Correlation analysis of the estimated patterns translates the structured noise into spurious bias structure in the similarity matrix. The bias is especially severe with low signal-to-noise ratio and when task states cannot be random- ized, as in a Markov decision process. To overcome this bias, we propose an alternative Bayesian framework for computing representational similarity, an extension of the pattern component model (Diedrichsen et al. , 2011). We treat the covariance structure of the states and their neural representation as a hyper-parameter in a generative model of the fMRI data, and directly estimate this covariance structure from data while marginalizing over the unknown activity patterns. Converting the estimated covariance structure into a cor- relation matrix offers a much less biased estimate of representational similarity, and therefore the structure of the neural representation of a task. The method can also learn a shared similarity structure across multiple participants. Our tool is freely available in Brain Imaging Analysis Kit (BrainIAK).

RLDM Conference 2017 Conference Abstract

Characterizing people’s priors over naturalistic task structure

  • Gecia Bravo Hermsdorff
  • Yael Niv

Humans perform a large and diverse array of tasks (e. g. navigating, reading, cooking) with relative ease. However, if we think about all the possible tasks one could formulate, we would not be good at most of them (e. g. doing the Stroop task). This basic observation leads to the question: what are the essential properties of tasks that our brain is good at solving? From a computational perspective, the curse of dimensionality suggests that task representations should be compact, filtering out redundancies. However, reduced representations also constrain the set of tasks an agent can efficiently solve. Thus, for organisms to behave adaptively, they must adapt to leverage the relevant statistics of naturalistic tasks, i. e. tasks that the organism encounters in everyday life and have been relevant over evolutionary time-scales. What do people assume about the structure of naturalistic tasks? To study people’s priors over naturalistic task structure, we first map tasks structure to graphs, which enable us to quantify task structure using graph theoretical tools. We then use iterated learning (a process whereby an agent learns from data generated by another agent, who themselves learned it in the same way) to estimate peoples’ priors over task graphs. Specifically, 0) we create a task graph for the first subject; 1) the subject learns partial information about the task structure; 2) the subject infers the unshown task structure; 3) we construct a new graph from these responses for the next subject in this chain; we repeat steps 1-3 until the chain converges. Although we have promising preliminary results for priors over navigation graphs of university campuses, our experiments are still running. We believe that they will help understand people’s priors over the abstract structures of naturalistic tasks; how these priors depend on the number of objects (vertices), task domain (e. g. navigation vs. social), and contexts; and the individual variability amongst these priors.

RLDM Conference 2017 Conference Abstract

Latent Cause Inference in Social Biases

  • Yeon Soon Shin
  • Yael Niv

When making decisions in a social environment, how do we form impressions about a group of people whose members are diverse? If the majority of members are similar to one another with few members who are dissimilar from other people in the group, would experiences with those rare members influence the overall impression? Here, we explore how seemingly irrational biases where rare events gain prominence in overall estimation may result from normative inference of latent causes–causal structures of the world that generate a set of observations. We hypothesized that sparsity of events may lead to inference of unique latent causes for such events. This tendency to separate rare events to small latent causes, while grouping common events in large latent causes that explain multiple events, can cause overweighting of rare events in learning, if averaging is across latent causes rather than individual events. We tested this hypothesis by manipulating sparsity of non-overlapping event distributions. We first simulated the inference process, and showed the predicted effects of our theory. We then tested these predictions empirically in four decision-making experiments. Subjects observed a sequence of coin donations and were subsequently asked to estimate the average donation. As predicted by the latent-cause model, average estimation was biased toward sparse distributions (Exp 1 and 2). This bias was not explained by correctly averaging log- transformation of the donations (Exp 3), and disappeared when we interrupted the latent cause inference process by introducing step-by-step average estimation (Exp 4). These results suggest that social biases that have been found in empirical social cognition research may be the results of a semi-rational Bayesian latent cause inference process. Our theory also applies to formation of impressions about an individual on the basis of multiple interactions, and not only to evaluations of groups of people.

RLDM Conference 2017 Conference Abstract

Learning State Representations

  • Yael Niv

On the face of it, most real-world world tasks are hopelessly complex from the point of view of reinforce- ment learning mechanisms. In particular, due to the ”curse of dimensionality”, even the simple task of crossing the street should, in principle, take thousands of trials to learn to master. But we are better than that. . How does our brain do it? In this talk, I will argue that the hardest part of learning is not assigning values or learning policies, but rather deciding on the boundaries of similarity between experiences, that define the ”states” that we learn about. I will show behavioral evidence that humans and animals are con- stantly engaged in this representation learning process, and suggest that in a not too far future, we may be able to read out these representations from the brain, and therefore find out how the brain has mastered this complex problem. I will formalize the problem of learning a state representation in terms of Bayesian inference with infinite capacity models, and suggest that an understanding of the computational problem of representation learning can lead to insights into the machine learning problem of transfer learning, and psychological/neuroscientific questions about the interplay between memory and learning.

RLDM Conference 2017 Conference Abstract

Reinforcement learning predicts attention and memory in a multidimensional probabilistic task

  • Alana Jaskir
  • Yael Niv

Evidence suggests that attention and learning interact to help identify and learn about relevant dimensions that predict reward in a high dimensional environment. How exactly this changes the internal representation of the environment is still unclear. We tested human participants on a task in which they had to learn through trial and error to maximize reward in a probabilistic task in which reward probability was determined only by a single dimension (color, orientation or frequency) of the available visual stimuli. Occasionally, we probed participants’ memory of features in an entire dimension to gauge how their atten- tion changed during learning. We found a positive correlation between learning and memory on stimuli that subjects selected during each trial, suggesting that learned feature values influence how those features are encoded in an internal representation. Analyses also suggests that participants attend to the whole dimen- sion of higher rewarding features. This work paves the way for developing better models of how the brain compresses the state space of a high dimensional environment in relationship to ongoing value learning.

RLDM Conference 2017 Conference Abstract

Time-adaptive temporal difference reinforcement learning

  • Angela Langdon
  • Yael Niv

Anticipating the timing of rewards is as crucial to adaptive behavior as predicting what those rewards will be. In the brain, reward learning is thought to depend on dopamine signals that convey a prediction error whenever reward predictions do not accord with reality. Research further suggests that neural circuits in the basal ganglia use these error signals to learn reward predictions. This well-articulated behavioral and neural story nevertheless leaves an important question unaddressed: how does the neural machinery in the basal ganglia support learning of exactly when rewards should be expected? Prominent temporal difference reinforcement learning models purport to explain how the timing of rewards is learned, but rely for this on simplifying assumptions that are not tenable in the biological circuits that support reward prediction and learning in the brain. To address this question, we introduce time-adaptive temporal differ- ence reinforcement learning (time-adaptive TDRL), in which both the value and duration of underlying task states are learned concurrently. This model builds on the theory of partially-observable semi-Markov deci- sion processes, and introduces a mechanism for learning the duration of hidden task states by tracking the elapsed time between observations. We use this model to reproduce a number of features of dopaminergic reward prediction error signals to manipulations in the timing of reward delivery in simple learning tasks and interpret these response patterns as the reflection of an inference process that unfolds in time over the true underlying state of the task. In this framework, expectations about the likely time of state transitions are used to gate prediction error signaling, thereby concentrating learning at the time of predicted changes in the underlying state of the task, making testable predictions about the neural computation of dopaminergic prediction errors during learning.

NeurIPS Conference 2016 Conference Paper

A Bayesian method for reducing bias in neural representational similarity analysis

  • Mingbo Cai
  • Nicolas Schuck
  • Jonathan Pillow
  • Yael Niv

In neuroscience, the similarity matrix of neural activity patterns in response to different sensory stimuli or under different cognitive states reflects the structure of neural representational space. Existing methods derive point estimations of neural activity patterns from noisy neural imaging data, and the similarity is calculated from these point estimations. We show that this approach translates structured noise from estimated patterns into spurious bias structure in the resulting similarity matrix, which is especially severe when signal-to-noise ratio is low and experimental conditions cannot be fully randomized in a cognitive task. We propose an alternative Bayesian framework for computing representational similarity in which we treat the covariance structure of neural activity patterns as a hyper-parameter in a generative model of the neural data, and directly estimate this covariance structure from imaging data while marginalizing over the unknown activity patterns. Converting the estimated covariance structure into a correlation matrix offers a much less biased estimate of neural representational similarity. Our method can also simultaneously estimate a signal-to-noise map that informs where the learned representational structure is supported more strongly, and the learned covariance matrix can be used as a structured prior to constrain Bayesian estimation of neural activity patterns. Our code is freely available in Brain Imaging Analysis Kit (Brainiak) (https: //github. com/IntelPNI/brainiak), a python toolkit for brain imaging analysis.

RLDM Conference 2015 Conference Abstract

A learning mechanism for variability-sensitive reinforcement learning

  • Angela Langdon
  • Yael Niv

Variability in reward outcome is known to influence motivated behavior in humans and animals. While this sensitivity to so-called risk is a well-established behavioral phenomenon, the neural mechanisms that underlie its action are not well understood. We propose a model of reinforcement learning in the stria- tum that is sensitive to both the average of rewards and their variability, thereby outlining a putative neural mechanism for the influence of risk on learning and decision-making. Current theories of reinforcement learning in the basal ganglia propose a central role for dopamine in signaling errors in the prediction of reward, and hypothesize a central role for dopamine-mediated plasticity in the striatum in learning the as- sociation between states of the environment and the average future rewards they predict. We extend such a model of striatal reinforcement learning by introducing a parallel learning circuit that monitors ongoing dopaminergic prediction errors as a proxy for variability in reward outcomes around their mean. The spe- cific pattern of risk learnt from probabilistic rewards in the environment is dictated by nonlinearities in the response of the variability learning system and the step-size of its update rule. Coupling between the vari- ability learning system and the primary average reinforcement learning circuit allows learnt risk to affect the iterative update of state value, driving differentiation between states of equal expected future reward according to the weighting on their variability. This model demonstrates how parallel update systems tied to the same dopaminergically-mediated prediction error signal can interact locally in a neural circuit to pro- duce adaptive learning based on the experienced variability of rewards. We discuss the striatal cholinergic system as a putative neural substrate of the variability learning system and consider its potential role in the modulation of reinforcement learning in the striatum.

RLDM Conference 2015 Conference Abstract

Human Orbitofrontal Cortex Represents a Cognitive Map of State Space

  • Nicolas Schuck
  • Yael Niv

If Bob would buy or sell his stocks based on whether he sees his neighbor walking the dog or not, he won’t be very successful. Obviously, making a choice based on the wrong information will lead to wrong decisions. Reinforcement learning presupposes a compilation of all decision-relevant information into a single Markovian ‘state’ of the environment. Where do these states reside in the human brain? We have previously hypothesized that the orbitofrontal cortex (OFC) may play a key role in representing task states, especially when these are partially observable. Here we test this idea in humans, using multivariate decoding and representational similarity analysis of fMRI signals. In line with our hypothesis, we find evidence for a state representation in OFC. Moreover, we show that the fidelity of the state information in OFC, and the similarity between different states as they are represented neurally, robustly relate to performance differences. Our results suggest that internal state representations can be ‘read out’ for a variety of tasks, and indicate that the geometry of the individual state space can be used to make predictions about individual performance characteristics.

RLDM Conference 2015 Conference Abstract

Learning in multidimensional environments: Computational and neural processes across the lifespan

  • Reka Daniel
  • Yael Niv
  • Angela Radulescu

In order to behave efficiently in multidimensional environments, we have to learn to focus at- tention to only those dimensions of the environment that are currently predictive of reward. Unfortunately, both core components of this process, focusing attention and learning from rewards, have been shown to be compromised with healthy human aging. Here we investigate how learning and attention interact on the computational and neural level in older adults, and how these mechanisms differ from younger adults. To this end we collected behavioral and functional magnetic resonance imaging (fMRI) data from both older (M = 70. 0; range = 61-80) and younger (M = 22. 7; range = 18-35) adults performing a multidimensional probabilistic learning task. In this task, essentially a multi-dimensional bandit task, older adults showed worse performance; however, the same reinforcement learning model fit behavior in both groups. In fact, the model accounted better for older adults’ data than it did for younger adults. Neurally, activation in the Default Mode Network (DMN), a set of brain regions that is known to be deactivated during cognitively demanding tasks, was negatively correlated with the model-derived attentional focus in younger adults, sug- gesting that for younger adults the DMN was deactivated more at the beginning than at the end of games. This correlation was significantly weaker in older adults, indicating that older adults were not as successful in deactivating the DMN in accordance with the attentional demands of the task. In line with this, DMN deactivation during the first five trials of the task predicted higher performance in older adults, but not in younger adults. We conclude that computational mechanisms employed to optimize learning in multidimen- sional tasks do not change qualitatively across the human lifespan; however, older adults fail to selectively disengage their DMN as per task demands, leading to impaired behavioral performance. Poster T42*: Dopamine type 2 receptors control inverse temperature beta for transition from perceptual inference to reinforcement learning Eunjeong Lee*, NIMH/NIH; Olga Dal Monte, NIMH/NIH; Bruno Averbeck, NIH Decisions are based on a combination of immediate perception and previous experience. If the mapping between actions and outcomes in a context is unpredictable over time, decisions must be made on the basis of immediately available information. Alternatively, if action-outcome mappings can be learned by reinforcement, then this information can be combined with immediately available information. Previous neurophysiological results suggest that frontal-striatal circuits may be involved in the interaction between these processes. The role of dopamine, however, has not been examined directly. We injected locally dopamine type 1 (D1A; SCH23390) or type 2 (D2A; Eticlopride) antagonists or saline into the dorsal stria- tum while macaques performed an oculomotor sequential decision making task. Choices in the task were driven by perceptual inference and/or reinforcement of past choices. We found that the D2A affected deci- sions based on previous outcomes. When we fit Rescorla-Wagner models, the inverse temperature decreased after D2A injections into the dorsal striatum compared with a pre-injection period. We found that neither the D1A nor saline injections affected behavior. Overall, our results suggest D2Rs in the striatum control the inverse temperature in reinforcement learning.

RLDM Conference 2015 Conference Abstract

Model Comparison via Real-Time Manipulation of Human Learning

  • Andra Geana
  • Yael Niv

How do we learn what features of our multidimensional environment are relevant in a given task? To study the computational process underlying this type of ‘representation learning’, we propose a novel method of causal model comparison. Participants played a probabilistic learning task that required them to identify one relevant feature among several irrelevant ones. To compare between two models of this learning process, we fit the models to each participant’s initial behavioral data, and then ran the model alongside the participant during task performance, making predictions regarding the values underlying the participant’s choices in real time. To test these predictions, we used each model to try to perturb the participant’s learning process: based on the model’s predictions, we crafted the available stimuli so as to either obscure infor- mation regarding which feature is more relevant to solving the task, or make this information more readily available. A model whose predictions coincide with the true learned values in the participant’s mind, is expected to be effective in perturbing learning in this way, whereas a model whose predictions stray from the true learning process should not. Indeed, we show that in our task a feature-level reinforcement-learning (RL) model can be used to causally help or hinder participants’ learning, while a Bayesian ideal observer model cannot exert such an effect on learning. In particular, games in which we used the RL model to manipulate learning had significantly higher percentage of learned games and average scores in the ‘Help’ as compared to the ‘Hurt’ condition. Games that were manipulated using predictions of the Bayesian model showed no differences for helping versus hurting conditions. Beyond informing us about the computational substrates of representation learning, our manipulation represents a sensitive method for model comparison, which allows us to change the course of people’s learning in real-time.

RLDM Conference 2015 Conference Abstract

Modeling the Hemodynamic Response Function for Prediction Errors in the Human Ventral

  • Gecia Bravo Hermsdorff
  • Yael Niv

Recent years have seen a proliferation of studies in which computational models are used to spec- ify precisely a set of hypotheses regarding reinforcement learning and decision making in humans, which are then tested against data from functional magnetic resonance imaging (fMRI). fMRI research proceeds by using information provided by the blood oxygenation level dependent (BOLD) signal to make inferences about the underlying neural activation. The focus of much of this model-based fMRI effort has been on the ventral striatum (VS), where the BOLD response has been shown to reflect reward prediction error signals (momentary differences between expected and obtained outcomes) from dopaminergic afferents. To make sensible inferences from fMRI data it is important to accurately model the hemodynamic response function (HRF), i. e. , the hemodynamic response evoked by a punctate neural event. A canonical HRF, mapped for sensory cortical regions, is commonly used for analyzing activity throughout the brain despite the fact that hemodynamics are known to vary across regions, in particular in subcortical areas such as the VS. Here we use data from an experiment focused on learning from prediction errors (Niv et al. , 2010) to fit a VS-specific HRF function. Our results show that the VS HRF differs significantly from the canonical HRF, most im- portantly peaking at 6 sec rather than at 5 sec. We demonstrate the superiority of the VS HRF in modeling data by showing that it increases statistical power. This result is particularly relevant to fMRI studies of reinforcement learning and decision making as many of these rely on fine analysis of the VS BOLD activity to distinguish between important but subtle differences in computational models of learning and choice. We therefore recommend the use of this new HRF for future fMRI studies of the ventral striatum.

RLDM Conference 2015 Conference Abstract

Neural representations of posterior distributions over latent causes

  • Stephanie Chan
  • Kenneth Norman
  • Yael Niv

The world is governed by unobserved ‘causes’, which generate the events that we do observe. In reinforcement learning, and in particular, in partially observable Markov decision processes (POMDPs), these hidden causes are the ‘states’ of the task. Accurate inference about the current state, based on the agent’s observations, is critical for optimal decision making and learning. Here we investigate the neural basis of this type of inference about hidden causes in the human brain. In particular, we are interested in the neural substrates that allow humans to maintain, approximately or exactly, a belief distribution over the hidden states, which assigns varying levels of probability to each. We conducted an experiment in which participants viewed sequences of animals drawn from one of four ‘sectors’ in a safari. They were tasked with guessing which sector the animals were from, based on previous experience with the likelihood of each animal in each sector. We used functional magnetic resonance imaging (fMRI) to investigate brain representations of the posterior distribution P(sector — animals). Our results suggest that neural patterns in the lateral orbitofrontal cortex, angular gyrus, and precuneus correspond to a posterior distribution over ‘sectors’. We also show that a complementary set of areas are involved in the updating of the posterior distribution. These results are consistent with previous work implicating these areas in the representation of ‘state’ from reinforcement learning (Wilson et al, 2014) and ‘schemas’ or ‘situation models’ (Ranganath & Ritchey, 2012).

RLDM Conference 2013 Conference Abstract

A Reinforcement Learning Theory of Mood Instability

  • Eran Eldar
  • Yael Niv

A propensity to experience cycles of good and bad mood characterizes the emotional life of pa- tients suffering from bipolar disorder, as well as of healthy but susceptible individuals. What neural mecha- nism brings about mood cycles? Here, we show that oscillations of mood naturally emerge within a standard reinforcement learning framework as a result of two plausible assumptions - that reward prediction error affects mood, and that mood affects perception of reward. We then provide behavioral and neural evidence that supports the validity of these assumptions, specifically in individuals that are susceptible to mood fluc- tuations. We conducted a trial-and-error learning experiment, in which participants were asked to choose between slot machines that yielded small monetary rewards with fixed probabilities. In the middle of the experiment, participants took part in a “wheel of fortune” draw, in which they either won or lost a (rela- tively) large sum ($7). Post-experiment, we tested the effect of the wheel of fortune draw on participants’ valuations by asking them to choose between equally-rewarding slot machines that they had encountered be- fore and after the draw. As predicted, both subjective reports of mood and valuations of slot machines were significantly affected by the wheel of fortune draw for those participants susceptible to mood fluctuations (assessed using a self-report questionnaire). Specifically, these participants tended to favor the slot machines that appeared after the draw if the draw was successful, and the slot machines that appeared before the draw if the draw was unsuccessful. The results were replicated in a different group of participants performing the experiment in an MRI scanner. Supporting the interpretation of these results in terms of biased perception of rewards in susceptible participants, striatal BOLD responses to slot-machine rewards became stronger after a successful draw and weaker after an unsuccessful one.

RLDM Conference 2013 Conference Abstract

Age-related Differences in Learning to Selectively Attend

  • Angela Radulescu
  • Reka Daniel
  • Yael Niv

When confronted with many stimuli in a complex world, how does the aging brain learn where to focus attention? Previous work shows that older adults have more difficulty switching between different task-relevant dimensions. It remains unclear, however, whether and how the cognitive strategies they use differ from those employed by younger adults. Here we focus on age-related differences in the dynamics of representation learning, where participants learn which stimulus features are relevant to each task through trial and error. We compare the behavior of older and younger adults in a multidimensional reinforcement learning task designed to study how subjects update their representations on-line, and propose a series of models that implement various forms of selective attention. Model-based analysis of choice patterns shows that both younger and older adults employ attention during learning. However, older adults seem to maintain a narrower attentional filter, a cognitive strategy that might reflect an adaptation to changes in the interaction between the dopaminergic system and the prefrontal cortex.

RLDM Conference 2013 Conference Abstract

Human reinforcement learning processes act on learned attentionally-filtered representations of the world

  • Yuan Chang Leong
  • Yael Niv

Reinforcement learning (RL) models are often applied to study human learning and decision- making. However, simple RL algorithms do not fare well in explaining learning behavior in real world situations where the environment is high-dimensional and the relevant states are not known. As a solution, we propose that RL processes act on an attentionally-filtered representation of the environment. This im- proves the computational efficiency of RL by constraining the state-space that the learning agent has to consider. We further propose that the attention filter is learned and is dynamically modulated according to the outcomes of ongoing decisions. To test our hypotheses, we had participants perform a decision-making task with multi-dimensional stimuli and probabilistic awards. Model-based analysis of participants’ choices suggests that participants prefer strategies that favor computational efficiency at the expense of statistical optimality. To better study the dynamics of attention, we had a group of participants perform a variant of the task in which they had to select the dimensions they wanted to view before making their choice. We treat- ed the viewed dimensions as a proxy for participants’ attention filter. Our models fit the data better when learning was restricted to attended dimensions, suggesting that participants do indeed constrain choice and learning to a subset of dimensions. Finally, attention dynamics themselves were best explained by a mod- el that preferentially attended to dimensions with features that have acquired high value over the course of learning. This result provides evidence that the attention filter is dynamically modulated as participants receive feedback from ongoing decisions.

RLDM Conference 2013 Conference Abstract

Humans employ selective attention when learning in complex environments: evidence from computational modeling and neuroimaging

  • Reka Daniel
  • Vivian DeWoskin
  • Yuan Chang Leong
  • Angela Radulescu
  • Yael Niv

In two experiments, we show how humans employ selective attention to enable efficient rein- forcement learning in naturalistic environments. Despite the invaluable contribution of current reinforce- ment learning models to our understanding of animal and human learning, they do not scale well to complex environments with many multidimensional stimuli. Fortunately, in ecologically valid settings only a few fea- tures of the environment are relevant for maximizing rewards, while other features can be safely ignored in order to facilitate learning and generalization. In Experiment 1 we propose a function approximation model that can scale to multidimensional environments and test it on choice data in a multidimensional reinforce- ment learning task. We show that humans learn about the distinct features of stimuli separately, and that during successful learning they assign value only to the features that are important for predicting reward. The differential weighting of stimulus features in our model can be interpreted as an attentional filter. Using functional neuroimaging (fMRI) data we demonstrate that the width of this attentional filter is correlated with activation in the intraparietal sulcus (IPS), a structure known to be involved in the control of attention. In Experiment 2 we introduce a novel method for decoding the focus of attention on a trial-by-trial basis. Using a combination of eye-tracking and multivariate pattern analysis of fMRI data, we successfully deter- mine which of multiple simultaneously presented stimulus features participants attend to on each trial. We compare this data to the model-derived focus of attention and provide further evidence that learning operates on an attentionally filtered representation of the environment. Taken together, our findings demonstrate that humans maximize rewards in complex multidimensional environments by focusing attention only on the reward-predicting features of each stimulus.

RLDM Conference 2013 Conference Abstract

Is model fitting necessary for model-based fMRI?

  • Robert Wilson
  • Yael Niv

Model-based analysis of functional magnetic resonance imaging (fMRI) data is an important tool for investigating the computational role of different brain regions. With this method, theoretical models of behavior can be leveraged to find the brain structures underlying latent variables that are key to specific algorithms, such as prediction errors in temporal difference learning. A key step in this type of analysis is model fitting. Most commonly, a model is first fit to behavioral data to establish ‘good’ parameters. These are then used to generate model-based regressors of the quantity of interest, for regressing against brain activations acquired using fMRI. While such model fitting may intuitively seem like good practice, in this work we ask whether it is really necessary. We focus on the classic reinforcement learning regressors for value and prediction error and examine their sensitivity to perturbations of the learning rate parameter both in theory and in a previously published dataset. Surprisingly, in many cases, we find that fitting the learning rate is not necessary to generate good regressors and in some situations, even use of the worst possible parameter settings affects the model-based analysis only marginally. Our results suggest that precise model fitting is not necessary for model-based fMRI, thereby freeing experimental design from the constraint of allowing precise fits. They also highlight the limited use of fMRI data for arbitrating between different (correlated) models or model parameters.

RLDM Conference 2013 Conference Abstract

Optimal Task Decomposition

  • Alec Solway
  • Natalia Cordova
  • Debbie Yee
  • Andrew Barto
  • Yael Niv
  • Matthew Botvinick

Reinforcement learning has provided a rich framework for understanding the computational sub- strates underlying human decision making. Most work has so far has focused on simple decision problems with small state spaces. More recently researchers have begun applying ideas from hierarchical reinforce- ment learning, and the options framework in particular, to address how human decision making may scale. This framework specifies how the computational complexity associated with both learning and planning in high-dimensional state spaces may be reduced through the use of temporal abstraction. In addition to primitive actions that lead to transitions between adjacent states, the agent can execute options that lead to transitions between distant states. While there is now evidence that humans make use of options, it is unclear how they come to select which options are useful in the first place. We present option selection as a Bayesian model comparison problem and show that the options people select are those corresponding to the maximal model evidence.

RLDM Conference 2013 Conference Abstract

“Identity prediction errors” and model-based learning

  • Stephanie Chan
  • Nina Lopatina
  • Yael Niv

It is known that humans and animals perform model-based reinforcement learning, in which decision-making uses a full model of the environment, including the transition probabilities between states. This is in contrast to model-free reinforcement learning, which relies on a “value” for each state or state- action pair. In model-free reinforcement learning, state values are thought to be learned using “value predic- tion errors” the difference between the expected value and the actual value observed. This is a computational- ly efficient means of learning, and, furthermore, these predictions errors famously seem to be represented by the activity of midbrain dopamine neurons. How are the full models, needed for model based reinforcement learning, learned? It has been hypothesized that there exist analogous “identity (state) prediction errors” in the brain, which are used to learn the transition probabilities between states. Identity prediction errors are elicited when a specific outcome (or state) is unexpected, regardless of whether the utility value of the outcome is different from what was expected. Here, we use fMRI to find signals corresponding to identity prediction errors in the brain. Based on prior work, we expected to find such signals in the orbitofrontal cortex.

NeurIPS Conference 2008 Conference Paper

Learning to Use Working Memory in Partially Observable Environments through Dopaminergic Reinforcement

  • Michael Todd
  • Yael Niv
  • Jonathan Cohen

Working memory is a central topic of cognitive neuroscience because it is critical for solving real world problems in which information from multiple temporally distant sources must be combined to generate appropriate behavior. However, an often neglected fact is that learning to use working memory effectively is itself a difficult problem. The Gating" framework is a collection of psychological models that show how dopamine can train the basal ganglia and prefrontal cortex to form useful working memory representations in certain types of problems. We bring together gating with ideas from machine learning about using finite memory systems in more general problems. Thus we present a normative Gating model that learns, by online temporal difference methods, to use working memory to maximize discounted future rewards in general partially observable settings. The model successfully solves a benchmark working memory problem, and exhibits limitations similar to those observed in human experiments. Moreover, the model introduces a concise, normative definition of high level cognitive concepts such as working memory and cognitive control in terms of maximizing discounted future rewards. "

NeurIPS Conference 2005 Conference Paper

How fast to work: Response vigor, motivation and tonic dopamine

  • Yael Niv
  • Nathaniel Daw
  • Peter Dayan

Reinforcement learning models have long promised to unify computa- tional, psychological and neural accounts of appetitively conditioned be- havior. However, the bulk of data on animal conditioning comes from free-operant experiments measuring how fast animals will work for rein- forcement. Existing reinforcement learning (RL) models are silent about these tasks, because they lack any notion of vigor. They thus fail to ad- dress the simple observation that hungrier animals will work harder for food, as well as stranger facts such as their sometimes greater produc- tivity even when working for irrelevant outcomes such as water. Here, we develop an RL framework for free-operant behavior, suggesting that subjects choose how vigorously to perform selected actions by optimally balancing the costs and benefits of quick responding. Motivational states such as hunger shift these factors, skewing the tradeoff. This accounts normatively for the effects of motivation on response rates, as well as many other classic findings. Finally, we suggest that tonic levels of dopamine may be involved in the computation linking motivational state to optimal responding, thereby explaining the complex vigor-related ef- fects of pharmacological manipulation of dopamine.

v2026.09.13