Arrow Research search

Author name cluster

Geoffrey Schoenbaum

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
1 author row

Possible papers

3

RLDM Conference 2015 Conference Abstract

Contingency and Correlation in Reversal Learning

  • Bradley Pietras
  • Peter Dayan
  • Thomas Stalnaker
  • Geoffrey Schoenbaum
  • Tzu-Lan Yu

Reversal learning is one of the most venerable paradigms for studying the acquisition, extinction, and reacquisition of knowledge in humans and other animals. It has been of particular value in asking questions about the roles played by prefrontal structures such as the orbitofrontal cortex (OFC). Indeed, evidence from rats and monkeys suggests that these areas are involved in various forms of context-sensitive inference about the contingencies linking cues and actions over time to the value and identity of predicted outcomes. In order to explore these roles in depth, we fit data from a substantial behavioural neuroscience study in rodents who experienced blocks of free- and forced-choice instrumental learning trials with identity or value reversals at each block transition. We constructed two classes of models, fit their parameters using a random effects treatment, tested their generative competence, and selected between them based on a complexity-sensitive integrated Bayesian Information Criteria score. One class of ‘return’-based models was based on elaborations of a standard Q-learning algorithm, including parameters such as different learning rates or combination rules for forced- and fixed-choice trials, behavioural lapses, and eligibility traces. The other novel class of ‘income’-based models exploited the weak notion of contingency over time advocated by Walton et al (2010) in their analysis of the choices of monkeys with OFC lesions. We show that income- based and return-based models are both able to predict the behaviour well, and examine their performance and implications for reinforcement learning. The outcome of this study sets the stage for the next phase of the research that will attempt to correlate the values of the parameters to neural recordings taken in the rats while performing the task.

RLDM Conference 2015 Conference Abstract

Signaling prediction for size versus value of rewards in rodent orbitofrontal cortex during Pavlo- vian unblocking

  • Geoffrey Schoenbaum
  • Nina Lopatina
  • Brian Sadacca
  • Michael Mc-

Modern reinforcement learning models and learning theories distinguish at least two different forms of reward prediction: specific features or properties of rewards, and value or general utility of rewards. Formation of specific goals requires intact prediction of reward features. Maximizing the value of these goals requires intact prediction of reward value. While reward size and reward value are inextricably linked, the changes of neural activity of individual units in response to rewards of different sizes can shed light on whether individual neurons’ encoding reflects reward size or value. The current study examined changes in orbitofrontal cortex (OFC) neural activity using single-unit electrophysiological recording, measuring activity during a novel Pavlovian unblocking procedure that assesses excitatory and inhibitory cue learning driven by upshifts or downshifts in expected reward size. We have recorded hundreds of OFC neurons during the task. Preliminary analyses show that cue-related activity is regulated by the unblocking paradigm used, with findings of differential firing to blocked, size-downshift and size-upshift cues in individual neurons. A more comprehensive analysis of these neural data will be presented. Poster T8*: Ensembles of Shapings Tim Brys*, Vrije Universiteit Brussel; Anna Harutyunyan, Vrije Universiteit Brussel; Matthew Taylor, Washington State University; Ann Nowé, Vrije Universiteit Brussel Many reinforcement learning algorithms try to solve a problem from scratch, i. e. , without a pri- ori knowledge. This works for small and simple problems, but quickly becomes impractical as problems of growing complexity are tackled. The reward function with which the agent evaluates its behaviour of- ten is sparse and uninformative, which leads to the agent requiring large amounts of exploration before feedback is discovered and good behaviour can be generated. Reward shaping is one approach to address this problem, by enriching the reward signal with extra intermediate rewards, often of a heuristic nature. These intermediate rewards may be derived from expert knowledge, knowledge transferred from a previous task, demonstrations provided to the agent, etc. In many domains, multiple such pieces of knowledge are available, and could all potentially benefit the agent during its learning process. We investigate the use of ensemble techniques to automatically combine these various sources of information, helping the agent learn faster than with any of the individual pieces of information alone. We empirically show that the use of such ensembles alleviates two tuning problems: (1) the problem of selecting which (combination of) heuris- tic knowledge to use, and (2) the problem of tuning the scaling of this information as it is injected in the original reward function. We show that ensembles are both robust against bad information and bad scalings.

RLDM Conference 2013 Conference Abstract

VTA neurons show value prediction signals for cues possessing inferred value

  • Brian Sadacca
  • Geoffrey Schoenbaum

In recent years, the application of reinforcement learning models to neuroscientific data has led to an understanding that multiple decision making systems do, in fact, coexist in the brain, with these two particular threads of the cannon described in terms of ”model-free’ and ”modal based’ decision making, respectively. Dopamine (DA) released from the ventral tegmental area (VTA) has often been related to model-free learning; actions that occur just prior to increases in DA release are more likely to occur again, and cues that occur just prior to increases in DA release are more likely to be sought. However, it is unclear if dopamine neurons of the VTA have access to information about cues whose relationship to reward is inferred. To test if VTA neurons can predict this value, rats were run in a sensory pre-conditioning task, while the activity of VTA neurons recorded extracellularly. In this task, rats first learn a timing relationship within two pairs of cues in the absence of reward (A before B, and C before D). Rats then learn that one of the cues (B+) predicts the availability of reward while another (D-) predicts the absence of reward. In a final phase, rat behavior is monitored while cues A and C occur; here rats infer the learned relationship between cues A and B, and try to access reward in the presence of cue A alone. In the test, neurons showed early phasic responses to cues with inferred value mimicking the responses to cues explicitly paired with reward, clearly demonstrating that the VTA can access values beyond the model-free systems of the brain. In addition, cue A (predicting the explicitly rewarded cue), also evoked a late phasic response, that seems to predict the timing of the explicitly rewarded cue. These finding demonstrate a role for the VTA beyond simply driving learning based on immediate experiences, and suggests that there may be further effects of drug-induced plasticity in VTA beyond changes to simple cue-reward learning.

v2026.09.13