Arrow Research search

Author name cluster

Matthew Botvinick

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
1 author row

Possible papers

16

TMLR Journal 2023 Journal Article

Adaptive patch foraging in deep reinforcement learning agents

  • Nathan Wispinski
  • Andrew Butcher
  • Kory Wallace Mathewson
  • Craig S Chapman
  • Matthew Botvinick
  • Patrick M. Pilarski

Patch foraging is one of the most heavily studied behavioral optimization challenges in biology. However, despite its importance to biological intelligence, this behavioral optimization problem is understudied in artificial intelligence research. Patch foraging is especially amenable to study given that it has a known optimal solution, which may be difficult to discover given current techniques in deep reinforcement learning. Here, we investigate deep reinforcement learning agents in an ecological patch foraging task. For the first time, we show that machine learning agents can learn to patch forage adaptively in patterns similar to biological foragers, and approach optimal patch foraging behavior when accounting for temporal discounting. Finally, we show emergent internal dynamics in these agents that resemble single-cell recordings from foraging non-human primates, which complements experimental and theoretical work on the neural mechanisms of biological foraging. This work suggests that agents interacting in complex environments with ecologically valid pressures arrive at common solutions, suggesting the emergence of foundational computations behind adaptive, intelligent behavior in both biological and artificial agents.

RLDM Conference 2019 Conference Abstract

Learned human-agent decision-making, communication and joint action in a virtual reality environment

  • Patrick M. Pilarski
  • Andrew Butcher
  • Matthew Botvinick
  • Andrew Bolt
  • Adam Parker

Humans make decisions and act alongside other humans to pursue both short-term and long-term goals. As a result of ongoing progress in areas such as computing science and automation, humans now also interact with non-human agents of varying complexity as part of their day-to-day activities; substantial work is being done to integrate increasingly intelligent machine agents into human work and play. With increases in the cognitive, sensory, and motor capacity of these agents, intelligent machinery for human assistance can now reasonably be considered to engage in joint action with humans—i. e. , two or more agents adapt- ing their behaviour and their understanding of each other so as to progress in shared objectives or goals. The mechanisms, conditions, and opportunities for skillful joint action in human-machine partnerships is of great interest to multiple communities. Despite this, human-machine joint action is as yet under-explored, especially in cases where a human and an intelligent machine interact in a persistent way during the course of real-time, daily-life experience (as opposed to specialized, episodic, or time-limited settings such as game play, teaching, or task-focused personal computing applications). In this work, we contribute a virtual reality environment wherein a human and an agent can adapt their predictions, their actions, and their communica- tion so as to pursue a simple foraging task. In a case study with a single participant, we provide an example of human-agent coordination and decision-making involving prediction learning on the part of the human and the machine agent, and control learning on the part of the machine agent wherein audio communication signals are used to cue its human partner in service of acquiring shared reward. These comparisons suggest the utility of studying human-machine coordination in a virtual reality environment, and identify further research that will expand our understanding of persistent human-machine joint action.

RLDM Conference 2019 Conference Abstract

Understanding Emergent Structure-Based Learning in Recurrent Neural

  • Kevin J Miller
  • Jane Wang
  • Zeb Kurth-Nelson
  • Matthew Botvinick

Humans and animals possess rich knowledge about the structure of the environments that they inhabit, and use this knowledge to scaffold ongoing learning about changes in those environments. This structural knowledge allows for more meaningful credit assignment, which improves the efficiency with which new experiential data can be used. Such data-efficiency continues to challenge modern artificial in- telligence methods, especially those combining reinforcement learning with deep neural networks. One promising route to improving the data-efficiency of these networks is learning-to-learn, in which a hand- crafted, general-purpose, data-inefficient learning process operating over the weights of a recurrent neural network gives rise to an emergent, environment-specific, data-efficient learning process operating in the ac- tivity dynamics controlled by those weights. Here, we apply learning-to-learn to several structured tasks and characterize the emergent learning algorithms that result, using tools inspired by cognitive neuroscience. We show that the behavior of these algorithms, like that of human and animal subjects, shows clear evidence of structure-based credit assignment. We further show that the dynamical system which embodies these algorithms is relatively low-dimensional, and exhibits interpretable dynamics. This work provides an im- proved understanding of how structure-based learning can take place in artificial systems, as well as a set of hypotheses about how it may be implemented in the brain.

RLDM Conference 2015 Conference Abstract

Reward-based decision making with infinite choice sets

  • Jonathan Berliner
  • Matthew Botvinick

Behavioral and neuroscientific research has given rise to detailed, experimentally validated mod- els of both perceptual and reward-based decision making. In the case of binary choice, these models are underpinned by a precise normative account. Recently, considerable effort has been given towards the chal- lenge of extending this normative account to de- cision problems involving more than two choices. In the present work, we address decision problems with uncountably infinite candidate choices. Specifically, we present initial attempts to develop a normative account of reward-based deci- sion making in continuous choice domains. To develop this account, using tools from the fields of Bayesian optimization and Gaussian process theory, we can assess human behavior in a decision making task for which normative predictions can be precisely specified. The present work confirms two predictions critical to the development of the frame- work. First, we find that choice behavior adapts to domain structure, as required by the normative account. Second, we find that choice behavior during information collection can be better described as function max- imization rather than function approximation, and so fits the framework of Bayesian optimization more so than that of active learning.

RLDM Conference 2015 Conference Abstract

The successor representation in human reinforcement learning: evidence from retrospective revaluation

  • Ida Momennejad
  • Jin Cheong
  • Matthew Botvinick
  • Samuel Gershman

Reinforcement learning (RL) has been posed as a competition between model-free (MF) and model-based (MB) learning. MF learning cannot solve problems such as revaluation and latent learning, hallmarks of MB behavior. However, we suggest that there are varieties of MB-like behavior that are also predicted by an alternative solution to the RL problem that lies between the MF and MB strategies. In particular, the successor representation (SR) can account for certain kinds of retrospective revaluation of rewards, a behavior that has traditionally been ascribed to the MB system. We conducted two experiments to test this hypothesis and compared the classic ‘reward devaluation’ (reward structure changes, transition structure stays the same) with ‘transition devaluation’ (reward structure stays the same, transition structure changes). A pure SR strategy will only be sensitive to reward but not transition devaluation, because the SR effectively ‘compiles’ the transition structure and therefore cannot adapt quickly to changes. Behav- ioral results from Study 1 showed that subjects were more sensitive to reward than transition devaluation (n=58, p<. 05). However, subjects still showed some sensitivity to transition devaluation, inconsistent with a pure SR strategy. These results point to the possibility that subjects may employ a mixed SR-MB strategy, whereby the value function for a MB strategy is initialized using the SR. With more processing time, the in- fluence of the initialization diminishes, causing behavior to resemble a pure MB strategy. Study 2 tested the hybrid SR-MB model of retrospective revaluation. Consistent with our predictions, fast responses showed greater sensitivity to reward than transition devaluation, while slower responses displayed equal sensitivity to both (n=52, p<. 05). Very slow responses showed low sensitivity to both types of devaluation; consistent with the hypothesis that noise accumulates in the MB computation, impairing performance.

NeurIPS Conference 2014 Conference Paper

Design Principles of the Hippocampal Cognitive Map

  • Kimberly Stachenfeld
  • Matthew Botvinick
  • Samuel Gershman

Hippocampal place fields have been shown to reflect behaviorally relevant aspects of space. For instance, place fields tend to be skewed along commonly traveled directions, they cluster around rewarded locations, and they are constrained by the geometric structure of the environment. We hypothesize a set of design principles for the hippocampal cognitive map that explain how place fields represent space in a way that facilitates navigation and reinforcement learning. In particular, we suggest that place fields encode not just information about the current location, but also predictions about future locations under the current transition distribution. Under this model, a variety of place field phenomena arise naturally from the structure of rewards, barriers, and directional biases as reflected in the transition policy. Furthermore, we demonstrate that this representation of space can support efficient reinforcement learning. We also propose that grid cells compute the eigendecomposition of place fields in part because is useful for segmenting an enclosure along natural boundaries. When applied recursively, this segmentation can be used to discover a hierarchical decomposition of space. Thus, grid cells might be involved in computing subgoals for hierarchical reinforcement learning.

RLDM Conference 2013 Conference Abstract

A seven parameter mixture model that describes steady-state rodent behavior on a two-armed bandit task nearly as well as it can be described; Applications to orbitofrontal cortex inactivations

  • Kevin Miller
  • Jeffery Erlich
  • Charles Kopec
  • Matthew Botvinick
  • Carlos Brody

Simple reinforcement learning models are widely used to interpret human and animal behavior on decision-making tasks in dynamic environments. These models have the advantage of simplicity, but provide only an incomplete description of choice behavior. Regression models and Markov models provide much more complete descriptions of behavior, but come with the cost of having dozens, hundreds, or even thousands of free parameters. This makes their results difficult to interpret, and also makes them applicable only to relatively large datasets. We present a mixture model which we believe to be an ideal compromise. This model contains contributions from a variety of behavioral strategies, including temporal difference learning, win-stay/lose-switch, and perseveration, and combines them to determine choice probability. We show that this model is nearly as good as a regression model (within 0. 2 % of variance explained) at describing rat behavior on a two-armed bandit task. In turn, we show that the regression models are nearly as good (within 0. 1 %) as maximally complete Markov models. This supports the idea that our mixture model describes behavior on the two- armed bandit task nearly as well as any stationary model possibly could. Our model contains only seven free parameters, making it applicable to datasets of the size typically found in neuroscience experiments. We have collected data from rats whose orbitofrontal cortex (OFC) has been inactivated using the GABA agonist muscimol. Most rats are impaired at the task during OFC inactivation, and the parameter fits of the model suggest insights into the specific nature of the impairments.

RLDM Conference 2013 Conference Abstract

Assessing Structure Learning in Motor Tasks

  • Jonathan Berliner
  • Matthew Botvinick
  • Jordan Taylor

There is mounting evidence that humans utilize “structure learning, ” the identification and utiliza- tion of the latent parameters driving action outcomes in a given environment, in motor learning tasks. There is also accumulating evidence suggesting that people engage in “active learning, ” selectively sampling their environment in order to quickest reduce their “hypothesis space” of sets of variables that may underlie the environment. We sought to directly assess whether people actively sample their environment in order to best learn its latent structure. Subjects made non-rewarded “training reaches, ” which they used to inform their movements on rewarded “test reaches, ” in a rapidly changing environment. We assessed whether participants would learn to prefer to make training reaches towards “information-bearing” areas that most reduced the hypothesis space of candidate environment structures. Participants learned to selectively sample the more information-bearing areas of the task environment. Further, given equal information-yield across the train- ing space, participants preferred to train in areas near those in which they expected to later be tested, a trait not predicted by certain implementations of structure learning in the motor domain. We provide evidence suggesting that, when engaging in motor tasks, people may employ heuristic-based movement strategies that are more agnostic to the environment than strategies utilizing hypothesized latent structure would predict.

RLDM Conference 2013 Conference Abstract

Optimal Task Decomposition

  • Alec Solway
  • Natalia Cordova
  • Debbie Yee
  • Andrew Barto
  • Yael Niv
  • Matthew Botvinick

Reinforcement learning has provided a rich framework for understanding the computational sub- strates underlying human decision making. Most work has so far has focused on simple decision problems with small state spaces. More recently researchers have begun applying ideas from hierarchical reinforce- ment learning, and the options framework in particular, to address how human decision making may scale. This framework specifies how the computational complexity associated with both learning and planning in high-dimensional state spaces may be reduced through the use of temporal abstraction. In addition to primitive actions that lead to transitions between adjacent states, the agent can execute options that lead to transitions between distant states. While there is now evidence that humans make use of options, it is unclear how they come to select which options are useful in the first place. We present option selection as a Bayesian model comparison problem and show that the options people select are those corresponding to the maximal model evidence.

AIJ Journal 2013 Journal Article

Using Wikipedia to learn semantic feature representations of concrete concepts in neuroimaging experiments

  • Francisco Pereira
  • Matthew Botvinick
  • Greg Detre

In this paper we show that a corpus of a few thousand Wikipedia articles about concrete or visualizable concepts can be used to produce a low-dimensional semantic feature representation of those concepts. The purpose of such a representation is to serve as a model of the mental context of a subject during functional magnetic resonance imaging (fMRI) experiments. A recent study by Mitchell et al. (2008) [19] showed that it was possible to predict fMRI data acquired while subjects thought about a concrete concept, given a representation of those concepts in terms of semantic features obtained with human supervision. We use topic models on our corpus to learn semantic features from text in an unsupervised manner, and show that these features can outperform those in Mitchell et al. (2008) [19] in demanding 12-way and 60-way classification tasks. We also show that these features can be used to uncover similarity relations in brain activation for different concepts which parallel those relations in behavioral data from human subjects.

NeurIPS Conference 2012 Conference Paper

A systematic approach to extracting semantic information from functional MRI data

  • Francisco Pereira
  • Matthew Botvinick

This paper introduces a novel classification method for functional magnetic resonance imaging datasets with tens of classes. The method is designed to make predictions using information from as many brain locations as possible, instead of resorting to feature selection, and does this by decomposing the pattern of brain activation into differently informative sub-regions. We provide results over a complex semantic processing dataset that show that the method is competitive with state-of-the-art feature selection and also suggest how the method may be used to perform group or exploratory analyses of complex class structure.

YNIMG Journal 2011 Journal Article

Information mapping with pattern classifiers: A comparative study

  • Francisco Pereira
  • Matthew Botvinick

Information mapping using pattern classifiers has become increasingly popular in recent years, although without a clear consensus on which classifier(s) ought to be used or how results should be tested. This paper addresses each of these questions, both analytically and through comparative analyses on five empirical datasets. We also describe how information maps in multiple class situations can provide information concerning the content of neural representations. Finally, we introduce a publically available software toolbox designed specifically for information mapping.

YNIMG Journal 2009 Journal Article

Machine learning classifiers and fMRI: A tutorial overview

  • Francisco Pereira
  • Tom Mitchell
  • Matthew Botvinick

Interpreting brain image experiments requires analysis of complex, multivariate data. In recent years, one analysis approach that has grown in popularity is the use of machine learning algorithms to train classifiers to decode stimuli, mental states, behaviours and other variables of interest from fMRI data and thereby show the data contain information about them. In this tutorial overview we review some of the key choices faced in using this approach as well as how to derive statistically significant results, illustrating each point from a case study. Furthermore, we show how, in addition to answering the question of ‘is there information about a variable of interest’ (pattern discrimination), classifiers can be used to tackle other classes of question, namely ‘where is the information’ (pattern localization) and ‘how is that information encoded’ (pattern characterization).

NeurIPS Conference 2008 Conference Paper

Goal-directed decision making in prefrontal cortex: a computational framework

  • Matthew Botvinick
  • James An

Research in animal learning and behavioral neuroscience has distinguished between two forms of action control: a habit-based form, which relies on stored action values, and a goal-directed form, which forecasts and compares action outcomes based on a model of the environment. While habit-based control has been the subject of extensive computational research, the computational principles underlying goal-directed control in animals have so far received less attention. In the present paper, we advance a computational framework for goal-directed control in animals and humans. We take three empirically motivated points as founding premises: (1) Neurons in dorsolateral prefrontal cortex represent action policies, (2) Neurons in orbitofrontal cortex represent rewards, and (3) Neural computation, across domains, can be appropriately understood as performing structured probabilistic inference. On a purely computational level, the resulting account relates closely to previous work using Bayesian inference to solve Markov decision problems, but extends this work by introducing a new algorithm, which provably converges on optimal plans. On a cognitive and neuroscientific level, the theory provides a unifying framework for several different forms of goal-directed action selection, placing emphasis on a novel form, within which orbitofrontal reward representations directly drive policy selection.

YNIMG Journal 2005 Journal Article

Viewing facial expressions of pain engages cortical areas involved in the direct experience of pain

  • Matthew Botvinick
  • Amishi P. Jha
  • Lauren M. Bylsma
  • Sara A. Fabian
  • Patricia E. Solomon
  • Kenneth M. Prkachin

Recent neuroimaging and neuropsychological work has begun to shed light on how the brain responds to the viewing of facial expressions of emotion. However, one important category of facial expression that has not been studied on this level is the facial expression of pain. We investigated the neural response to pain expressions by performing functional magnetic resonance imaging (fMRI) as subjects viewed short video sequences showing faces expressing either moderate pain or, for comparison, no pain. In alternate blocks, the same subjects received both painful and non-painful thermal stimulation. Facial expressions of pain were found to engage cortical areas also engaged by the first-hand experience of pain, including anterior cingulate cortex and insula. The reported findings corroborate other work in which the neural response to witnessed pain has been examined from other perspectives. In addition, they lend support to the idea that common neural substrates are involved in representing one's own and others' affective states.

v2026.09.13