Arrow Research search

Author name cluster

Nathan Wispinski

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
1 author row

Possible papers

3

TMLR Journal 2023 Journal Article

Adaptive patch foraging in deep reinforcement learning agents

  • Nathan Wispinski
  • Andrew Butcher
  • Kory Wallace Mathewson
  • Craig S Chapman
  • Matthew Botvinick
  • Patrick M. Pilarski

Patch foraging is one of the most heavily studied behavioral optimization challenges in biology. However, despite its importance to biological intelligence, this behavioral optimization problem is understudied in artificial intelligence research. Patch foraging is especially amenable to study given that it has a known optimal solution, which may be difficult to discover given current techniques in deep reinforcement learning. Here, we investigate deep reinforcement learning agents in an ecological patch foraging task. For the first time, we show that machine learning agents can learn to patch forage adaptively in patterns similar to biological foragers, and approach optimal patch foraging behavior when accounting for temporal discounting. Finally, we show emergent internal dynamics in these agents that resemble single-cell recordings from foraging non-human primates, which complements experimental and theoretical work on the neural mechanisms of biological foraging. This work suggests that agents interacting in complex environments with ecologically valid pressures arrive at common solutions, suggesting the emergence of foundational computations behind adaptive, intelligent behavior in both biological and artificial agents.

RLDM Conference 2015 Conference Abstract

Independent Biases in Human Decision Making from Experience Revealed by Action Dynamics

  • Nathan Wispinski
  • Christopher Madan
  • Craig Chapman

When acting in complex environments, humans often need to make decisions involving risky and ambiguous options. That is, decisions frequently involve options that can have multiple outcomes (risk), and information about those outcomes and/or their respective probabilities of occurrence can be uncertain (ambiguity). We investigated biases in a reaching task using both implicit (reaction times and reach trajecto- ries) and explicit (choices, personality inventories, and probability estimates) measures while subjects made decisions involving options for which they were given perfect information, and those for which they only had information about potential outcomes and not their associated probabilities of occurrence. However, participants were given feedback about selected options on every trial, and thus learned about ambiguous options through experience. Overall, we found that each measure revealed distinct results: probability esti- mates were relatively accurate; early choices were biased by novelty-seeking and later choices were biased toward described information; and reaction times and reaching movements were primarily driven by reward and differences in expected value, respectively. Overall, our results demonstrate that how information about options is acquired, how decisions are physically made, and the individual differences between participants are important, though often overlooked, components of learning and decision making, which can reveal important aspects about human cognitive processing when integrated. This research presents novel experi- mental data showing distinct behavioral biases at different levels of cognition during decision making which can be used to constrain plausible models of human reinforcement learning, and also shows that the use of multiple methods may provide valuable information for future human reinforcement-learning research. Poster T28*: Utility-weighted sampling in decisions from experience Falk Lieder*, UC Berkeley; Thomas Griffiths, UC Berkeley; Ming Hsu, UC Berkeley People overweight extreme events in decision-making and overestimate their frequency. Previous theoretical work has shown that this apparently irrational bias could result from utility-weighted sampling-a decision mechanism that makes rational use of limited computational resources (Lieder, Hsu, & Griffiths, 2014). Here, we show that utility-weighted sampling can emerge from a neurally plausible associative learning mechanism. Our model explains the over-weighting of extreme outcomes in repeated decisions from experience (Ludvig, Madan, & Spetch, 2014), as well as the overestimation of their frequency and the underlying memory biases (Madan, Ludvig, & Spetch, 2014). Our results support the conclusion that utility drives probability-weighting by biasing the neural simulation of potential consequences towards extreme values.

RLDM Conference 2015 Conference Abstract

Separating value from selection frequency in rapid reaching biases to visual targets

  • Craig Chapman
  • Jason Gallivan
  • Nathan Wispinski
  • James Enns

Stimuli associated with positive rewards in one task often receive preferential processing in a subsequent task, even when those associations are no longer relevant. Here we use a rapid reaching task to investigate these biases. In Experiment 1 we first replicated the learning procedure of Raymond and O’Brien (2009), for a set of arbitrary shapes that varied in value (positive, negative) and probability (20%, 80%). In a subsequent task, participants rapidly reached toward one of two shapes, except now the previously learned associations were irrelevant. As in the previous studies, we found significant reach biases toward shapes previously associated with a high probable, positive outcome. Unexpectedly, we also found a bias toward shapes previously associated with a low probable, negative outcome. Closer inspection of the learning task revealed a potential second factor that might account for these results; since a low probable negative shape was always paired with a high probable negative shape, it was selected with disproportionate frequency. To assess how selection frequency and reward value might both contribute to reaching biases we performed a second experiment. The results of this experiment at a group level replicated the reach-bias toward positively rewarding stimuli, but also revealed a separate bias toward stimuli that had been more frequently selected. At the level of individual participants, we observed a variety of preference profiles, with some participants biased primarily by reward value, others by frequency, and a few actually biased away from both highly rewarding and high frequency targets. These findings highlight that: (1) Rapid reaching provides a sensitive readout of preferential processing. (2) Target reward value and target selection frequency are separate sources of bias. (3) Group-level analyses in complex decision-making tasks can obscure important and varied individual differences in preference profiles.

v2026.09.13