Arrow Research search

Author name cluster

Jonathan J. Hunt

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
2 author rows

Possible papers

5

ICLR Conference 2023 Conference Paper

Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

  • Zhendong Wang 0005
  • Jonathan J. Hunt
  • Mingyuan Zhou

Offline reinforcement learning (RL), which aims to learn an optimal policy using a previously collected static dataset, is an important paradigm of RL. Standard RL methods often perform poorly in this regime due to the function approximation errors on out-of-distribution actions. While a variety of regularization methods have been proposed to mitigate this issue, they are often constrained by policy classes with limited expressiveness that can lead to highly suboptimal solutions. In this paper, we propose representing the policy as a diffusion model, a recent class of highly-expressive deep generative models. We introduce Diffusion Q-learning (Diffusion-QL) that utilizes a conditional diffusion model to represent the policy. In our approach, we learn an action-value function and we add a term maximizing action-values into the training loss of the conditional diffusion model, which results in a loss that seeks optimal actions that are near the behavior policy. We show the expressiveness of the diffusion model-based policy, and the coupling of the behavior cloning and policy improvement under the diffusion model both contribute to the outstanding performance of Diffusion-QL. We illustrate the superiority of our method compared to prior works in a simple 2D bandit example with a multimodal behavior policy. We then show that our method can achieve state-of-the-art performance on the majority of the D4RL benchmark tasks.

ICLR Conference 2023 Conference Paper

Hyperbolic Deep Reinforcement Learning

  • Edoardo Cetin
  • Ben Chamberlain 0001
  • Michael M. Bronstein
  • Jonathan J. Hunt

In deep reinforcement learning (RL), useful information about the state is inherently tied to its possible future successors. Consequently, encoding features that capture the hierarchical relationships between states into the model's latent representations is often conducive to recovering effective policies. In this work, we study a new class of deep RL algorithms that promote encoding such relationships by using hyperbolic space to model latent representations. However, we find that a naive application of existing methodology from the hyperbolic deep learning literature leads to fatal instabilities due to the non-stationarity and variance characterizing common gradient estimators in RL. Hence, we design a new general method that directly addresses such optimization challenges and enables stable end-to-end learning with deep hyperbolic representations. We empirically validate our framework by applying it to popular on-policy and off-policy RL algorithms on the Procgen and Atari 100K benchmarks, attaining near universal performance and generalization benefits. Given its natural fit, we hope this work will inspire future RL research to consider hyperbolic representations as a standard tool.

ICML Conference 2019 Conference Paper

Composing Entropic Policies using Divergence Correction

  • Jonathan J. Hunt
  • AndrĂ© Barreto 0001
  • Timothy P. Lillicrap
  • Nicolas Heess

Composing skills mastered in one task to solve novel tasks promises dramatic improvements in the data efficiency of reinforcement learning. Here, we analyze two recent works composing behaviors represented in the form of action-value functions and show that they perform poorly in some situations. As part of this analysis, we extend an important generalization of policy improvement to the maximum entropy framework and introduce an algorithm for the practical implementation of successor features in continuous action spaces. Then we propose a novel approach which addresses the failure cases of prior work and, in principle, recovers the optimal policy during transfer. This method works by explicitly learning the (discounted, future) divergence between base policies. We study this approach in the tabular case and on non-trivial continuous control problems with compositional structure and show that it outperforms or matches existing methods across all tasks considered.

YNIMG Journal 2014 Journal Article

Stripe-rearing changes multiple aspects of the structure of primary visual cortex

  • Nicholas J. Hughes
  • Jonathan J. Hunt
  • Shaun L. Cloherty
  • Michael R. Ibbotson
  • Frank Sengpiel
  • Geoffrey J. Goodhill

An important example of brain plasticity is the change in the structure of the orientation map in mammalian primary visual cortex in response to a visual environment consisting of stripes of one orientation. In principle there are many different ways in which the structure of a normal map could change to accommodate increased preference for one orientation. However, until now these changes have been characterised only by the relative sizes of the areas of primary visual cortex representing different orientations. Here we extend to the stripe-reared case a recently proposed Bayesian method for reconstructing orientation maps from intrinsic signal optical imaging data. We first formulated a suitable prior for the stripe-reared case, and developed an efficient method for maximising the marginal likelihood of the model in order to determine the optimal parameters. We then applied this to a set of orientation maps from normal and stripe-reared cats. This analysis revealed that several parameters of overall map structure, specifically the difference between wavelength, scaling and mean of the two vector components of maps, changed in response to stripe-rearing, which together give a more nuanced assessment of the effect of rearing condition on map structure than previous measures. Overall this work expands our understanding of the effects of the environment on brain structure.

YNIMG Journal 2009 Journal Article

Natural scene statistics and the structure of orientation maps in the visual cortex

  • Jonathan J. Hunt
  • Clare E. Giacomantonio
  • Huajin Tang
  • Duncan Mortimer
  • Sajjida Jaffer
  • Vasily Vorobyov
  • Geoffery Ericksson
  • Frank Sengpiel

Visual activity after eye-opening influences feature map structure in primary visual cortex (V1). For instance, rearing cats in an environment of stripes of one orientation yields an over-representation of that orientation in V1. However, whether such changes also affect the higher-order statistics of orientation maps is unknown. A statistical bias of orientation maps in normally raised animals is that the probability of the angular difference in orientation preference between each pair of points in the cortex depends on the angle of the line joining those points relative to a fixed but arbitrary set of axes. Natural images show an analogous statistical bias; however, whether this drives the development of comparable structure in V1 is unknown. We examined these statistics for normal, stripe-reared and dark-reared cats, and found that the biases present were not consistently related to those present in the input, or to genetic relationships. We compared these results with two computational models of orientation map development, an analytical model and a Hebbian model. The analytical model failed to reproduce the experimentally observed statistics. In the Hebbian model, while orientation difference statistics could be strongly driven by the input, statistics similar to those seen in experimental maps arose only when symmetry breaking was allowed to occur spontaneously. These results suggest that these statistical biases of orientation maps arise primarily spontaneously, rather than being governed by either input statistics or genetic mechanisms.

v2026.09.13