Arrow Research search

Author name cluster

Nathan Monette

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
1 author row

Possible papers

2

RLC Conference 2025 Conference Paper

An Optimisation Framework for Unsupervised Environment Design

  • Nathan Monette
  • Alistair Letcher
  • Michael Beukman
  • Matthew Thomas Jackson
  • Alexander Rutherford
  • Alexander David Goldie
  • Jakob Nicolaus Foerster

For reinforcement learning agents to be deployed in high-risk settings, they must achieve a high level of robustness to unfamiliar scenarios. One method for improving robustness is unsupervised environment design (UED), a suite of methods aiming to maximise an agent's generalisability across configurations of an environment. In this work, we study UED from an optimisation perspective, providing stronger theoretical guarantees for practical settings than prior work. Whereas previous methods relied on guarantees *if* they reach convergence, our framework employs a nonconvex-strongly-concave objective for which we provide a *provably convergent* algorithm in the zero-sum setting. We empirically verify the efficacy of our method, outperforming prior methods in a number of environments with varying difficulties.

RLJ Journal 2025 Journal Article

An Optimisation Framework for Unsupervised Environment Design

  • Nathan Monette
  • Alistair Letcher
  • Michael Beukman
  • Matthew Thomas Jackson
  • Alexander Rutherford
  • Alexander David Goldie
  • Jakob Nicolaus Foerster

For reinforcement learning agents to be deployed in high-risk settings, they must achieve a high level of robustness to unfamiliar scenarios. One method for improving robustness is unsupervised environment design (UED), a suite of methods aiming to maximise an agent's generalisability across configurations of an environment. In this work, we study UED from an optimisation perspective, providing stronger theoretical guarantees for practical settings than prior work. Whereas previous methods relied on guarantees *if* they reach convergence, our framework employs a nonconvex-strongly-concave objective for which we provide a *provably convergent* algorithm in the zero-sum setting. We empirically verify the efficacy of our method, outperforming prior methods in a number of environments with varying difficulties.

v2026.09.13