Arrow Research search

Author name cluster

Cameron Foale

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
1 author row

Possible papers

6

KER Journal 2025 Journal Article

An empirical investigation of value-based multi-objective reinforcement learning for stochastic environments

  • Kewen Ding
  • Peter Vamplew
  • Cameron Foale
  • Richard Dazeley

Abstract One common approach to solve multi-objective reinforcement learning (MORL) problems is to extend conventional Q-learning by using vector Q-values in combination with a utility function. However issues can arise with this approach in the context of stochastic environments, particularly when optimising for the scalarised expected reward (SER) criterion. This paper extends prior research, providing a detailed examination of the factors influencing the frequency with which value-based MORL Q-learning algorithms learn the SER-optimal policy for an environment with stochastic state transitions. We empirically examine several variations of the core multi-objective Q-learning algorithm as well as reward engineering approaches and demonstrate the limitations of these methods. In particular, we highlight the critical impact of the noisy Q-value estimates issue on the stability and convergence of these algorithms.

AAMAS Conference 2024 Conference Paper

Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning

  • Peter Vamplew
  • Cameron Foale
  • Conor F. Hayes
  • Patrick Mannion
  • Enda Howley
  • Richard Dazeley
  • Scott Johnson
  • Johan Källström

Research in multi-objective reinforcement learning (MORL) has introduced the utility-based paradigm, which makes use of both environmental rewards and a function that defines the utility derived by the user from those rewards. In this paper we extend this paradigm to the context of single-objective reinforcement learning (RL), and outline multiple potential benefits including the ability to perform multi-policy learning across tasks relating to uncertain objectives, risk-aware RL, discounting, and safe RL. We also examine the algorithmic implications of adopting a utility-based approach.

AAMAS Conference 2023 Conference Paper

Scalar Reward is Not Enough

  • Peter Vamplew
  • Benjamin J. Smith
  • Johan Källström
  • Gabriel Ramos
  • Roxana Rădulescu
  • Diederik M. Roijers
  • Conor F. Hayes
  • Friedrik Hentz

Silver et al. [14] posit that scalar reward maximisation is sufficient to underpin all intelligence and provides a suitable basis for artificial general intelligence (AGI). This extended abstract summarises the counter-argument from our JAAMAS paper[19].

JAAMAS Journal 2022 Journal Article

Scalar reward is not enough: a response to Silver, Singh, Precup and Sutton (2021)

  • Peter Vamplew
  • Benjamin J. Smith
  • Cameron Foale

Abstract The recent paper “Reward is Enough” by Silver, Singh, Precup and Sutton posits that the concept of reward maximisation is sufficient to underpin all intelligence, both natural and artificial, and provides a suitable basis for the creation of artificial general intelligence. We contest the underlying assumption of Silver et al. that such reward can be scalar-valued. In this paper we explain why scalar rewards are insufficient to account for some aspects of both biological and computational intelligence, and argue in favour of explicitly multi-objective models of reward maximisation. Furthermore, we contend that even if scalar reward functions can trigger intelligent behaviour in specific cases, this type of reward is insufficient for the development of human-aligned artificial general intelligence due to unacceptable risks of unsafe or unethical behaviour.

AIJ Journal 2021 Journal Article

Levels of explainable artificial intelligence for human-aligned conversational explanations

  • Richard Dazeley
  • Peter Vamplew
  • Cameron Foale
  • Charlotte Young
  • Sunil Aryal
  • Francisco Cruz

Over the last few years there has been rapid research growth into eXplainable Artificial Intelligence (XAI) and the closely aligned Interpretable Machine Learning (IML). Drivers for this growth include recent legislative changes and increased investments by industry and governments, along with increased concern from the general public. People are affected by autonomous decisions every day and the public need to understand the decision-making process to accept the outcomes. However, the vast majority of the applications of XAI/IML are focused on providing low-level ‘narrow’ explanations of how an individual decision was reached based on a particular datum. While important, these explanations rarely provide insights into an agent's: beliefs and motivations; hypotheses of other (human, animal or AI) agents' intentions; interpretation of external cultural expectations; or, processes used to generate its own explanation. Yet all of these factors, we propose, are essential to providing the explanatory depth that people require to accept and trust the AI's decision-making. This paper aims to define levels of explanation and describe how they can be integrated to create a human-aligned conversational explanation system. In so doing, this paper will survey current approaches and discuss the integration of different technologies to achieve these levels with Broad eXplainable Artificial Intelligence (Broad-XAI), and thereby move towards high-level ‘strong’ explanations.

EAAI Journal 2021 Journal Article

Potential-based multiobjective reinforcement learning approaches to low-impact agents for AI safety

  • Peter Vamplew
  • Cameron Foale
  • Richard Dazeley
  • Adam Bignold

The concept of impact-minimisation has previously been proposed as an approach to addressing the safety concerns that can arise from utility-maximising agents. An impact-minimising agent takes into account the potential impact of its actions on the state of the environment when selecting actions, so as to avoid unacceptable side-effects. This paper proposes and empirically evaluates an implementation of impact-minimisation within the framework of multiobjective reinforcement learning. The key contributions are a novel potential-based approach to specifying a measure of impact, and an examination of a variety of non-linear action-selection operators so as to achieve an acceptable trade-off between achieving the agent’s primary task and minimising environmental impact. These experiments also highlight a previously unreported issue with noisy estimates for multiobjective agents using non-linear action-selection, which has broader implications for the application of multiobjective reinforcement learning.

v2026.09.13