Arrow Research search
Back to AAMAS

AAMAS 2026

Enhanced Deep Q-Learning with Gaussian Mixtures

Conference Paper Extended Abstracts Autonomous Agents and Multiagent Systems

Abstract

Value-basedreinforcementlearningmethods, likeDeepQ-Networks (DQNs), typically estimate returns by minimizing the mean squared error between predicted and target values. From a Bayesian standpoint, this procedure implicitly assumes that returns follow a unimodalGaussiandistribution, withparameterslearnedviamaximum likelihood estimation. However, this assumption can be limiting in environments characterized by high stochasticity or complex reward dynamics, where capturing uncertainty and multi-modality in the return distribution is critical for robust decision-making. We propose Gaussian Mixture Q-Networks (GQN), a novel extension of Q-learning that models return distribution as a mixture of Gaussians. Architecturally, GQN can be interpreted as a mixtureof-experts Q-learning algorithm, where each Gaussian component acts as an expert head and mixture weights are adaptively updated via temporal-difference responsibilities inspired by Expectation–Maximization. We evaluate GQN on the Atari benchmark suite and observe improvements in both learning stability and final performance compared to standard DQN baselines.

Authors

Keywords

  • Deep Q-Learning
  • Gaussian Mixture Model

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
1070428544185070026
v2026.09.13