AAMAS 2026
Enhanced Deep Q-Learning with Gaussian Mixtures
Abstract
Value-basedreinforcementlearningmethods, likeDeepQ-Networks (DQNs), typically estimate returns by minimizing the mean squared error between predicted and target values. From a Bayesian standpoint, this procedure implicitly assumes that returns follow a unimodalGaussiandistribution, withparameterslearnedviamaximum likelihood estimation. However, this assumption can be limiting in environments characterized by high stochasticity or complex reward dynamics, where capturing uncertainty and multi-modality in the return distribution is critical for robust decision-making. We propose Gaussian Mixture Q-Networks (GQN), a novel extension of Q-learning that models return distribution as a mixture of Gaussians. Architecturally, GQN can be interpreted as a mixtureof-experts Q-learning algorithm, where each Gaussian component acts as an expert head and mixture weights are adaptively updated via temporal-difference responsibilities inspired by Expectation–Maximization. We evaluate GQN on the Atari benchmark suite and observe improvements in both learning stability and final performance compared to standard DQN baselines.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 1070428544185070026