AAMAS Conference 2026 Conference Paper
Quality-Diversity for Multi-Agent Reinforcement Learning
- Hao Chen
- Pengyi Li
- Bin Zhang
- Hu Fu
- Zhiwei Xu
- Ce Zhang
- Xinyue Lu
- Guoliang Fan
Quality–diversity optimization (QD) in multi-agent reinforcement learning (MARL) aims to evolve a population of team policies that are both high-performing and behaviorally diverse, enabling effective coordination in complex cooperative tasks. However, existing QD approaches often depend on random exploration to encourage diversity, resulting in unstable learning and limited coverage in high-dimensional environments. We propose MIQD, a mutualinformation–enhanced QD framework that integrates fragmentbased behavioral descriptors into the critic to capture short-term patterns and guide policy updates. Mutual information measures alignment between policy behavior and target descriptors; its steplevel decomposition yields intrinsic rewards that promote alignment at each state–action pair. Experimental results show that our method consistently outperforms strong baselines across multiple metrics, demonstrating its effectiveness in jointly enhancing policy quality and diversity.