AAMAS Conference 2026 Conference Paper
Pareto-Guided Exploration for Multi-Objective Multiagent Learning
- Gaurav Dixit
- Kagan Tumer
Cooperativemulti-objectivemultiagentsettingsrequirediscovering teams that realize diverse trade-offs across conflicting objectives. In multi-objective reinforcement learning, policies are typically optimized via scalarization of expected returns, while Pareto-based methods approximate sets of non-dominated solutions in objective space. In cooperative settings, however, these perspectives are typically treated in isolation. We introduce Multiagent Pareto- Led Exploration (MAPLE), a framework that couples preferencealignedactor-criticlearningwithParetodominance-basedselection over expected team returns. MAPLE trains policies under multiple preference vectors while preserving and reusing agents based on their contribution to non-dominated teams through a shared archive. Empirical results in cooperative continuous-control benchmarks demonstrate improved Pareto front coverage and greater policy composability across trade-offs. These findings suggest that coupling scalarized learning with Pareto-level selection provides a principled mechanism for multi-objective multiagent learning.