AAMAS Conference 2026 Conference Paper
Autonomous Vehicles need Social Awareness to Find Optima in Multi-agent Reinforcement Learning Routing Games
- Anastasia Psarou
- Łukasz Gorczyca
- Dominik Gawel
- Rafal Kucharski
Previous work has shown that when multiple selfish Autonomous Vehicles(AVs)simultaneouslylearnoptimalroutingstrategiesusing Multi-Agent Reinforcement Learning (MARL), they may require a significant amount of time to converge to the optimal solution, equivalent to years of real-world commuting. We demonstrate that moving beyond the selfish component in the reward significantly relieves this issue. In particular, we introduce a reward signal based on the marginal cost matrix, which quantifies the impact of each individual action (route-choice) on the system (total travel time). This formulation reduces training time and improves convergence reliability. Experiments on both a toy network and the real-world Saint-Arnoult network show that the proposed reward improves individual and system travel times over the selfish reward baseline, and in thetoy network, enablesagents to reachthe optimal solution faster, indicating that incorporating social awareness (i. e. , including marginal costs in routing decisions) can enhance both system-wide and individual outcomes in future urban systems with AVs.