Arrow Research search
Back to AAMAS

AAMAS 2026

A Causality-Inspired Spatial-Temporal Return Decomposition Approach for Multi-Agent Reinforcement Learning

Conference Paper Extended Abstracts Autonomous Agents and Multiagent Systems

Abstract

Cooperativemulti-agentreinforcementlearning(MARL)hasachieved strongperformance, butitremainslimitedinexplaininghowagents’ decisions contribute to outcomes. This limitation is especially acute under delayed, episodic rewards, where credit must be assigned across both time and agents. We propose CAusally-inspired Spatial- Temporal return decomposition (CAST) for episodic cooperative MARL. CAST provides an interpretable decomposition while relaxing common assumptions on multi-agent reward structure. Temporally, the episodic return is expressed as a sum of per-timestep team rewards. Spatially, team rewards are modeled as general nonlinear mixtures of individual rewards rather than simple additive forms, enabling more flexible and accurate credit assignment. We show that team rewards, individual rewards, and the underlying causal relations are identifiable under our framework, yielding structural constraints that improve interpretability. Experiments on MPE and variants demonstrate state-of-the-art performance and qualitative visualizations that reveal meaningful causal structure.

Authors

Keywords

  • Multi-agent
  • Spatial-Temporal Credit Assignment
  • Causality

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
876543956521312314
v2026.09.13