Arrow Research search
Back to AAMAS

AAMAS 2026

The Multi-Agent Off-Switch Game

Conference Paper Research Paper Track Autonomous Agents and Multiagent Systems

Abstract

The off-switch game framework has been instrumental in understanding corrigibility—the property that AI agents should allow human oversight and intervention. In single-agent settings, uncertainty about human preferences naturally incentivizes agents to defer to human judgment. However, as AI systems increasingly operate in multi-agent environments, a crucial question arises: does corrigibility compose across multiple agents? We introduce the multi-agent off-switch game and demonstrate that individually corrigible agents can become collectively incorrigible when strategic interactions are considered. Through formal analysis and illustrative examples, we show that corrigibility is not compositional and identify conditions under which group incorrigibility emerges. Our results highlight fundamental challenges for AI safety in multi-agent settings and suggest the need for new approaches that explicitly address collective dynamics.

Authors

Keywords

  • Corrigibility
  • Off-switch game
  • Alignment
  • Multi-agent systems

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
894788104076187982
v2026.09.13