Arrow Research search
Back to AAMAS

AAMAS 2026

Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models

Conference Paper Extended Abstracts Autonomous Agents and Multiagent Systems

Abstract

Policy-Space Response Oracles (PSRO) have enabled the computation of approximate Nash equilibria in complex games. However, standard implementations rely on Deep Reinforcement Learning oracles, producing "black-box" neural network policies that are opaque, difficult to verify, and sample-inefficient. We introduce Code-Space Response Oracles (CSRO), a framework that tasks a Large Language Model (LLM) to synthesize code policies. CSRO reframes best-response computation as a code generation task, producing policies as executable, human-readable Python code. We demonstrate that CSRO, particularly when augmented with evolutionary refinement (AlphaEvolve), achieves performance competitive with baselines while offering superior interpretability and leveraging the LLM’s pretraining knowledge.

Authors

Keywords

  • Multi-Agent Reinforcement Learning
  • Game Theory
  • Large Language Models
  • Code Synthesis

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
877924371818703017
v2026.09.13