Arrow Research search
Back to AAMAS

AAMAS 2026

Token-level Advantage Policy Optimization from Negative Feedback in Multi-Turn Agents

Conference Paper Research Paper Track Autonomous Agents and Multiagent Systems

Abstract

Trainingmulti-turnagentsforcomplextasksischallengedbysparse rewards. Existing methods are inefficient: they either learn exclusively from successes, discarding valuable failure data, or require rigid win-loss pairs, limiting data utilization. We propose Tokenlevel Advantage Policy Optimization (TAPO), a flexible, pair-free methodthatleveragesalltrajectories. TAPOtranslatesatrajectory’s terminal reward into token-level advantages, effectively reinforcing the entire sequence of actions in successful trajectories while penalizing those in failed ones. Furthermore, TAPO concentrates updates on high-entropy tokens, which represent pivotal moments of model uncertainty and are thus crucial for efficient exploration and policy improvement. As a post-training optimization, TAPO substantially boostsabaselineSFTagent’saveragescorefrom74. 2to89. 4(+20. 5% relative) on three challenging multi-turn benchmarks, outperformingRFTandnegative-awarebaselinesanddemonstratingconsistent gains in both seen and unseen settings. The code is available at https: //github. com/Sunrepe/TAPO.

Authors

Keywords

  • Multi-turn agents
  • Token-level Policy optimization
  • Entropy-guided Learning
  • GAAI

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
794473401559680917
v2026.09.13