AAMAS 2026
Token-level Advantage Policy Optimization from Negative Feedback in Multi-Turn Agents
Abstract
Trainingmulti-turnagentsforcomplextasksischallengedbysparse rewards. Existing methods are inefficient: they either learn exclusively from successes, discarding valuable failure data, or require rigid win-loss pairs, limiting data utilization. We propose Tokenlevel Advantage Policy Optimization (TAPO), a flexible, pair-free methodthatleveragesalltrajectories. TAPOtranslatesatrajectory’s terminal reward into token-level advantages, effectively reinforcing the entire sequence of actions in successful trajectories while penalizing those in failed ones. Furthermore, TAPO concentrates updates on high-entropy tokens, which represent pivotal moments of model uncertainty and are thus crucial for efficient exploration and policy improvement. As a post-training optimization, TAPO substantially boostsabaselineSFTagent’saveragescorefrom74. 2to89. 4(+20. 5% relative) on three challenging multi-turn benchmarks, outperformingRFTandnegative-awarebaselinesanddemonstratingconsistent gains in both seen and unseen settings. The code is available at https: //github. com/Sunrepe/TAPO.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 794473401559680917