Non-linear Dynamics in Multiagent Reinforcement Learning Algorithms

Sherief Abdallah; Victor Lesser

Back to AAMAS

AAMAS 2008

Non-linear Dynamics in Multiagent Reinforcement Learning Algorithms

Conference Paper Agent and Multi-Agent Learning (Short Papers) Autonomous Agents and Multiagent Systems

PDF

Abstract

Several multiagent reinforcement learning (MARL) algorithms have been proposed to optimize agents’ decisions. Only a subset of these MARL algorithms both do not require agents to know the underlying environment and can learn a stochastic policy (a policy that chooses actions according to a probability distribution). Weighted Policy Learner (WPL) is a MARL algorithm that belongs to this subset and was shown, experimentally in previous work, to converge and outperform previous MARL algorithms belonging to the same subset. The main contribution of this paper is analyzing the dynamics of WPL and showing the effect of its non-linear nature, as opposed to previous MARL algorithms that had linear dynamics. First, we represent the WPL algorithm as a set of differential equations. We then solve the equations and show that it is consistent with experimental results reported in previous work. We finally compare the dynamics of WPL with earlier MARL algorithms and discuss the interesting differences and similarities we have discovered.

Authors

Keywords

Reinforcement Learning
Multiagent Systems
Dynamics
Convergence Analysis

Context

Venue: International Conference on Autonomous Agents and Multiagent Systems
Archive span: 2002-2025
Indexed papers: 7403
Paper id: 454841143033433239