Arrow Research search
Back to ICRA

ICRA 2007

Adaptive Play Q-Learning with Initial Heuristic Approximation

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

The problem of an effective coordination of multiple autonomous robots is one of the most important tasks of the modern robotics. In turn, it is well known that the learning to coordinate multiple autonomous agents in a multiagent system is one of the most complex challenges of the state-of-the-art intelligent system design. Principally, this is because of the exponential growth of the environment's dimensionality with the number of learning agents. This challenge is known as "curse of dimensionality", and relates to the fact that the dimensionality of the multiagent coordination problem is exponential in the number of learning agents, because each state of the system is a joint state of all agents and each action is a joint action composed of actions of each agent. In this paper, we address this problem for the restricted class of environments known as goal-directed stochastic games with action-penalty representation. We use a single-agent problem solution as a heuristic approximation of the agents' initial preferences and, by so doing, we restrict to a great extent the space of multiagent learning. We show theoretically the correctness of such an initialization, and the results of experiments in a well-known two-robot grid world problem show that there is a significant reduction of complexity of the learning process.

Authors

Keywords

  • Robot kinematics
  • Multiagent systems
  • Stochastic processes
  • State-space methods
  • Autonomous agents
  • Intelligent systems
  • Intelligent agent
  • Intelligent robots
  • Robotics and automation
  • Orbital robotics
  • Learning Process
  • System State
  • Joint Action
  • Multiple Agents
  • Multi-agent Systems
  • Curse Of Dimensionality
  • Coordination Problems
  • Learning Agent
  • Multi-agent Reinforcement Learning
  • Learning Algorithms
  • State Space
  • Equilibrium Point
  • Heuristic Algorithm
  • Optimal Policy
  • Best Response
  • Heuristic Search
  • Goal State
  • Reward Function
  • Markov Decision Process
  • Mixed Strategy
  • Nash Equilibrium
  • Transition Function
  • Stage Of The Game
  • Pure Strategy
  • State-action Pair
  • Value Iteration
  • Joint Strategy
  • Reward Structure
  • Trajectory Optimization
  • Complex Learning

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
868462305124471685
v2026.09.13