Arrow Research search

Author name cluster

Adrian Agogino

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
1 author row

Possible papers

16

AAMAS Conference 2013 Conference Paper

Addressing Hard Constraints in the Air Traffic Problem through Partitioning and Difference Rewards

  • William Curran
  • Adrian Agogino
  • Kagan Tumer

In the US alone, weather hazards and airport congestion cause thousands of hours of delay, costing billions of dollars annually. The task of managing delay may be modeled as a multiagent congestion problem with tightly coupled agents who collectively impact the system. Reward shaping has been effective at reducing noise caused by agent interaction and improving learning in soft constraint problems. We extend those results to hard constraints that cannot be easily learned, and must be algorithmically enforced. We present an agent partitioning algorithm in conjunction with reward shaping to simplify the learning domain. Our results show that a partitioning of the agents using system features leads to up to a 1000x speed up over the straight reward shaping approach, as well as up to a 30% improvement in performance over a greedy scheduling solution, corresponding to hundreds of hours of delay saved in a single day.

AAMAS Conference 2013 Conference Paper

CLEAN Rewards for Improving Multiagent Coordination in the Presence of Exploration

  • Chris HolmesParker
  • Adrian Agogino
  • Kagan Tumer

In cooperative multiagent systems, coordinating the jointactions of agents is difficult. One of the fundamental difficulties in such multiagent systems is the slow learning process where an agent may not only need to learn how to behave in a complex environment, but may also need to account for the actions of the other learning agents. Here, the inability of agents to distinguish the true environmental dynamics from those caused by the stochastic exploratory actions of other agents creates noise on each agent’s reward signal. To address this, we introduce Coordinated Learning without Exploratory Action Noise (CLEAN) rewards, which are agent-specific shaped rewards that effectively remove such learning noise from each agent’s reward signal. We demonstrate their performance with up to 1000 agents in a standard congestion problem.

AAMAS Conference 2013 Conference Paper

Exploiting Structure and Utilizing Agent-Centric Rewards to Promote Coordination in Large Multiagent Systems

  • Chris HolmesParker
  • Adrian Agogino
  • Kagan Tumer

A goal within the field of multiagent systems is to achieve scaling to large systems involving hundreds or thousands of agents. In such systems the communication requirements for agents as well as the individual agents’ ability to make decisions both play critical roles in performance. We take an incremental step towards improving scalability in such systems by introducing a novel algorithm that conglomerates three well-known existing techniques to address both agent communication requirements as well as decision making within large multiagent systems. In particular, we couple a Factored-Action Factored Markov Decision Process (FA-FMDP) framework which exploits problem structure and establishes localized rewards for agents (reducing communication requirements) with reinforcement learning using agent-centric difference rewards which addresses agent decision making and promotes coordination by addressing the structural credit assignment problem. We demonstrate our algorithms performance compared to two other popular reward techniques (global, local) with up to 10, 000 agents.

AAMAS Conference 2013 Conference Paper

Learning to Control Complex Tensegrity Robots

  • Atil Iscen (Oregon State University, USA)
  • Adrian Agogino
  • Vytas Sun Spiral
  • Kagan Tumer

Tensegrity robots are based on the idea of tensegrity structures that provides many advantages critical to robotics such as being lightweight and impact tolerant. Unfortunately tensegrity robots are hard to control due to overall complexity. We use multiagent learning to learn controls of a ball-shaped tensegrity with 6 rods and 24 cables. Our simulation results show that multiagent learning can be used to learn an efficient rolling behavior and test its robustness to actuation noise.

AAMAS Conference 2011 Conference Paper

Agent-Based Resource Allocation in Dynamically Formed CubeSat Constellations

  • Chris HolmesParker
  • Adrian Agogino

In the near future, there is potential for a tremendous expansion in the number of Earth-orbiting CubeSats, due to reduced cost associated with platform standardization, availability of standardized parts for CubeSats, and reduced launching costs due to improved packaging methods and lower cost launchers. However, software algorithms capable of efficiently coordinating CubeSats have not kept up with their hardware gains, making it likely that these CubSats will be severely underutilized. Fortunately, these coordination issues can be addressed with multiagent algorithms. In this paper, we show how a multiagent system can be used to address the particular problem of how a third party should bid for use of existing Earth-observing CubeSats so that it can achieve optical coverage over a key geographic region of interest. In this model, an agent is assigned to every CubeSat from which observations may be purchased, and agents must decide how much to offer for these services. We address this problem by having agents use reinforcement learning algorithms with agent-specific shaped rewards. The results show an eight fold improvement over a simple strawman allocation algorithm and a two fold improvement over a multiagent system using standard reward functions.

IS Journal 2009 Journal Article

Improving Air Traffic Management with a Learning Multiagent System

  • Kagan Tumer
  • Adrian Agogino

A fundamental challenge facing the aerospace industry is efficient, safe, and reliable air traffic management (ATM). On a typical day, more than 40, 000 commercial flights operate in US airspace, and the number of flights is increasing rapidly. This paper shows how learning multiagent system helps improve ATM.

AAMAS Conference 2008 Conference Paper

Aligning social welfare and agent preferences to alleviate traffic congestion

  • Kagan Tumer
  • Zach Welch
  • Adrian Agogino

Multiagent coordination algorithms provide unique insights into the challenging problem of alleviating traffic congestion. What is particularly interesting in this class of problem is that no individual action (e. g. , leave at a given time) is intrinsically “bad” but that combinations of actions among agents lead to undesirable outcomes. As a consequence, agents need to learn how to coordinate their actions with those of other agents, rather than learn a particular set of ”good” actions. In general, the traffic problem can be approached from two distinct perspectives: (i) from a city manager’s point of view, where the aim is to optimize a city wide objective function (e. g. , minimize total city wide delays), and (ii) from the individual driver’s point of view, where each driver is aiming to optimize a personal objective function (e. g. , a“timeliness”function that minimizes the difference desired and actual arrival times at a destination). In many cases, these two objective functions are at odds with one another, where drivers aiming to optimize their own objectives yield to congestion and poor values of city objective functions. In this paper we present an objective shaping approach to both types of problems and study the system behavior that arises from the drivers’ choices. We first show a topdown approach that provides incentives to drivers and leads to good values of the city manager’s objective function. We then present a bottom-up approach that shows that drivers aiming to optimize their own personal timeliness objective lead to poor performance with respect to a city manager’s objective function. Finally, we present the intriguing result that drivers that aim to optimize a modified version of their own timeliness function not only perform well in terms of the city manager’s objective function, but also perform better with respect to their own original timeliness functions.

AAMAS Conference 2008 Conference Paper

Regulating Air Traffic Flow with Coupled Agents

  • Adrian Agogino
  • Kagan Tumer

The ability to provide flexible, automated management of air traffic is critical to meeting the ever increasing needs of the next generation air transportation systems. This problem is particularly complex as it requires the integration of many factors including, updated information (e. g. , changing weather info), conflicting priorities (e. g. , different airlines), limited resources (e. g. , air traffic controllers) and very heavy traffic volume (e. g. , over 40, 000 daily flights over the US airspace). Furthermore, because the Federal Flight Administration will not accept black-box solutions, algorithmic improvements need to be consistent with current operating practices and provide explanations for each new decision. Unfortunately current methods provide neither flexibility for future upgrades, nor high enough performance in complex coupled air traffic flow problems. This paper extends agent-based methods for controlling air traffic flow to more realistic domains that have coupled flow patterns and need to be controlled through a variety of mechanisms. First, we explore an agent control structure that allows agents to control air traffic flow through one of three mechanisms (miles in trail, ground delays and rerouting). Second, we explore a new agent learning algorithm that can efficiently handle coupled flow patterns. We then test this agent solution on a series of congestion problems, showing that it is flexible enough to achieve high performance with different control mechanisms. In addition the results show that the new solution is able to achieve up to a 20% increase in performance over previous methods that did not account for the agent coupling.

AAMAS Conference 2007 Conference Paper

Distributed Agent-Based Air Traffic Flow Management

  • Kagan Tumer
  • Adrian Agogino

Air traffic flow management is one of the fundamental challenges facing the Federal Aviation Administration (FAA) today. The FAA estimates that in 2005 alone, there were over 322, 000 hours of delays at a cost to the industry in excess of three billion dollars. Finding reliable and adaptive solutions to the flow management problem is of paramount importance if the Next Generation Air Transportation Systems are to achieve the stated goal of accommodating three times the current traffic volume. This problem is particularly complex as it requires the integration and/or coordination of many factors including: new data (e. g. , changing weather info), potentially conflicting priorities (e. g. , different airlines), limited resources (e. g. , air traffic controllers) and very heavy traffic volume (e. g. , over 40, 000 flights over the US airspace).

AAAI Conference 2006 Conference Paper

QUICR-Learning for Multi-Agent Coordination

  • Adrian Agogino

Coordinating multiple agents that need to perform a sequence of actions to maximize a system level reward requires solving two distinct credit assignment problems. First, credit must be assigned for an action taken at time step t that results in a reward at time step t > t. Second, credit must be assigned for the contribution of agent i to the overall system performance. The first credit assignment problem is typically addressed with temporal difference methods such as Q-learning. The second credit assignment problem is typically addressed by creating custom reward functions. To address both credit assignment problems simultaneously, we propose the “Q Updates with Immediate Counterfactual Rewards-learning” (QUICR-learning) designed to improve both the convergence properties and performance of Q-learning in large multi-agent problems. QUICR-learning is based on previous work on single-time-step counterfactual rewards described by the collectives framework. Results on a traffic congestion problem shows that QUICR-learning is significantly better than a Qlearner using collectives-based (single-time-step counterfactual) rewards. In addition QUICR-learning provides significant gains over conventional and local Q-learning. Additional results on a multi-agent grid-world problem show that the improvements due to QUICR-learning are not domain specific and can provide up to a ten fold increase in performance over existing methods.

v2026.09.13