Arrow Research search

Author name cluster

Thomas Gabor

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAMAS Conference 2024 Conference Paper

Quantum Circuit Design: A Reinforcement Learning Challenge

  • Philipp Altmann
  • Adelina Bärligea
  • Jonas Stein
  • Michael Kölle
  • Thomas Gabor
  • Thomy Phan
  • Claudia Linnhof-Popien

To assess the prospects of using reinforcement learning (RL) for selecting and parameterizing quantum gates to build viable circuit architectures, we introduce the quantum circuit designer (QCD). By considering quantum control a decision-making problem, we strive to profit from advanced RL exploration mechanisms to overcome the need for granular specification and hand-crafted architectures. To evaluate current state-of-the-art RL algorithms, we define generic objectives that arise from quantum architecture search and circuit optimization. Those evaluation results reveal challenges inherent to learning optimal quantum control.

ICML Conference 2023 Conference Paper

Attention-Based Recurrence for Multi-Agent Reinforcement Learning under Stochastic Partial Observability

  • Thomy Phan
  • Fabian Ritz
  • Philipp Altmann
  • Maximilian Zorn
  • Jonas Nüßlein
  • Michael Kölle 0001
  • Thomas Gabor
  • Claudia Linnhoff-Popien

Stochastic partial observability poses a major challenge for decentralized coordination in multi-agent reinforcement learning but is largely neglected in state-of-the-art research due to a strong focus on state-based centralized training for decentralized execution (CTDE) and benchmarks that lack sufficient stochasticity like StarCraft Multi-Agent Challenge (SMAC). In this paper, we propose Attention-based Embeddings of Recurrence In multi-Agent Learning (AERIAL) to approximate value functions under stochastic partial observability. AERIAL replaces the true state with a learned representation of multi-agent recurrence, considering more accurate information about decentralized agent decisions than state-based CTDE. We then introduce MessySMAC, a modified version of SMAC with stochastic observations and higher variance in initial states, to provide a more general and configurable benchmark regarding stochastic partial observability. We evaluate AERIAL in Dec-Tiger as well as in a variety of SMAC and MessySMAC maps, and compare the results with state-based CTDE. Furthermore, we evaluate the robustness of AERIAL and state-based CTDE against various stochasticity configurations in MessySMAC.

AAMAS Conference 2023 Conference Paper

Attention-Based Recurrency for Multi-Agent Reinforcement Learning under State Uncertainty

  • Thomy Phan
  • Fabian Ritz
  • Jonas Nüßlein
  • Michael Kölle
  • Thomas Gabor
  • Claudia Linnhoff-Popien

State uncertainty poses a major challenge for decentralized coordination. However, state uncertainty is largely neglected in multiagent reinforcement learning research due to a strong focus on state-based centralized training for decentralized execution (CTDE) and benchmarks that lack sufficient stochasticity like StarCraft Multi-Agent Challenge (SMAC). In this work, we propose Attentionbased Embeddings of Recurrence In multi-Agent Learning (AERIAL) to approximate value functions under agent-wise state uncertainty. AERIAL uses a learned representation of multi-agent recurrence, considering more accurate information about decentralized agent decisions than state-based CTDE. We then introduce MessySMAC, a modified version of SMAC with stochastic observations and higher variance in initial states, to provide a more general and configurable benchmark. We evaluate AERIAL in a variety of MessySMAC maps, and compare the results with state-based CTDE.

AAAI Conference 2021 Conference Paper

Resilient Multi-Agent Reinforcement Learning with Adversarial Value Decomposition

  • Thomy Phan
  • Lenz Belzner
  • Thomas Gabor
  • Andreas Sedlmeier
  • Fabian Ritz
  • Claudia Linnhoff-Popien

We focus on resilience in cooperative multi-agent systems, where agents can change their behavior due to udpates or failures of hardware and software components. Current stateof-the-art approaches to cooperative multi-agent reinforcement learning (MARL) have either focused on idealized settings without any changes or on very specialized scenarios, where the number of changing agents is fixed, e. g. , in extreme cases with only one productive agent. Therefore, we propose Resilient Adversarial value Decomposition with Antagonist-Ratios (RADAR). RADAR offers a value decomposition scheme to train competing teams of varying size for improved resilience against arbitrary agent changes. We evaluate RADAR in two cooperative multi-agent domains and show that RADAR achieves better worst case performance w. r. t. arbitrary agent changes than state-of-the-art MARL.

NeurIPS Conference 2021 Conference Paper

VAST: Value Function Factorization with Variable Agent Sub-Teams

  • Thomy Phan
  • Fabian Ritz
  • Lenz Belzner
  • Philipp Altmann
  • Thomas Gabor
  • Claudia Linnhoff-Popien

Value function factorization (VFF) is a popular approach to cooperative multi-agent reinforcement learning in order to learn local value functions from global rewards. However, state-of-the-art VFF is limited to a handful of agents in most domains. We hypothesize that this is due to the flat factorization scheme, where the VFF operator becomes a performance bottleneck with an increasing number of agents. Therefore, we propose VFF with variable agent sub-teams (VAST). VAST approximates a factorization for sub-teams which can be defined in an arbitrary way and vary over time, e. g. , to adapt to different situations. The sub-team values are then linearly decomposed for all sub-team members. Thus, VAST can learn on a more focused and compact input representation of the original VFF operator. We evaluate VAST in three multi-agent domains and show that VAST can significantly outperform state-of-the-art VFF, when the number of agents is sufficiently large.

IJCAI Conference 2019 Conference Paper

Adaptive Thompson Sampling Stacks for Memory Bounded Open-Loop Planning

  • Thomy Phan
  • Thomas Gabor
  • Robert Müller
  • Christoph Roch
  • Claudia Linnhoff-Popien

We propose Stable Yet Memory Bounded Open-Loop (SYMBOL) planning, a general memory bounded approach to partially observable open-loop planning. SYMBOL maintains an adaptive stack of Thompson Sampling bandits, whose size is bounded by the planning horizon and can be automatically adapted according to the underlying domain without any prior domain knowledge beyond a generative model. We empirically test SYMBOL in four large POMDP benchmark problems to demonstrate its effectiveness and robustness w. r. t. the choice of hyperparameters and evaluate its adaptive memory consumption. We also compare its performance with other open-loop planning algorithms and POMCP.

AAMAS Conference 2019 Conference Paper

Distributed Policy Iteration for Scalable Approximation of Cooperative Multi-Agent Policies

  • Thomy Phan
  • Kyrill Schmid
  • Lenz Belzner
  • Thomas Gabor
  • Sebastian Feld
  • Claudia Linnhoff-Popien

We propose Strong Emergent Policy (STEP) approximation, a scalable approach to learn strong decentralized policies for cooperative MAS with a distributed variant of policy iteration. For that, we use function approximation to learn from action recommendations of a decentralized multi-agent planning algorithm. STEP combines decentralized multi-agent planning with centralized learning, only requiring a generative model for distributed black box optimization. We experimentally evaluate STEP in two challenging and stochastic domains with large state and joint action spaces and show that STEP is able to learn stronger policies than standard multi-agent reinforcement learning algorithms, when combining multi-agent open-loop planning with centralized function approximation. The learned policies can be reintegrated into the multi-agent planning process to further improve performance.

IJCAI Conference 2019 Conference Paper

Subgoal-Based Temporal Abstraction in Monte-Carlo Tree Search

  • Thomas Gabor
  • Jan Peter
  • Thomy Phan
  • Christian Meyer
  • Claudia Linnhoff-Popien

We propose an approach to general subgoal-based temporal abstraction in MCTS. Our approach approximates a set of available macro-actions locally for each state only requiring a generative model and a subgoal predicate. For that, we modify the expansion step of MCTS to automatically discover and optimize macro-actions that lead to subgoals. We empirically evaluate the effectiveness, computational efficiency and robustness of our approach w. r. t. different parameter settings in two benchmark domains and compare the results to standard MCTS without temporal abstraction.

AAMAS Conference 2018 Conference Paper

Leveraging Statistical Multi-Agent Online Planning with Emergent Value Function Approximation

  • Thomy Phan
  • Lenz Belzner
  • Thomas Gabor
  • Kyrill Schmid

Making decisions is a great challenge in distributed autonomous environments due to enormous state spaces and uncertainty. Many online planning algorithms rely on statistical sampling to avoid searching the whole state space, while still being able to make acceptable decisions. However, planning often has to be performed under strict computational constraints making online planning in multi-agent systems highly limited, which could lead to poor system performance, especially in stochastic domains. In this paper, we propose Emergent Value function Approximation for Distributed Environments (EVADE), an approach to integrate global experience into multi-agent online planning in stochastic domains to consider global effects during local planning. For this purpose, a value function is approximated online based on the emergent system behaviour by using methods of reinforcement learning. We empirically evaluated EVADE with two statistical multi-agent online planning algorithms in a highly complex and stochastic smart factory environment, where multiple agents need to process various items at a shared set of machines. Our experiments show that EVADE can effectively improve the performance of multiagent online planning while offering efficiency w. r. t. the breadth and depth of the planning process.

v2026.09.13