Arrow Research search

Author name cluster

Michael M. Zavlanos

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

18 papers
2 author rows

Possible papers

18

ICML Conference 2025 Conference Paper

Distributionally Robust Multi-Agent Reinforcement Learning for Dynamic Chute Mapping

  • Guangyi Liu
  • Suzan Iloglu
  • Michael Caldara
  • Joseph W. Durham
  • Michael M. Zavlanos

In Amazon robotic warehouses, the destination-to-chute mapping problem is crucial for efficient package sorting. Often, however, this problem is complicated by uncertain and dynamic package induction rates, which can lead to increased package recirculation. To tackle this challenge, we introduce a Distributionally Robust Multi-Agent Reinforcement Learning (DRMARL) framework that learns a destination-to-chute mapping policy that is resilient to adversarial variations in induction rates. Specifically, DRMARL relies on group distributionally robust optimization (DRO) to learn a policy that performs well not only on average but also on each individual subpopulation of induction rates within the group that capture, for example, different seasonality or operation modes of the system. This approach is then combined with a novel contextual bandit-based estimator of the worst-case induction distribution for each state-action pair, significantly reducing the cost of exploration and thereby increasing the learning efficiency and scalability of our framework. Extensive simulations demonstrate that DRMARL achieves robust chute mapping in the presence of varying induction distributions, reducing package recirculation by an average of 80% in the simulation scenario.

ICML Conference 2025 Conference Paper

Enhancing Cooperative Multi-Agent Reinforcement Learning with State Modelling and Adversarial Exploration

  • Andreas Kontogiannis
  • Konstantinos Papathanasiou
  • Yi Shen 0011
  • Giorgos Stamou
  • Michael M. Zavlanos
  • George A. Vouros

Learning to cooperate in distributed partially observable environments with no communication abilities poses significant challenges for multi-agent deep reinforcement learning (MARL). This paper addresses key concerns in this domain, focusing on inferring state representations from individual agent observations and leveraging these representations to enhance agents’ exploration and collaborative task execution policies. To this end, we propose a novel state modelling framework for cooperative MARL, where agents infer meaningful belief representations of the non-observable state, with respect to optimizing their own policies, while filtering redundant and less informative joint state information. Building upon this framework, we propose the MARL SMPE$^2$ algorithm. In SMPE$^2$, agents enhance their own policy’s discriminative abilities under partial observability, explicitly by incorporating their beliefs into the policy network, and implicitly by adopting an adversarial type of exploration policies which encourages agents to discover novel, high-value states while improving the discriminative abilities of others. Experimentally, we show that SMPE$^2$ outperforms a plethora of state-of-the-art MARL algorithms in complex fully cooperative tasks from the MPE, LBF, and RWARE benchmarks.

TMLR Journal 2025 Journal Article

LAPP: Large Language Model Feedback for Preference-Driven Reinforcement Learning

  • Pingcheng Jian
  • Xiao Wei
  • Yanbaihui Liu
  • Samuel A. Moore
  • Michael M. Zavlanos
  • Boyuan Chen

We introduce Large Language Model-Assisted Preference Prediction (LAPP), a novel framework for robot learning that enables efficient, customizable, and expressive behavior acquisition with minimum human effort. Unlike prior approaches that rely heavily on reward engineering, human demonstrations, motion capture, or expensive pairwise preference labels, LAPP leverages large language models (LLMs) to automatically generate preference labels from raw state-action trajectories collected during reinforcement learning (RL). These labels are used to train an online preference predictor, which in turn guides the policy optimization process toward satisfying high-level behavioral specifications provided by humans. Our key technical contribution is the integration of LLMs into the RL feedback loop through trajectory-level preference prediction, enabling robots to acquire complex skills including subtle control over gait patterns and rhythmic timing. We evaluate LAPP on a diverse set of quadruped locomotion and dexterous manipulation tasks and show that it achieves efficient learning, higher final performance, faster adaptation, and precise control of high-level behaviors. Notably, LAPP enables robots to master highly dynamic and expressive tasks such as quadruped backflips, which remain out of reach for standard LLM-generated or handcrafted rewards. Our results highlight LAPP as a promising direction for scalable preference-driven robot learning.

NeurIPS Conference 2024 Conference Paper

Outlier-Robust Distributionally Robust Optimization via Unbalanced Optimal Transport

  • Zifan Wang
  • YI Shen
  • Michael M. Zavlanos
  • Karl H. Johansson

Distributionally Robust Optimization (DRO) accounts for uncertainty in data distributions by optimizing the model performance against the worst possible distribution within an ambiguity set. In this paper, we propose a DRO framework that relies on a new distance inspired by Unbalanced Optimal Transport (UOT). The proposed UOT distance employs a soft penalization term instead of hard constraints, enabling the construction of an ambiguity set that is more resilient to outliers. Under smoothness conditions, we establish strong duality of the proposed DRO problem. Moreover, we introduce a computationally efficient Lagrangian penalty formulation for which we show that strong duality also holds. Finally, we provide empirical results that demonstrate that our method offers improved robustness to outliers and is computationally less demanding for regression and classification tasks.

TMLR Journal 2024 Journal Article

Perception Stitching: Zero-Shot Perception Encoder Transfer for Visuomotor Robot Policies

  • Pingcheng Jian
  • Easop Lee
  • Zachary I. Bell
  • Michael M. Zavlanos
  • Boyuan Chen

Vision-based imitation learning has shown promising capabilities of endowing robots with various motion skills given visual observation. However, current visuomotor policies fail to adapt to drastic changes in their visual observations. We present Perception Stitching that enables strong zero-shot adaptation to large visual changes by directly stitching novel combinations of visual encoders. Our key idea is to enforce modularity of visual encoders by aligning the latent visual features among different visuomotor policies. Our method disentangles the perceptual knowledge with the downstream motion skills and allows the reuse of the visual encoders by directly stitching them to a policy network trained with partially different visual conditions. We evaluate our method in various simulated and real-world manipulation tasks. While baseline methods failed at all attempts, our method could achieve zero-shot success in real-world visuomotor tasks. Our quantitative and qualitative analysis of the learned features of the policy network provides more insights into the high performance of our proposed method.

ICRA Conference 2022 Conference Paper

Formal Verification of Stochastic Systems with ReLU Neural Network Controllers

  • Shiqi Sun
  • Yan Zhang 0043
  • Xusheng Luo
  • Panagiotis Vlantis
  • Miroslav Pajic
  • Michael M. Zavlanos

In this work, we address the problem of formal safety verification for stochastic cyber-physical systems (CPS) equipped with ReLU neural network (NN) controllers. Our goal is to find the set of initial states from where, with a predetermined confidence, the system will not reach an unsafe configuration within a specified time horizon. Specifically, we consider discrete-time LTI systems with Gaussian noise, which we abstract by a suitable graph. Then, we formulate a Satisfiability Modulo Convex (SMC) problem to estimate upper bounds on the transition probabilities between nodes in the graph. Using this abstraction, we propose a method to compute tight bounds on the safety probabilities of nodes in this graph, despite possible over-approximations of the transition probabilities between these nodes. Additionally, using the proposed SMC formula, we devise a heuristic method to refine the abstraction of the system in order to further improve the estimated safety bounds. Finally, we corroborate the efficacy of the proposed method with simulation results considering a robot navigation example and comparison against a state-of-the-art verification scheme.

ICRA Conference 2022 Conference Paper

Receding Horizon Tracking of an Unknown Number of Mobile Targets using a Bearings-Only Sensor

  • James D. Turner
  • James McMahon
  • Michael M. Zavlanos

Planning the motion of bearings-only sensors is critical for enabling accurate tracking of the positions of moving targets. In this paper, we demonstrate planning the observer's motion over horizons greater than one step for estimating an unknown and varying number of indistinguishable, maneuvering targets of interest using a probability hypothesis density (PHD) filter, with a Rériyi divergence reward for selecting actions. We describe approximations to make this approach computationally feasible, and we propose using Monte Carlo tree search (MCTS) to further reduce the cost. Finally, we present simulation results showing that longer planning horizons reduce the error in the estimates and that MCTS can reduce the cost of planning without sacrificing the quality of the estimates.

ICML Conference 2022 Conference Paper

Risk-Averse No-Regret Learning in Online Convex Games

  • Zifan Wang 0002
  • Yi Shen 0011
  • Michael M. Zavlanos

We consider an online stochastic game with risk-averse agents whose goal is to learn optimal decisions that minimize the risk of incurring significantly high costs. Specifically, we use the Conditional Value at Risk (CVaR) as a risk measure that the agents can estimate using bandit feedback in the form of the cost values of only their selected actions. Since the distributions of the cost functions depend on the actions of all agents that are generally unobservable, they are themselves unknown and, therefore, the CVaR values of the costs are difficult to compute. To address this challenge, we propose a new online risk-averse learning algorithm that relies on one-point zeroth-order estimation of the CVaR gradients computed using CVaR values that are estimated by appropriately sampling the cost functions. We show that this algorithm achieves sub-linear regret with high probability. We also propose two variants of this algorithm that improve performance. The first variant relies on a new sampling strategy that uses samples from the previous iteration to improve the estimation accuracy of the CVaR values. The second variant employs residual feedback that uses CVaR values from the previous iteration to reduce the variance of the CVaR gradient estimates. We theoretically analyze the convergence properties of these variants and illustrate their performance on an online market problem that we model as a Cournot game.

ICRA Conference 2021 Conference Paper

Model-Free Reinforcement Learning for Stochastic Games with Linear Temporal Logic Objectives

  • Alper Kamil Bozkurt
  • Yu Wang 0044
  • Michael M. Zavlanos
  • Miroslav Pajic

We study synthesis of control strategies from linear temporal logic (LTL) objectives in unknown environments. We model this problem as a turn-based zero-sum stochastic game between the controller and the environment, where the transition probabilities and the model topology are fully unknown. The winning condition for the controller in this game is the satisfaction of the given LTL specification, which can be captured by the acceptance condition of a deterministic Rabin automaton (DRA) directly derived from the LTL specification. We introduce a model-free reinforcement learning (RL) methodology to find a strategy that maximizes the probability of satisfying a given LTL specification when the Rabin condition of the derived DRA has a single accepting pair. We then generalize this approach to any LTL formulas, for which the Rabin accepting condition may have more than one pairs, providing a lower bound on the satisfaction probability. Finally, we show applicability of our RL method on two planning case studies.

ICRA Conference 2020 Conference Paper

Bio-Inspired Distance Estimation using the Self-Induced Acoustic Signature of a Motor-Propeller System

  • Luke Calkins
  • Joseph F. Lingevitch
  • Loy McGuire
  • Jason Geder
  • Matthew Kelly 0003
  • Michael M. Zavlanos
  • Donald Sofge
  • Daniel M. Lofaro

In this paper we propose an algorithm to actively control the distance of a motor-propeller system (MPS) to a large obstacle using data from a single microphone. The method is based upon a broadband constructive/destructive interference pattern across the audible frequency band that is present when the MPS is near an obstacle. By taking the difference between the power spectrum in the obstacle-free case and the spectrum when recording near an obstacle, a broadband oscillation with respect to frequency is revealed. The frequency of this oscillation is linearly-related to the distance from the microphone to the wall. We present both static and dynamic experiments showcasing the ability of the proposed method to estimate the distance to a wall as well as actively control it.

ICRA Conference 2020 Conference Paper

Control Synthesis from Linear Temporal Logic Specifications using Model-Free Reinforcement Learning

  • Alper Kamil Bozkurt
  • Yu Wang 0044
  • Michael M. Zavlanos
  • Miroslav Pajic

We present a reinforcement learning (RL) frame-work to synthesize a control policy from a given linear temporal logic (LTL) specification in an unknown stochastic environment that can be modeled as a Markov Decision Process (MDP). Specifically, we learn a policy that maximizes the probability of satisfying the LTL formula without learning the transition probabilities. We introduce a novel rewarding and discounting mechanism based on the LTL formula such that (i) an optimal policy maximizing the total discounted reward effectively maximizes the probabilities of satisfying LTL objectives, and (ii) a model-free RL algorithm using these rewards and discount factors is guaranteed to converge to such a policy. Finally, we illustrate the applicability of our RL-based synthesis approach on two motion planning case studies.

ICRA Conference 2020 Conference Paper

Deep Imitative Reinforcement Learning for Temporal Logic Robot Motion Planning with Noisy Semantic Observations

  • Qitong Gao
  • Miroslav Pajic
  • Michael M. Zavlanos

In this paper, we propose a Deep Imitative Q-learning (DIQL) method to synthesize control policies for mobile robots that need to satisfy Linear Temporal Logic (LTL) specifications using noisy semantic observations of their surroundings. The robot sensing error is modeled using probabilistic labels defined over the states of a Labeled Transition System (LTS) and the robot mobility is modeled using a Labeled Markov Decision Process (LMDP) with unknown transition probabilities. We use existing product-based model checkers (PMCs) as experts to guide the Q-learning algorithm to convergence. To the best of our knowledge, this is the first approach that models noise in semantic observations using probabilistic labeling functions and employs existing model checkers to provide suboptimal instructions to the Q-learning agent.

ICRA Conference 2018 Conference Paper

Distributed Intermittent Communication Control of Mobile Robot Networks Under Time-Critical Dynamic Tasks

  • Yiannis Kantaros
  • Michael M. Zavlanos

In this paper, we develop a distributed intermittent communication framework for teams of mobile robots that are responsible for accomplishing time-critical dynamic tasks and sharing the collected information with all other robots and possibly also with a user. Specifically, we consider situations where the robot communication capabilities are not sufficient to maintain reliable and connected networks while the robots move to accomplish their tasks. In this case, intermittent communication protocols are necessary that allow the robots to temporarily disconnect from the network in order to accomplish their tasks free of communication constraints. We assume that the robots can only communicate with each other when they meet at common locations in space. Our proposed distributed control framework determines offline schedules of communication events and integrates them online with task planning. The resulting paths ensure task accomplishment and exchange of information among robots infinitely often at locations that minimize a user-specified metric. Simulation results corroborate the proposed distributed control framework.

ICRA Conference 2017 Conference Paper

Distributed data gathering with buffer constraints and intermittent communication

  • Meng Guo 0002
  • Michael M. Zavlanos

We consider a team of multiple dynamical and heterogeneous robots which are deployed for gathering different types of data within a common workspace. The robots have different roles due to different capabilities: some gather data from the workspace (Type-A robots) and others receive data from Type-A robots and upload them to a data center (Type-B robots). The data-gathering tasks are specified locally to each Type-A robot as high-level Linear Temporal Logic (LTL) formulas. All robots have a limited buffer to store the data. Thus the data gathered by Type-A robots should be transferred to Type-B robots before the buffers overflow, respecting at the same time limited communication range for all robots. The main contribution of this work is a distributed task coordination and intermittent meeting scheme that guarantees the satisfaction of all local tasks while obeying the above constraints. We present numerical simulations to demonstrate the advantages of the proposed method over most existing approaches that require all-time network connectivity.

IROS Conference 2014 Conference Paper

Three-dimensional multirobot formation control for target enclosing

  • Miguel Aranda
  • Gonzalo López-Nicolás
  • Carlos Sagüés
  • Michael M. Zavlanos

This paper presents a novel method that enables a team of aerial robots to enclose a target in 3D space by attaining a desired geometric formation around it. We propose an approach in which each robot obtains its motion commands using measurements of the relative position of the other agents and of the target, without the need for a central coordinator. As contribution, our method permits any desired 3D target enclosing configuration to be defined, in contrast with the planar circular patterns commonly encountered in the literature. The proposed control strategy relies on the minimization of a cost function that captures the collective motion objective. In our method, the robots do not need to use a common reference frame. This coordinate independence is achieved through the introduction in the cost function of a rotation matrix computed locally by each robot. We prove that our motion controller is exponentially stable, and illustrate its performance through simulations.

ICRA Conference 2013 Conference Paper

A hybrid control approach to the Next-Best-View problem using stereo vision

  • Charles Freundlich
  • Philippos Mordohai
  • Michael M. Zavlanos

In this paper, we consider the problem of precisely localizing a group of stationary targets using a single stereo camera mounted on a mobile robot. In particular, assuming that at least one pair of stereo images of the targets is available, we seek to determine where to move the stereo camera so that the localization uncertainty of the targets is minimized. We call this problem the Next-Best-View problem. The advantage of using a stereo camera is that, using triangulation, the two simultaneous images can yield range and bearing measurements of the targets, as well as their uncertainty. We use a Kalman filter to fuse location and uncertainty estimates as more measurements are acquired. Our solution to the Next-Best-View problem is to iteratively minimize the fused uncertainty of the targets' locations subject to field-of-view constraints. We capture these objectives by appropriate artificial potentials on the camera's relative frame and the global frame, respectively. In particular, with every new observation, the mobile stereo camera computes the new next best view on the relative frame and subsequently realizes this view in the global frame via gradient descent on the space of robot positions and orientations, until a new observation is made. Integration of next best view with motion planning results in a hybrid system, which we illustrate in computer simulations.

ICRA Conference 2008 Conference Paper

Distributed multi-robot task assignment and formation control

  • Nathan Michael
  • Michael M. Zavlanos
  • Vijay Kumar 0001
  • George J. Pappas

Distributed task assignment for multiple agents raises fundamental and novel problems in control theory and robotics. A new challenge is the development of distributed algorithms that dynamically assign tasks to multiple agents, not relying on a priori assignment information. We address this challenge using market-based coordination protocols where the agents are able to bid for task assignment with the assumption that every agent has knowledge of the maximum number of agents that any given task can accommodate. We show that our approach always achieves the desired assignment of agents to tasks after exploring at most a polynomial number of assignments, dramatically reducing the combinatorial nature of discrete assignment problems. We verify our algorithm through both simulation and experimentation on a team of non-holonomic robots performing distributed formation stabilization and group splitting and merging.

ICRA Conference 2007 Conference Paper

Sensor-Based Dynamic Assignment in Distributed Motion Planning

  • Michael M. Zavlanos
  • George J. Pappas

Distributed motion planning of multiple agents raises fundamental and novel problems in control theory and robotics. Recently, one such great challenge has been the development of motion planning algorithms that dynamically assign targets or destinations to multiple homogeneous agents, not relying on any a priori assignment of agents to destinations. In this paper, we address this challenge using two novel ideas. First, we develop distributed multi-destination potential fields able to drive every agent to any available destination for almost all initial conditions. Second, we propose sensor-based coordination protocols that ensure that distinct agents are assigned to distinct destinations. Integration of the overall system results in a distributed, multi-agent, hybrid system for which we show that the mutual exclusion property of the final assignment is guaranteed for almost all initial conditions. Moreover, we show that our dynamic assignment algorithm converges after exploring at most a polynomial number of assignments, dramatically reducing the combinatorial nature of purely discrete assignment problems. Our scalable approach is illustrated with nontrivial computer simulations.

v2026.09.13