Arrow Research search

Author name cluster

Brian M. Sadler

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

15 papers
2 author rows

Possible papers

15

EAAI Journal 2025 Journal Article

Agent deception via polynomial path planning

  • Nolan B. Gutierrez
  • Brian M. Sadler
  • William J. Beksi

Deceptive behavior involves an intelligent agent creating plans that conceal its true intentions while appearing to pursue different goals. It is crucial for inducing confusion in various applications, including security, military, and competitive environments, where the ability to conceal true intentions can lead to significant strategic advantages. Developing better artificial intelligence (AI) for adversarial environments or strategic decision-making scenarios (e. g. , more realistic testing of human or AI decision-making capabilities in games and simulations) requires effective deceptive planning. In this paper, we propose a novel polynomial path planner that enables an agent to deceive its observers. Our contributions include the following: (i) we develop a framework for obtaining ambiguous functions; (ii) we introduce new deception metrics; (iii) we present a method for standardizing trajectories to enable shape-based comparisons independent of speed; (iv) we conduct a human survey to evaluate deception and its relation to goal recognition; and (v) we outperform the state of the art on a multitude of deception metrics. Furthermore, our findings show that using more complex functions and increasing the level of misdirection greatly enhances agent deception effectiveness.

IROS Conference 2025 Conference Paper

Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation

  • Bhrij Patel
  • Kasun Weerakoon
  • Wesley A. Suttle
  • Alec Koppel
  • Brian M. Sadler
  • Tianyi Zhou 0001
  • Dinesh Manocha
  • Amrit Singh Bedi

Reinforcement learning (RL) is a promising approach for robotic navigation, allowing robots to learn through trial and error. However, real-world robotic tasks often suffer from sparse rewards, leading to inefficient exploration and suboptimal policies due to sample inefficiency of RL. In this work, we introduce Confidence-Controlled Exploration (CCE), a novel method that improves sample efficiency in RL-based robotic navigation without modifying the reward function. Unlike existing approaches, such as entropy regularization and reward shaping, which can introduce instability by altering rewards, CCE dynamically adjusts trajectory length based on policy entropy. Specifically, it shortens trajectories when uncertainty is high to enhance exploration and extends them when confidence is high to prioritize exploitation. CCE is a principled and practical solution inspired by a theoretical connection between policy entropy and gradient estimation. It integrates seamlessly with on-policy and off-policy RL methods and requires minimal modifications. We validate CCE across REINFORCE, PPO, and SAC in both simulated and real-world navigation tasks. CCE outperforms fixed-trajectory and entropy-regularized baselines, achieving an 18% higher success rate, 20-38% shorter paths, and 9. 32% lower elevation costs under a fixed training sample budget. Finally, we deploy CCE on a Clearpath Husky robot, demonstrating its effectiveness in complex outdoor environments.

IROS Conference 2025 Conference Paper

On the Vulnerability of LLM/VLM-Controlled Robotics

  • Xiyang Wu
  • Souradip Chakraborty
  • Ruiqi Xian
  • Jing Liang 0006
  • Tianrui Guan
  • Fuxiao Liu
  • Brian M. Sadler
  • Dinesh Manocha

In this work, we highlight vulnerabilities in robotic systems integrating large language models (LLMs) and vision-language models (VLMs) due to input modality sensitivities. While LLM/VLM-controlled robots show impressive performance across various tasks, their reliability under slight input variations remains underexplored yet critical. These models are highly sensitive to instruction or perceptual input changes, which can trigger misalignment issues, leading to execution failures with severe real-world consequences. To study this issue, we analyze the misalignment-induced vulnerabilities within LLM/VLM-controlled robotic systems and present a mathematical formulation for failure modes arising from variations in input modalities. We propose empirical perturbation strategies to expose these vulnerabilities and validate their effectiveness through experiments on multiple robot manipulation tasks. Our results show that simple input perturbations reduce task execution success rates by 22. 2% and 14. 6% in two representative LLM/VLM-controlled robotic systems. These findings underscore the importance of input modality robustness and motivate further research to ensure the safe and reliable deployment of advanced LLM/VLM-controlled robotic systems.

AAMAS Conference 2024 Conference Paper

Deceptive Path Planning via Reinforcement Learning with Graph Neural Networks

  • Michael Y. Fatemi
  • Wesley A. Suttle
  • Brian M. Sadler

Deceptive path planning (DPP) is the problem of designing a path that hides its true goal from an outside observer. Existing methods for DPP rely on unrealistic assumptions, such as global state observability and perfect model knowledge, and therefore do not generalize to unseen problem instances, lack scalability to realistic problem sizes, and preclude both on-the-fly tunability of deception levels and real-time adaptivity to changing environments. In this paper, we propose a reinforcement learning (RL)-based scheme that overcomes these issues. Through extensive experimentation we show that, without additional fine-tuning, at test time the resulting policies successfully generalize, scale, enjoy tunable levels of deception, and adapt in real-time to changes in the environment.

ICML Conference 2024 Conference Paper

PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling

  • Utsav Singh
  • Wesley A. Suttle
  • Brian M. Sadler
  • Vinay P. Namboodiri
  • Amrit Singh Bedi

In this work, we introduce PIPER: Primitive-Informed Preference-based Hierarchical reinforcement learning via Hindsight Relabeling, a novel approach that leverages preference-based learning to learn a reward model, and subsequently uses this reward model to relabel higher-level replay buffers. Since this reward is unaffected by lower primitive behavior, our relabeling-based approach is able to mitigate non-stationarity, which is common in existing hierarchical approaches, and demonstrates impressive performance across a range of challenging sparse-reward tasks. Since obtaining human feedback is typically impractical, we propose to replace the human-in-the-loop approach with our primitive-in-the-loop approach, which generates feedback using sparse rewards provided by the environment. Moreover, in order to prevent infeasible subgoal prediction and avoid degenerate solutions, we propose primitive-informed regularization that conditions higher-level policies to generate feasible subgoals for lower-level policies. We perform extensive experiments to show that PIPER mitigates non-stationarity in hierarchical reinforcement learning and achieves greater than 50$\%$ success rates in challenging, sparse-reward robotic environments, where most other baselines fail to achieve any significant progress.

ICML Conference 2024 Conference Paper

Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles

  • Bhrij Patel
  • Wesley A. Suttle
  • Alec Koppel
  • Vaneet Aggarwal
  • Brian M. Sadler
  • Dinesh Manocha
  • Amrit Singh Bedi

In the context of average-reward reinforcement learning, the requirement for oracle knowledge of the mixing time, a measure of the duration a Markov chain under a fixed policy needs to achieve its stationary distribution, poses a significant challenge for the global convergence of policy gradient methods. This requirement is particularly problematic due to the difficulty and expense of estimating mixing time in environments with large state spaces, leading to the necessity of impractically long trajectories for effective gradient estimation in practical applications. To address this limitation, we consider the Multi-level Actor-Critic (MAC) framework, which incorporates a Multi-level Monte-Carlo (MLMC) gradient estimator. With our approach, we effectively alleviate the dependency on mixing time knowledge, a first for average-reward MDPs global convergence. Furthermore, our approach exhibits the tightest available dependence of $\mathcal{O}(\sqrt{\tau_{mix}})$ known from prior work. With a 2D grid world goal-reaching navigation experiment, we demonstrate that MAC outperforms the existing state-of-the-art policy gradient-based method for average reward settings.

ICML Conference 2023 Conference Paper

Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-Critic

  • Wesley A. Suttle
  • Amrit Singh Bedi
  • Bhrij Patel
  • Brian M. Sadler
  • Alec Koppel
  • Dinesh Manocha

Many existing reinforcement learning (RL) methods employ stochastic gradient iteration on the back end, whose stability hinges upon a hypothesis that the data-generating process mixes exponentially fast with a rate parameter that appears in the step-size selection. Unfortunately, this assumption is violated for large state spaces or settings with sparse rewards, and the mixing time is unknown, making the step size inoperable. In this work, we propose an RL methodology attuned to the mixing time by employing a multi-level Monte Carlo estimator for the critic, the actor, and the average reward embedded within an actor-critic (AC) algorithm. This method, which we call M ulti-level A ctor- C ritic (MAC), is developed specifically for infinite-horizon average-reward settings and neither relies on oracle knowledge of the mixing time in its parameter selection nor assumes its exponential decay; it is therefore readily applicable to applications with slower mixing times. Nonetheless, it achieves a convergence rate comparable to SOTA actor-critic algorithms. We experimentally show that these alleviated restrictions on the technical conditions required for stability translate to superior performance in practice for RL problems with sparse rewards.

ICML Conference 2022 Conference Paper

On the Hidden Biases of Policy Mirror Ascent in Continuous Action Spaces

  • Amrit Singh Bedi
  • Souradip Chakraborty
  • Anjaly Parayil
  • Brian M. Sadler
  • Pratap Tokekar
  • Alec Koppel

We focus on parameterized policy search for reinforcement learning over continuous action spaces. Typically, one assumes the score function associated with a policy is bounded, which {fails to hold even for Gaussian policies. } To properly address this issue, one must introduce an exploration tolerance parameter to quantify the region in which it is bounded. Doing so incurs a persistent bias that appears in the attenuation rate of the expected policy gradient norm, which is inversely proportional to the radius of the action space. To mitigate this hidden bias, heavy-tailed policy parameterizations may be used, which exhibit a bounded score function, but doing so can cause instability in algorithmic updates. To address these issues, in this work, we study the convergence of policy gradient algorithms under heavy-tailed parameterizations, which we propose to stabilize with a combination of mirror ascent-type updates and gradient tracking. Our main theoretical contribution is the establishment that this scheme converges with constant batch sizes, whereas prior works require these parameters to respectively shrink to null or grow to infinity. Experimentally, this scheme under a heavy-tailed policy parameterization yields improved reward accumulation across a variety of settings as compared with standard benchmarks.

IROS Conference 2020 Conference Paper

Game Theoretic Formation Design for Probabilistic Barrier Coverage

  • Daigo Shishika
  • Douglas G. Macharet
  • Brian M. Sadler
  • Vijay Kumar 0001

We study strategies to deploy defenders/sensors to detect intruders that approach a targeted region. This scenario is formulated as a barrier coverage, which aims to minimize the number of unseen paths. The problem becomes challenging when the number of defenders is insufficient for a full coverage, requiring us to find the most effective location to deploy them. To this end, we use ideas from game theory to account for various paths that the intruders may take. Specifically, we propose an iterative algorithm to refine the set of candidate defender formations, which uses the payoff matrix to directly evaluate the utility of different formations. Given the set of candidate formations, a mixed Nash equilibrium gives a stochastic policy to deploy the defenders. The efficacy of the proposed strategy is demonstrated by a numerical analysis that compares our method with an existing graph-theoretic method.

JMLR Journal 2019 Journal Article

Decentralized Dictionary Learning Over Time-Varying Digraphs

  • Amir Daneshmand
  • Ying Sun
  • Gesualdo Scutari
  • Francisco Facchinei
  • Brian M. Sadler

This paper studies Dictionary Learning problems wherein the learning task is distributed over a multi-agent network, modeled as a time-varying directed graph. This formulation is relevant, for instance, in Big Data scenarios where massive amounts of data are collected/stored in different locations (e.g., sensors, clouds) and aggregating and/or processing all data in a fusion center might be inefficient or unfeasible, due to resource limitations, communication overheads or privacy issues. We develop a unified decentralized algorithmic framework for this class of nonconvex problems, which is proved to converge to stationary solutions at a sublinear rate. The new method hinges on Successive Convex Approximation techniques, coupled with a decentralized tracking mechanism aiming at locally estimating the gradient of the smooth part of the sum-utility. To the best of our knowledge, this is the first provably convergent decentralized algorithm for Dictionary Learning and, more generally, bi-convex problems over (time-varying) (di)graphs. [abs] [ pdf ][ bib ] &copy JMLR 2019. ( edit, beta )

TCS Journal 2017 Journal Article

Network connectivity assessment and improvement through relay node deployment

  • Maggie X. Cheng
  • Yi Ling
  • Brian M. Sadler

In wireless ad hoc networks, maintaining network connectivity is very important as high level network functions all depend on it. However, how to measure network connectivity remains a fundamental challenge. For example, a network can have good overall k-connectivity and yet still have a communication bottleneck. In this paper, we address how to locate bottlenecks and relieve them. A new connectivity measure based on the Cheeger's Constant is used for bottleneck discovery, and a partition algorithm that divides the network at the bottleneck is developed. After the network is partitioned, we consider deploying a relay node to increase the conductance of the network across the partition. The relay node deployment problem is formulated as an integer linear program to maximize the number of connections between the two sides of the cut, and then a convex optimization algorithm is used to find the precise location of the relay node, which is within the convex hull defined by the radio transmission ranges of all the nodes that can connect to the relay node. We show that the relay node significantly relieves the bottleneck and improves network connectivity, which is manifested by several network connectivity metrics. The partition and relay node deployment algorithms are then extended to the case where multiple relay nodes are available. Multiway partition and multiple relay node deployment algorithms are presented. Simulation results show this approach effectively enhances network connectivity with a small number of relay nodes.

ICRA Conference 2014 Conference Paper

Deployment of swarms of micro-aerial vehicles: From theory to practice

  • Aveek Purohit
  • Pei Zhang 0001
  • Brian M. Sadler
  • Stefano Carpin

We study the problem of deploying a high number of low-cost, low-complexity robots inside a known environment with the objective that at least one robotic platform reaches each of N preassigned goal locations. Our study is inspired by SensorFly, a micro-aerial vehicle successfully used for mobile sensor network applications. SensorFly nodes feature limited on-board sensors, so one has to rely on simple navigation strategies and increase performance through redundance in the team. We introduce a simple, fully scalable deployment algorithm exploiting the limited capabilities offered by the SensorFly platform, and we explore its performance by feeding the simulation system with parameters extracted from the real SensorFly platform.

IROS Conference 2014 Conference Paper

Rapid multirobot deployment with time constraints

  • Stefano Carpin
  • Marco Pavone 0001
  • Brian M. Sadler

In this paper we consider the problem of multirobot deployment under temporal deadlines. The objective is to compute strategies trading off safety for speed to maximize the probability of reaching a given set of target locations within a pre-assigned temporal deadline. We formulate this problem using the theory of Constrained Markov Decision Processes and we show that thanks to this framework it is possible to determine deploying strategies maximizing the probability of success while satisfying the deadline. Moreover, the formulation allows to exactly compute the failure probability of complex deployment tasks. Simulation results illustrate how the proposed method works in different scenarios and show how informed decisions can be made regarding the size of the robot team.

ICRA Conference 2013 Conference Paper

Theoretical foundations of high-speed robot team deployment

  • Stefano Carpin
  • Timothy H. Chung
  • Brian M. Sadler

In this paper we study the multi-robot deployment problem under hard temporal constraints. After proposing a model for this task, we consider the simplest deployment algorithm and we analyze the relationship between three fundamental parameters, the temporal deadline, the probability of success, and the number of robots. Because an exact analysis of even the simplest algorithm is computationally intractable, we derive an approximate bound leading to performance curves useful to answer design questions (how many robots are needed to get a certain performance guarantee?) or analysis questions (what is the probability of success given a certain deadline and number of robots?) Simulations show that the bounds are sharp and provide a useful tool to predict team deployment performance and tradeoffs.

ICRA Conference 2012 Conference Paper

RSS gradient-assisted frontier exploration and radio source localization

  • Jeffrey N. Twigg
  • Jonathan Fink
  • Paul L. Yu
  • Brian M. Sadler

We consider the combined problem of frontier exploration in a complex indoor environment while seeking a radio source. To do this in an efficient manner, we incorporate radio signal strength (RSS) information into the exploration algorithm by locally sampling the RSS and estimating the 2-D RSS gradient. The algorithm exploits the local motion to collect RSS samples for gradient estimation and seeks to explore in a way that brings the robot to the signal source. This strategy avoids random or exhaustive exploration. An indoor experiment demonstrates the exploration algorithm that uses this information to dynamically prioritize candidate frontiers and traverse to a radio source. Simulations, including radio propagation modeling with a ray-tracing algorithm, enable study of control algorithm tradeoffs and statistical performance.

v2026.09.13