Arrow Research search

Author name cluster

Florian Shkurti

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

34 papers
2 author rows

Possible papers

34

AAAI Conference 2026 Conference Paper

SICNav: Safe and Interactive Crowd Navigation Using Model Predictive Control and Bilevel Optimization (Abstract Reprint)

  • Sepehr Samavi
  • James R. Han
  • Florian Shkurti
  • Angela P. Schoellig

Robots need to predict and react to human motions to navigate through a crowd without collisions. Many existing methods decouple prediction from planning, which does not account for the interaction between robot and human motions and can lead to the robot getting stuck. We propose SICNav, a Model Predictive Control (MPC) method that jointly solves for robot motion and predicted crowd motion in closed-loop. We model each human in the crowd to be following an Optimal Reciprocal Collision Avoidance (ORCA) scheme and embed that model as a constraint in the robot's local planner, resulting in a bilevel nonlinear MPC optimization problem. We use a KKT-reformulation to cast the bilevel problem as a single level and use a nonlinear solver to optimize. Our MPC method can influence pedestrian motion while explicitly satisfying safety constraints in a single-robot multi-human environment. We analyze the performance of SICNav in two simulation environments and indoor experiments with a real robot to demonstrate safe robot motion that can influence the surrounding humans. We also validate the trajectory forecasting performance of ORCA on a human trajectory dataset.

ICRA Conference 2025 Conference Paper

Automated Planning Domain Inference for Task and Motion Planning

  • Jinbang Huang
  • Allen Tao
  • Rozilyn Marco
  • Miroslav Bogdanovic
  • Jonathan Kelly
  • Florian Shkurti

Task and motion planning (TAMP) frameworks address long and complex planning problems by integrating high-level task planners with low-level motion planners. However, existing TAMP methods rely heavily on the manual design of planning domains that specify the preconditions and postconditions of all high-level actions. This paper proposes a method to automate planning domain inference from a handful of test-time trajectory demonstrations, reducing the reliance on human design. Our approach incorporates a deep learning-based estimator that predicts the appropriate components of a domain for a new task and a search algorithm that refines this prediction, reducing the size and ensuring the utility of the inferred domain. Our method can generate new domains from minimal test time demonstrations, enabling robots to handle complex tasks more efficiently. We demonstrate that our approach outperforms behaviour cloning baselines, which directly imitate planner behaviour, in terms of planning performance and generalization across a variety of tasks. Additionally, our method reduces computational costs and data amount requirements at test time for inferring new planning domains.

ICRA Conference 2025 Conference Paper

Gaussian Splatting Visual MPC for Granular Media Manipulation

  • Wei-Cheng Tseng
  • Ellina Zhang
  • Krishna Murthy Jatavallabhula
  • Florian Shkurti

Recent advancements in learned 3D representations have enabled significant progress in solving complex robotic manipulation tasks, particularly for rigid-body objects. However, manipulating granular materials such as beans, nuts, and rice remains challenging due to the intricate physics of particle interactions, high-dimensional and partially observable state, inability to visually track individual particles in a pile, and the computational demands of accurate dynamics prediction. Current deep latent dynamics models often struggle to generalize in granular material manipulation due to a lack of inductive biases. In this work, we propose a novel approach that learns a visual dynamics model over Gaussian splatting repre-sentations of scenes and leverages this model for manipulating granular media via Model-Predictive Control. Our method enables efficient optimization for complex manipulation tasks on piles of granular media. We evaluate our approach in both simulated and real-world settings, demonstrating its ability to solve unseen planning tasks and generalize to new environments in a zero-shot transfer. We also show significant prediction and manipulation performance improvements compared to existing granular media manipulation methods.

ICLR Conference 2025 Conference Paper

Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis

  • Qi Chen
  • Jierui Zhu
  • Florian Shkurti

Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator structure. Leveraging recent information-theoretic tools, we propose a unified theoretical framework that provides guarantees for the generalization of both the encoder and generator by treating them as randomized mappings. This framework further enables (1) a refined analysis for VAEs, accounting for the generator's generalization, which was previously overlooked; (2) illustrating an explicit trade-off in generalization terms for DMs that depends on the diffusion time $T$; and (3) providing computable bounds for DMs based solely on the training data, allowing the selection of the optimal $T$ and the integration of such bounds into the optimization process to improve model performance. Empirical results on both synthetic and real datasets illustrate the validity of the proposed theory.

NeurIPS Conference 2025 Conference Paper

SAFE: Multitask Failure Detection for Vision-Language-Action Models

  • Qiao Gu
  • Yuanliang Ju
  • Shengxiang Sun
  • Igor Gilitschenski
  • Haruki Nishimura
  • Masha Itkina
  • Florian Shkurti

While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on novel tasks out of the box. To allow these policies to safely interact with their environments, we need a failure detector that gives a timely alert such that the robot can stop, backtrack, or ask for help. However, existing failure detectors are trained and tested only on one or a few specific tasks, while generalist VLAs require the detector to generalize and detect failures also in unseen tasks and novel environments. In this paper, we introduce the multitask failure detection problem and propose SAFE, a failure detector for generalist robot policies such as VLAs. We analyze the VLA feature space and find that VLAs have sufficient high-level knowledge about task success and failure, which is generic across different tasks. Based on this insight, we design SAFE to learn from VLA internal features and predict a single scalar indicating the likelihood of task failure. SAFE is trained on both successful and failed rollouts and is evaluated on unseen tasks. SAFE is compatible with different policy architectures. We test it on OpenVLA, $\pi_0$, and $\pi_0$-FAST in both simulated and real-world environments extensively. We compare SAFE with diverse baselines and show that SAFE achieves state-of-the-art failure detection performance and the best trade-off between accuracy and detection time using conformal prediction. More qualitative results and code can be found at the project webpage: https: //vla-safe. github. io/

NeurIPS Conference 2025 Conference Paper

STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation

  • Hossein Goli
  • Michael Gimelfarb
  • Nathan de Lara
  • Haruki Nishimura
  • Masha Itkina
  • Florian Shkurti

Off-policy evaluation (OPE) estimates the performance of a target policy using offline data collected from a behavior policy, and is crucial in domains such as robotics or healthcare where direct interaction with the environment is costly or unsafe. Existing OPE methods are ineffective for high-dimensional, long-horizon problems, due to exponential blow-ups in variance from importance weighting or compounding errors from learned dynamics models. To address these challenges, we propose STITCH-OPE, a model-based generative framework that leverages denoising diffusion for long-horizon OPE in high-dimensional state and action spaces. Starting with a diffusion model pre-trained on the behavior data, STITCH-OPE generates synthetic trajectories from the target policy by guiding the denoising process using the score function of the target policy. STITCH-OPE proposes two technical innovations that make it advantageous for OPE: (1) prevents over-regularization by subtracting the score of the behavior policy during guidance, and (2) generates long-horizon trajectories by stitching partial trajectories together end-to-end. We provide a theoretical guarantee that under mild assumptions, these modifications result in an exponential reduction in variance versus long-horizon trajectory diffusion. Experiments on the D4RL and OpenAI Gym benchmarks show substantial improvement in mean squared error, correlation, and regret metrics compared to state-of-the-art OPE methods.

IROS Conference 2025 Conference Paper

Synthetica: Large Scale Synthetic Data Generation for Robot Perception

  • Ritvik Singh
  • Jason Jingzhou Liu
  • Karl Van Wyk
  • Yu-Wei Chao
  • Jean-Francois Lafleche
  • Florian Shkurti
  • Nathan D. Ratliff
  • Ankur Handa

Vision-based object detectors are a crucial basis for robotics applications as they provide valuable information about object localization in the environment. These need to ensure high reliability in different lighting conditions, occlusions, and visual artifacts, all while running in real-time. Collecting and annotating real-world data for these networks is prohibitively time consuming and costly, especially for custom assets, such as industrial objects, making it untenable for generalization to in-the-wild scenarios. To this end, we present Synthetica, a method for large-scale synthetic data generation for training robust state estimators. This paper focuses on the task of object detection, an important problem which can serve as the front-end for most state estimation problems, such as pose estimation. Leveraging data from a photorealistic ray-tracing renderer, we scale up data generation, generating 2. 7 million images, to train highly accurate real-time detection transformers. We present a collection of rendering randomization and training-time data augmentation techniques conducive to robust sim-to-real performance for vision tasks. We demonstrate state-of-the-art performance on the task of object detection while having detectors that run at 50–100Hz which is 9 times faster than the prior state-of-the-art (SOTA). We further demonstrate the usefulness of our training methodology for robotics applications by showcasing a pipeline for use in the real world with custom objects for which there do not exist prior datasets. Our work highlights the importance of scaling synthetic data generation for robust sim-to-real transfer while achieving the fastest real-time inference speeds. Videos and supplementary information can be found at https://sites.google.com/view/synthetica-vision

ICRA Conference 2024 Conference Paper

ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

  • Qiao Gu
  • Ali Kuwajerwala
  • Sacha Morin
  • Krishna Murthy Jatavallabhula
  • Bipasha Sen
  • Aditya Agarwal
  • Corban Rivera
  • William Paul

For robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D representations. However, these approaches tend to produce maps with per-point feature vectors, which do not scale well in larger environments, nor do they contain semantic spatial relationships between entities in the environment, which are useful for downstream planning. In this work, we propose ConceptGraphs, an open-vocabulary graph-structured representation for 3D scenes. ConceptGraphs is built by leveraging 2D foundation models and fusing their output to 3D by multi-view association. The resulting representations generalize to novel semantic classes, without the need to collect large 3D datasets or finetune models. We demonstrate the utility of this representation through a number of downstream planning tasks that are specified through abstract (language) prompts and require complex reasoning over spatial and semantic concepts. To explore the full scope of our experiments and results, we encourage readers to visit our project webpage.

IROS Conference 2023 Conference Paper

Does Unpredictability Influence Driving Behavior?

  • Sepehr Samavi
  • Florian Shkurti
  • Angela P. Schoellig

In this paper we investigate the effect of the unpredictability of surrounding cars on an ego-car performing a driving maneuver. We use Maximum Entropy Inverse reinforcement Learning to model reward functions for an ego-car conducting a lane change in a highway setting. We define a new feature based on the unpredictability of surrounding cars and use it in the reward function. We learn two reward functions from human data: a baseline and one that incorporates our defined unpredictability feature, then compare their performance with a quantitative and qualitative evaluation. Our evaluation demonstrates that incorporating the unpredictability feature leads to a better fit of human-generated test data. These results encourage further investigation of the effect of unpredictability on driving behavior.

ICRA Conference 2023 Conference Paper

MVTrans: Multi-View Perception of Transparent Objects

  • Yi Ru Wang
  • Yuchi Zhao
  • Haoping Xu
  • Sagi Eppel
  • Alán Aspuru-Guzik
  • Florian Shkurti
  • Animesh Garg

Transparent object perception is a crucial skill for applications such as robot manipulation in household and laboratory settings. Existing methods utilize RGB-D or stereo inputs to handle a subset of perception tasks including depth and pose estimation. However transparent object perception remains to be an open problem. In this paper, we forgo the unreliable depth map from RGB-D sensors and extend the stereo based method. Our proposed method, MVTrans, is an end-to-end multi-view architecture with multiple perception capabilities, including depth estimation, segmentation, and pose estimation. Additionally, we establish a novel procedural photo-realistic dataset generation pipeline and create a large-scale transparent object detection dataset, Syn-TODD, which is suitable for training networks with all three modalities, RGB-D, stereo and multi-view RGB. https://ac-rad.github.io/MVTrans/

ICRA Conference 2023 Conference Paper

Policy-Guided Lazy Search with Feedback for Task and Motion Planning

  • Mohamed Khodeir
  • Atharv Sonwane
  • Ruthrash Hari
  • Florian Shkurti

PDDLStream solvers have recently emerged as viable solutions for Task and Motion Planning (TAMP) problems, extending PDDL to problems with continuous action spaces. Prior work has shown how PDDLStream problems can be reduced to a sequence of PDDL planning problems, which can then be solved using off-the-shelf planners. However, this approach can suffer from long runtimes. In this paper we propose LAZY, a solver for PDDLStream problems that maintains a single integrated search over action skeletons, which gets progressively more geometrically informed, as samples of possible motions are lazily drawn during motion planning. We explore how learned models of goal-directed policies and current motion sampling data can be incorporated in LAZY to adaptively guide the task planner. We show that this leads to significant speed-ups in the search for a feasible solution evaluated over unseen test environments of varying numbers of objects, goals, and initial conditions. We evaluate our TAMP approach by comparing to existing solvers for PDDLStream problems on a range of simulated 7DoF rearrangement/manipulation problems. Code can be found at https://rvl.cs.toronto.edu/learning-based-tamp.

ICRA Conference 2023 Conference Paper

Stochastic Planning for ASV Navigation Using Satellite Images

  • Yizhou Huang
  • Hamza Dugmag
  • Tim D. Barfoot
  • Florian Shkurti

Autonomous surface vessels (ASV) represent a promising technology to automate water-quality monitoring of lakes. In this work, we use satellite images as a coarse map and plan sampling routes for the robot. However, inconsistency between the satellite images and the actual lake, as well as environmental disturbances such as wind, aquatic vegetation, and changing water levels can make it difficult for robots to visit places suggested by the prior map. This paper presents a robust route-planning algorithm that minimizes the expected total travel distance given these environmental disturbances, which induce uncertainties in the map. We verify the efficacy of our algorithm in simulations of over a thousand Canadian lakes and demonstrate an application of our algorithm in a 3. 7 km-long real-world robot experiment on a lake in Northern Ontario, Canada.

ICRA Conference 2022 Conference Paper

Augmenting Imitation Experience via Equivariant Representations

  • Dhruv Sharma
  • Ali Kuwajerwala
  • Florian Shkurti

The robustness of visual navigation policies trained through imitation often hinges on the augmentation of the training image-action pairs. Traditionally, this has been done by collecting data from multiple cameras, by using standard data augmentations from computer vision, such as adding random noise to each image, or by synthesizing training images. In this paper we show that there is another practical alternative for data augmentation for visual navigation based on extrapolating viewpoint embeddings and actions nearby the ones observed in the training data. Our method makes use of the geometry of the visual navigation problem in 2D and 3D and relies on policies that are functions of equivariant embeddings, as opposed to images. Given an image-action pair from a training navigation dataset, our neural network model predicts the latent representations of images at nearby viewpoints, using the equivariance property, and augments the dataset. We then train a policy on the augmented dataset. Our simulation results indicate that policies trained in this way exhibit reduced cross-track error, and require fewer interventions compared to policies trained using standard augmentation methods. We also show similar results in autonomous visual navigation by a real ground robot along a path of over 500m.

ICLR Conference 2021 Conference Paper

Conservative Safety Critics for Exploration

  • Homanga Bharadhwaj
  • Aviral Kumar
  • Nicholas Rhinehart
  • Sergey Levine
  • Florian Shkurti
  • Animesh Garg

Safe exploration presents a major challenge in reinforcement learning (RL): when active data collection requires deploying partially trained policies, we must ensure that these policies avoid catastrophically unsafe regions, while still enabling trial and error learning. In this paper, we target the problem of safe exploration in RL, by learning a conservative safety estimate of environment states through a critic, and provably upper bound the likelihood of catastrophic failures at every training iteration. We theoretically characterize the tradeoff between safety and policy improvement, show that the safety constraints are satisfied with high probability during training, derive provable convergence guarantees for our approach which is no worse asymptotically then standard RL, and empirically demonstrate the efficacy of the proposed approach on a suite of challenging navigation, manipulation, and locomotion tasks. Our results demonstrate that the proposed approach can achieve competitive task performance, while incurring significantly lower catastrophic failure rates during training as compared to prior methods. Videos are at this URL https://sites.google.com/view/conservative-safety-critics/

ICRA Conference 2021 Conference Paper

Continual Model-Based Reinforcement Learning with Hypernetworks

  • Yizhou Huang
  • Kevin Xie
  • Homanga Bharadhwaj
  • Florian Shkurti

Effective planning in model-based reinforcement learning (MBRL) and model-predictive control (MPC) relies on the accuracy of the learned dynamics model. In many instances of MBRL and MPC, this model is assumed to be stationary and is periodically re-trained from scratch on state transition experience collected from the beginning of environment interactions. This implies that the time required to train the dynamics model - and the pause required between plan executions - grows linearly with the size of the collected experience. We argue that this is too slow for lifelong robot learning and propose HyperCRL, a method that continually learns the encountered dynamics in a sequence of tasks using task-conditional hypernetworks. Our method has three main attributes: first, it includes dynamics learning sessions that do not revisit training data from previous tasks, so it only needs to store the most recent fixed-size portion of the state transition experience; second, it uses fixed-capacity hypernetworks to represent non-stationary and task-aware dynamics; third, it outperforms existing continual learning alternatives that rely on fixed-capacity networks, and does competitively with baselines that remember an ever increasing coreset of past experience. We show that HyperCRL is effective in continual model-based reinforcement learning in robot locomotion and manipulation scenarios, such as tasks involving pushing and door opening. Our project website with videos is at this link http://rvl.cs.toronto.edu/blog/2020/hypercrl/

AAAI Conference 2021 Conference Paper

DIBS: Diversity Inducing Information Bottleneck in Model Ensembles

  • Samarth Sinha
  • Homanga Bharadhwaj
  • Anirudh Goyal
  • Hugo Larochelle
  • Animesh Garg
  • Florian Shkurti

Although deep learning models have achieved state-of-the art performance on a number of vision tasks, generalization over high dimensional multi-modal data, and reliable predictive uncertainty estimation are still active areas of research. Bayesian approaches including Bayesian Neural Nets (BNNs) do not scale well to modern computer vision tasks, as they are difficult to train, and have poor generalization under dataset-shift. This motivates the need for effective ensembles which can generalize and give reliable uncertainty estimates. In this paper, we target the problem of generating effective ensembles of neural networks by encouraging diversity in prediction. We explicitly optimize a diversity inducing adversarial loss for learning the stochastic latent variables and thereby obtain diversity in the output predictions necessary for modeling multi-modal data. We evaluate our method on benchmark datasets: MNIST, CIFAR100, TinyImageNet and MIT Places 2, and compared to the most competitive baselines show significant improvements in classification accuracy, under a shift in the data distribution and in out-of-distribution detection. : over 10% relative improvement in classification accuracy, over 5% relative improvement in generalizing under dataset shift, and over 5% better predictive uncertainty estimation as inferred by efficient out-of-distribution (OOD) detection.

ICLR Conference 2021 Conference Paper

gradSim: Differentiable simulation for system identification and visuomotor control

  • Krishna Murthy Jatavallabhula
  • Miles Macklin
  • Florian Golemo
  • Vikram Voleti
  • Linda Petrini
  • Martin Weiss
  • Breandan Considine
  • Jérôme Parent-Lévesque

In this paper, we tackle the problem of estimating object physical properties such as mass, friction, and elasticity directly from video sequences. Such a system identification problem is fundamentally ill-posed due to the loss of information during image formation. Current best solutions to the problem require precise 3D labels which are labor intensive to gather, and infeasible to create for many systems such as deformable solids or cloth. In this work we present gradSim, a framework that overcomes the dependence on 3D supervision by combining differentiable multiphysics simulation and differentiable rendering to jointly model the evolution of scene dynamics and image formation. This unique combination enables backpropagation from pixels in a video sequence through to the underlying physical attributes that generated them. Furthermore, our unified computation graph across dynamics and rendering engines enables the learning of challenging visuomotor control tasks, without relying on state-based (3D) supervision, while obtaining performance competitive to/better than techniques that require precise 3D labels.

IROS Conference 2021 Conference Paper

Latent Attention Augmentation for Robust Autonomous Driving Policies

  • Ran Cheng
  • Christopher Agia
  • Florian Shkurti
  • David Meger
  • Gregory Dudek

Model-free reinforcement learning has become a viable approach for vision-based robot control. However, sample complexity and adaptability to domain shifts remain persistent challenges when operating in high-dimensional observation spaces (images, LiDAR), such as those that are involved in autonomous driving. In this paper, we propose a flexible framework by which a policy’s observations are augmented with robust attention representations in the latent space to guide the agent’s attention during training. Our method encodes local and global descriptors of the augmented state representations into a compact latent vector, and scene dynamics are approximated by a recurrent network that processes the latent vectors in sequence. We outline two approaches for constructing attention maps; a supervised pipeline leveraging semantic segmentation networks, and an unsupervised pipeline relying only on classical image processing techniques. We conduct our experiments in simulation and test the learned policy against varying seasonal effects and weather conditions. Our design decisions are supported in a series of ablation studies. The results demonstrate that our state augmentation method both improves learning efficiency and encourages robust domain adaptation when compared to common end-to-end frameworks and methods that learn directly from intermediate representations.

ICLR Conference 2021 Conference Paper

Latent Skill Planning for Exploration and Transfer

  • Kevin Xie
  • Homanga Bharadhwaj
  • Danijar Hafner
  • Animesh Garg
  • Florian Shkurti

To quickly solve new tasks in complex environments, intelligent agents need to build up reusable knowledge. For example, a learned world model captures knowledge about the environment that applies to new tasks. Similarly, skills capture general behaviors that can apply to new tasks. In this paper, we investigate how these two approaches can be integrated into a single reinforcement learning agent. Specifically, we leverage the idea of partial amortization for fast adaptation at test time. For this, actions are produced by a policy that is learned over time while the skills it conditions on are chosen using online planning. We demonstrate the benefits of our design decisions across a suite of challenging locomotion tasks and demonstrate improved sample efficiency in single tasks as well as in transfer from one task to another, as compared to competitive baselines. Videos are available at: https://sites.google.com/view/latent-skill-planning/

ICRA Conference 2021 Conference Paper

LEAF: Latent Exploration Along the Frontier

  • Homanga Bharadhwaj
  • Animesh Garg
  • Florian Shkurti

Self-supervised goal proposal and reaching is a key component for exploration and efficient policy learning algorithms. Such a self-supervised approach without access to any oracle goal sampling distribution requires deep exploration and commitment so that long horizon plans can be efficiently discovered. In this paper, we propose an exploration framework, which learns a dynamics-aware manifold of reachable states. For a goal, our proposed method deterministically visits a state at the current frontier of reachable states (commit- ment/reaching) and then stochastically explores to reach the goal (exploration). This allocates exploration budget near the frontier of the reachable region instead of its interior. We target the challenging problem of policy learning from initial and goal states specified as images, and do not assume any access to the underlying ground-truth states of the robot and the environment. To keep track of reachable latent states, we propose a distance-conditioned reachability network that is trained to infer whether one state is reachable from another within the specified latent space distance. Given an initial state, we obtain a frontier of reachable states from that state. By incorporating a curriculum for sampling easier goals (closer to the start state) before more difficult goals, we demonstrate that the proposed self-supervised exploration algorithm, superior performance compared to existing baselines on a set of challenging robotic environments. https://sites.google.com/view/leaf-exploration

ICRA Conference 2021 Conference Paper

Shaping Rewards for Reinforcement Learning with Imperfect Demonstrations using Generative Models

  • Yuchen Wu 0001
  • Melissa Mozifian
  • Florian Shkurti

The potential benefits of model-free reinforcement learning to real robotics systems are limited by its uninformed exploration that leads to slow convergence, lack of data-efficiency, and unnecessary interactions with the environment. To address these drawbacks we propose a method that combines reinforcement and imitation learning by shaping the reward function with a state-and-action-dependent potential that is trained from demonstration data, using a generative model. We show that this accelerates policy learning by specifying high-value areas of the state and action space that are worth exploring first. Unlike the majority of existing methods that assume optimal demonstrations and incorporate the demonstration data as hard constraints on policy optimization, we instead incorporate demonstration data as advice in the form of a reward shaping potential trained as a generative model of states and actions. In particular, we examine both normalizing flows and Generative Adversarial Networks to represent these potentials. We show that, unlike many existing approaches that incorporate demonstrations as hard constraints, our approach is unbiased even in the case of suboptimal and noisy demonstrations. We present an extensive range of simulations, as well as experiments on the Franka Emika 7DOF arm, to demonstrate the practicality of our method.

IROS Conference 2020 Conference Paper

Catch the Ball: Accurate High-Speed Motions for Mobile Manipulators via Inverse Dynamics Learning

  • Ke Dong
  • Karime Pereida
  • Florian Shkurti
  • Angela P. Schoellig

Mobile manipulators consist of a mobile platform equipped with one or more robot arms and are of interest for a wide array of challenging tasks because of their extended workspace and dexterity. Typically, mobile manipulators are deployed in slow-motion collaborative robot scenarios. In this paper, we consider scenarios where accurate high-speed motions are required. We introduce a framework for this regime of tasks including two main components: (i) a bi-level motion optimization algorithm for real-time trajectory generation, which relies on Sequential Quadratic Programming (SQP) and Quadratic Programming (QP), respectively; and (ii) a learning-based controller optimized for precise tracking of high-speed motions via a learned inverse dynamics model. We evaluate our framework with a mobile manipulator platform through numerous high-speed ball catching experiments, where we show a success rate of 85. 33%. To the best of our knowledge, this success rate exceeds the reported performance of existing related systems [1], [2] and sets a new state of the art.

IROS Conference 2020 Conference Paper

One-Shot Informed Robotic Visual Search in the Wild

  • Karim Koreitem
  • Florian Shkurti
  • Travis Manderson
  • Wei-Di Chang
  • Juan Camilo Gamboa Higuera
  • Gregory Dudek

We consider the task of underwater robot navigation for the purpose of collecting scientifically relevant video data for environmental monitoring. The majority of field robots that currently perform monitoring tasks in unstructured natural environments navigate via path-tracking a pre-specified sequence of waypoints. Although this navigation method is often necessary, it is limiting because the robot does not have a model of what the scientist deems to be relevant visual observations. Thus, the robot can neither visually search for particular types of objects, nor focus its attention on parts of the scene that might be more relevant than the pre-specified waypoints and viewpoints. In this paper we propose a method that enables informed visual navigation via a learned visual similarity operator that guides the robot's visual search towards parts of the scene that look like an exemplar image, which is given by the user as a high-level specification for data collection. We propose and evaluate a weakly supervised video representation learning method that outperforms ImageNet embeddings for similarity tasks in the underwater domain. We also demonstrate the deployment of this similarity operator during informed visual navigation in collaborative environmental monitoring scenarios, in large-scale field trials, where the robot and a human scientist collaboratively search for relevant visual content. Code: https://github.com/rvl-lab-utoronto/visual_search_in_the_wild.

ICRA Conference 2019 Conference Paper

Generating Adversarial Driving Scenarios in High-Fidelity Simulators

  • Yasasa Abeysirigoonawardena
  • Florian Shkurti
  • Gregory Dudek

In recent years self-driving vehicles have become more commonplace on public roads, with the promise of bringing safety and efficiency to modern transportation systems. Increasing the reliability of these vehicles on the road requires an extensive suite of software tests, ideally performed on high-fidelity simulators, where multiple vehicles and pedestrians interact with the self-driving vehicle. It is therefore of critical importance to ensure that self-driving software is assessed against a wide range of challenging simulated driving scenarios. The state of the art in driving scenario generation, as adopted by some of the front-runners of the self-driving car industry, still relies on human input [1]. In this paper we propose to automate the process using Bayesian optimization to generate adversarial self-driving scenarios that expose poorly-engineered or poorly-trained self-driving policies, and increase the risk of collision with simulated pedestrians and vehicles. We show that by incorporating the generated scenarios into the training set of the self-driving policy, and by fine-tuning the policy using vision-based imitation learning we obtain safer self-driving behavior.

ICRA Conference 2018 Conference Paper

Model-Based Probabilistic Pursuit via Inverse Reinforcement Learning

  • Florian Shkurti
  • Nikhil Kakodkar
  • Gregory Dudek

We address the integrated prediction, planning, and control problem that enables a single follower robot (the photographer) to quickly re-establish visual contact with a moving target (the subject) that has escaped the follower's field of view. We deal with this scenario, which reactive controllers are typically ill-equipped to handle, by making plausible predictions about the long- and short-term behavior of the target, and planning pursuit paths that will maximize the chance of seeing the target again. At the core of our pursuit method is the use of predictive models of target behavior, which help narrow down the set of possible future locations of the target to a few discrete hypotheses, as well as the use of combinatorial search in physical space to check those hypotheses efficiently. We model target behavior in terms of a learned navigation reward function, using Inverse Reinforcement Learning, based on semantic terrain features of satellite maps. Our pursuit algorithm continuously predicts the latent destination of the target and its position in the future, and relies on efficient graph representation and search methods in order to navigate to locations at which the target is most likely to be seen at an anticipated time. We perform extensive evaluation of our predictive pursuit algorithm over multiple satellite maps, thousands of simulation scenarios, against state-of-the art MDP and POMDP solvers. We show that our method significantly outperforms them by exploiting domain-specific knowledge, while being able to run in real-time.

IROS Conference 2017 Conference Paper

Topologically distinct trajectory predictions for probabilistic pursuit

  • Florian Shkurti
  • Gregory Dudek

We address the integrated planning and control problem that enables a single follower robot (the “photographer”) to maintain a moving target (the “subject”) in its field of view for as long as possible. We propose a real-time pursuit algorithm that seamlessly handles the often neglected, yet unavoidable, scenario in which the target escapes the follower's field of view; a scenario that simple, reactive controllers are ill-equipped to handle. Our algorithm aims to minimize the expected time until visual contact is re-established, which enables the photographer to track the subject for as long as possible, even in the presence of loss of visibility. At the core of our pursuit algorithm is an efficient method for sampling plausible trajectories from different homotopy classes. We do this by generating topologically distinct shortest paths by using the Voronoi diagram. We use these paths to make informed, model-based predictions of the likely future locations of the target, given a history of observations. Given these predictions, our algorithm produces pursuit trajectories that approximately minimize the expected time to recover visual contact. We show that constraining the predictive pursuit problem to the space of homotopy classes condenses the expanse of possibilities that our algorithm must consider, which enables target tracking in large occupancy grids, as opposed to many POMDP methods that are constrained to small environments. We benchmark the tracking behavior of our algorithm against the baseline of human subjects who performed the same set of pursuit tasks in simulation, as well as against two other pursuit algorithms that only take into account paths from a single homotopy class. We show that considering homotopy alternatives in 2D pursuit improves the tracking performance and that our algorithm does at least as well as humans in most pursuit scenarios.

IROS Conference 2017 Conference Paper

Underwater multi-robot convoying using visual tracking by detection

  • Florian Shkurti
  • Wei-Di Chang
  • Peter Henderson 0002
  • Md Jahidul Islam
  • Juan Camilo Gamboa Higuera
  • Jimmy Li 0001
  • Travis Manderson
  • Anqi Xu 0003

We present a robust multi-robot convoying approach that relies on visual detection of the leading agent, thus enabling target following in unstructured 3-D environments. Our method is based on the idea of tracking-by-detection, which interleaves efficient model-based object detection with temporal filtering of image-based bounding box estimation. This approach has the important advantage of mitigating tracking drift (i. e. drifting away from the target object), which is a common symptom of model-free trackers and is detrimental to sustained convoying in practice. To illustrate our solution, we collected extensive footage of an underwater robot in ocean settings, and hand-annotated its location in each frame. Based on this dataset, we present an empirical comparison of multiple tracker variants, including the use of several convolutional neural networks, both with and without recurrent connections, as well as frequency-based model-free trackers. We also demonstrate the practicality of this tracking-by-detection strategy in real-world scenarios by successfully controlling a legged underwater robot in five degrees of freedom to follow another robot's independent motion.

IROS Conference 2014 Conference Paper

3D trajectory synthesis and control for a legged swimming robot

  • David Meger
  • Florian Shkurti
  • David Cortés Poza
  • Philippe Giguère
  • Gregory Dudek

Inspection and exploration of complex underwater structures requires the development of agile and easy to program platforms. In this paper, we describe a system that enables the deployment of an autonomous underwater vehicle in 3D environments proximal to the ocean bottom. Unlike many previous approaches, our solution: uses oscillating hydrofoil propulsion; allows for stable control of the robot's motion and sensor directions; allows human operators to specify detailed trajectories in a natural fashion; and has been successfully demonstrated as a holistic system in the open ocean near both coral reefs and a sunken cargo ship. A key component of our system is the 3D control of a hexapod swimming robot, which can move the vehicle through agile sequences of orientations despite challenging marine conditions. We present two methods to easily generate robot trajectories appropriate for deployments in close proximity to challenging contours of the sea floor. Both offline recording of trajectories using augmented reality and online placement of fiducial tags in the marine environment are shown to have desirable properties, with complementary strengths and weaknesses. Finally, qualitative and quantitative results of the 3D control system are presented.

IROS Conference 2014 Conference Paper

Ear-based exploration on hybrid metric/topological maps

  • Qiwen Zhang
  • David Whitney
  • Florian Shkurti
  • Ioannis M. Rekleitis

In this paper we propose a hierarchy of techniques for performing loop closure in indoor environments together with an exploration strategy designed to reduce uncertainty in the resulting map. We use the generalized Voronoi graph to represent the indoor environment and an extended Kalman filter to track the pose of the robot and the position of the junctions (vertices) of the topological graph. Every time a vertex is revisited, the robot re-localizes and updates the uncertainty estimate accordingly. Finally, since the reduction of the map uncertainty remains one of the main concerns, the robot will optimize its schedule of revisiting junctions in the environment in order to reduce the accumulated uncertainty. Experimental results from a mobile robot equipped with a laser range-finder and results from realistic simulations that validate our approach are presented.

ICRA Conference 2014 Conference Paper

Maximizing visibility in collaborative trajectory planning

  • Florian Shkurti
  • Gregory Dudek

In this paper we address the issue of coordinating the trajectories of two collaborating robots in environments with obstacles so that visibility between them is maximized in the presence of competing constraints. Specifically, we examine the problem of allowing one robot (the “photographer”) to follow another robot (“the subject”) through a planar environment while maintaining visual contact to the maximum degree consistent with an efficient traversal. This problem has numerous applications, for instance in scenarios where communication between robots requires line-of-sight. We formalize this problem in the context of centralized kinodynamic planning and we present solutions based on the asymptotically optimal sampling-based RRT* planner. We discuss connections to the traditional formulation of pursuit-evasion games where the analysis typically ends the moment the evader manages to escape the pursuer's visibility region. We also illustrate types of environments and other conditions under which allowing the pair of robots to break the line-of-sight is a better option than always requiring the presence of visual contact.

ICRA Conference 2013 Conference Paper

On the complexity of searching for an evader with a faster pursuer

  • Florian Shkurti
  • Gregory Dudek

In this paper we examine pursuit-evasion games in which the pursuer has higher speed than the evader. This scenario is motivated by visibility-based pursuit-evasion problems, particularly by the question of what happens when the pursuer loses visual track of the moving evader. In these cases the pursuer has two options for recovering visual contact with the evader: to perform search over the possible locations where the evader might be moving, or to clear the environment, in other words to progressively search it without allowing the evader to move into locations that have already been cleared. It has been shown that in sufficiently complex environments a single pursuer having the same speed as the evader cannot clear the environment. In this work we prove that computing the minimum speed which enables a faster pursuer to clear a graph environment is NP-hard. In light of this result we provide an experimental comparison of randomized and deterministic search strategies on planar graphs, which has practical significance in search and rescue settings.

IROS Conference 2012 Conference Paper

Multi-domain monitoring of marine environments using a heterogeneous robot team

  • Florian Shkurti
  • Anqi Xu 0003
  • Malika Meghjani
  • Juan Camilo Gamboa Higuera
  • Yogesh A. Girdhar
  • Philippe Giguère
  • Bir Bikram Dey
  • Jimmy Li 0001

In this paper we describe a heterogeneous multi-robot system for assisting scientists in environmental monitoring tasks, such as the inspection of marine ecosystems. This team of robots is comprised of a fixed-wing aerial vehicle, an autonomous airboat, and an agile legged underwater robot. These robots interact with off-site scientists and operate in a hierarchical structure to autonomously collect visual footage of interesting underwater regions, from multiple scales and mediums. We discuss organizational and scheduling complexities associated with multi-robot experiments in a field robotics setting. We also present results from our field trials, where we demonstrated the use of this heterogeneous robot team to achieve multi-domain monitoring of coral reefs, based on real-time interaction with a remotely-located marine biologist.

IROS Conference 2011 Conference Paper

MARE: Marine Autonomous Robotic Explorer

  • Yogesh A. Girdhar
  • Anqi Xu 0003
  • Bir Bikram Dey
  • Malika Meghjani
  • Florian Shkurti
  • Ioannis M. Rekleitis
  • Gregory Dudek

We present MARE, an autonomous airboat robot that is suitable for exploration-oriented tasks, such as inspection of coral reefs and shallow seabeds. The combination of this platform's particular mechanical properties and its powerful software framework enables it to function in a multitude of potential capacities, including autonomous surveillance, mapping, and search operations. In this paper we describe two different exploration strategies and their implementation using the MARE platform. First, we discuss the application of an efficient coverage algorithm, for the purpose of achieving systematic exploration of a known and bounded environment. Second, we present an exploration strategy driven by surprise, which steers the robot on a path that might lead to potentially surprising observations.

IROS Conference 2011 Conference Paper

State estimation of an underwater robot using visual and inertial information

  • Florian Shkurti
  • Ioannis M. Rekleitis
  • Milena Scaccia
  • Gregory Dudek

This paper presents an adaptation of a vision and inertial-based state estimation algorithm for use in an underwater robot. The proposed approach combines information from an Inertial Measurement Unit (IMU) in the form of linear accelerations and angular velocities, depth data from a pressure sensor, and feature tracking from a monocular downward facing camera to estimate the 6DOF pose of the vehicle. To validate the approach, we present extensive experimental results from field trials conducted in underwater environments with varying lighting and visibility conditions, and we demonstrate successful application of the technique underwater.

v2026.09.13