Arrow Research search

Author name cluster

Stan Birchfield

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

37 papers
2 author rows

Possible papers

37

NeurIPS Conference 2025 Conference Paper

RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion

  • Bardienus Duisterhof
  • Jan Oberst
  • Bowen Wen
  • Stan Birchfield
  • Deva Ramanan
  • Jeffrey Ichnowski

3D shape completion has broad applications in robotics, digital twin reconstruction, and extended reality (XR). Although recent advances in 3D object and scene completion have achieved impressive results, existing methods lack 3D consistency, are computationally expensive, and struggle to capture sharp object boundaries. Our work (RaySt3R) addresses these limitations by recasting 3D shape completion as a novel view synthesis problem. Specifically, given a single RGB-D image, and a novel viewpoint (encoded as a collection of query rays), we train a feedforward transformer to predict depth maps, object masks, and per-pixel confidence scores for those query rays. RaySt3R fuses these predictions across multiple query views to reconstruct complete 3D shapes. We evaluate RaySt3R on synthetic and real-world datasets, and observe it achieves state-of-the-art performance, outperforming the baselines on all datasets by up to 44% in 3D chamfer distance.

ICRA Conference 2025 Conference Paper

SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation

  • Cheng-Chun Hsu
  • Bowen Wen
  • Jie Xu 0028
  • Yashraj S. Narang
  • Xiaolong Wang 0004
  • Yuke Zhu
  • Joydeep Biswas
  • Stan Birchfield

We introduce SPOT, an object-centric imitation learning framework. The key idea is to capture each task by an object-centric representation, specifically the SE(3) object pose trajectory relative to the target. This approach decouples embodiment actions from sensory inputs, facilitating learning from various demonstration types, including both action-based and action-less human hand demonstrations, as well as crossembodiment generalization. Additionally, object pose trajectories inherently capture planning constraints from demonstrations without the need for manually-crafted rules. To guide the robot in executing the task, the object trajectory is used to condition a diffusion policy. We systematically evaluate our method on simulation and real-world tasks. In real-world evaluation, using only eight demonstrations shot on an iPhone, our approach completed all tasks while fully complying with task constraints. Project page: https://nvlabs.github.io/object_centric_diffusion

IROS Conference 2023 Conference Paper

HANDAL: A Dataset of Real-World Manipulable Object Categories with Pose Annotations, Affordances, and Reconstructions

  • Andrew Guo
  • Bowen Wen
  • Jianhe Yuan
  • Jonathan Tremblay
  • Stephen Tyree
  • Jeffrey Smith 0002
  • Stan Birchfield

We present the HANDAL dataset for category-level object pose estimation and affordance prediction. Unlike previous datasets, ours is focused on robotics-ready manipulable objects that are of the proper size and shape for functional grasping by robot manipulators, such as pliers, utensils, and screwdrivers. Our annotation process is streamlined, requiring only a single off-the-shelf camera and semi-automated processing, allowing us to produce high-quality 3D annotations without crowd-sourcing. The dataset consists of 308k annotated image frames from 2. 2k videos of 212 real-world objects in 17 categories. We focus on hardware and kitchen tool objects to facilitate research in practical scenarios in which a robot manipulator needs to interact with the environment beyond simple pushing or indiscriminate grasping. We outline the usefulness of our dataset for 6-DoF category-level pose+scale estimation and related tasks. We also provide 3D reconstructed meshes of all objects, and we outline some of the bottlenecks to be addressed for democratizing the collection of datasets like this one. Project website: https://nvlabs.github.io/HANDAL/

ICRA Conference 2023 Conference Paper

Parallel Inversion of Neural Radiance Fields for Robust Pose Estimation

  • Yunzhi Lin
  • Thomas Müller 0013
  • Jonathan Tremblay
  • Bowen Wen
  • Stephen Tyree
  • Alex Evans
  • Patricio A. Vela
  • Stan Birchfield

We present a parallelized optimization method based on fast Neural Radiance Fields (NeRF) for estimating 6-DoF pose of a camera with respect to an object or scene. Given a single observed RGB image of the target, we can predict the translation and rotation of the camera by minimizing the residual between pixels rendered from a fast NeRF model and pixels in the observed image. We integrate a momentum-based camera extrinsic optimization procedure into Instant Neural Graphics Primitives, a recent exceptionally fast NeRF implementation. By introducing parallel Monte Carlo sampling into the pose estimation task, our method overcomes local minima and improves efficiency in a more extensive search space. We also show the importance of adopting a more robust pixel-based loss function to reduce error. Experiments demonstrate that our method can achieve improved generalization and robustness on both synthetic and real-world benchmarks.

ICRA Conference 2023 Conference Paper

RGB-Only Reconstruction of Tabletop Scenes for Collision-Free Manipulator Control

  • Zhenggang Tang
  • Balakumar Sundaralingam
  • Jonathan Tremblay
  • Bowen Wen
  • Ye Yuan
  • Stephen Tyree
  • Charles T. Loop
  • Alexander G. Schwing

We present a system for collision-free control of a robot manipulator that uses only RGB views of the world. Perceptual input of a tabletop scene is provided by multiple images of an RGB camera (without depth) that is either handheld or mounted on the robot end effector. A NeRF-like process is used to reconstruct the 3D geometry of the scene, from which the Euclidean full signed distance function (ESDF) is computed. A model predictive control algorithm is then used to control the manipulator to reach a desired pose while avoiding obstacles in the ESDF. We show results on a real dataset collected and annotated in our lab. Our results are also available at https://ngp-mpc.github.io/.

IROS Conference 2022 Conference Paper

6-DoF Pose Estimation of Household Objects for Robotic Manipulation: An Accessible Dataset and Benchmark

  • Stephen Tyree
  • Jonathan Tremblay
  • Thang To
  • Jia Cheng
  • Terry Mosier
  • Jeffrey Smith 0002
  • Stan Birchfield

We present a new dataset for 6-DoF pose estimation of known objects, with a focus on robotic manipulation research. We propose a set of toy grocery objects, whose physical instantiations are readily available for purchase and are appropriately sized for robotic grasping and manipulation. We provide 3D scanned textured models of these objects, suitable for generating synthetic training data, as well as RGBD images of the objects in challenging, cluttered scenes exhibiting partial occlusion, extreme lighting variations, multiple instances per image, and a large variety of poses. Using semi-automated RGBD-to-model texture correspondences, the images are annotated with ground truth poses accurate within a few millimeters. We also propose a new pose evaluation metric called ADD-H based on the Hungarian assignment algorithm that is robust to symmetries in object geometry without requiring their explicit enumeration. We share pre-trained pose estimators for all the toy grocery objects, along with their baseline performance on both validation and test sets. We offer this dataset to the community to help connect the efforts of computer vision researchers with the needs of roboticists. 1 1 https://github.com/swtyree/hope-dataset

ICRA Conference 2022 Conference Paper

Keypoint-Based Category-Level Object Pose Tracking from an RGB Sequence with Uncertainty Estimation

  • Yunzhi Lin
  • Jonathan Tremblay
  • Stephen Tyree
  • Patricio A. Vela
  • Stan Birchfield

We propose a single-stage, category-level 6-DoF pose estimation algorithm that simultaneously detects and tracks instances of objects within a known category. Our method takes as input the previous and current frame from a monocular RGB video, as well as predictions from the previous frame, to predict the bounding cuboid and 6- DoF pose (up to scale). Internally, a deep network predicts distributions over object keypoints (vertices of the bounding cuboid) in image coordinates, after which a novel probabilistic filtering process integrates across estimates before computing the final pose using PnP. Our framework allows the system to take previous uncertainties into consideration when predicting the current frame, resulting in predictions that are more accurate and stable than single frame methods. Extensive experiments show that our method outperforms existing approaches on the challenging Objectron benchmark of annotated object videos. We also demonstrate the usability of our work in an augmented reality setting.

ICRA Conference 2022 Conference Paper

PredictionNet: Real-Time Joint Probabilistic Traffic Prediction for Planning, Control, and Simulation

  • Alexey Kamenev
  • Lirui Wang
  • Ollin Boer Bohan
  • Ishwar Kulkarni
  • Bilal Kartal
  • Artem Molchanov
  • Stan Birchfield
  • David Nistér

Predicting the future motion of traffic agents is crucial for safe and efficient autonomous driving. To this end, we present PredictionNet, a deep neural network (DNN) that predicts the motion of all surrounding traffic agents together with the ego-vehicle's motion. All predictions are probabilistic and are represented in a simple top-down rasterization that allows an arbitrary number of agents. Conditioned on a multi-layer map with lane information, the network outputs future positions, velocities, and backtrace vectors jointly for all agents including the ego-vehicle in a single pass. Trajectories are then extracted from the output. The network can be used to simulate realistic traffic, and it produces competitive results on popular benchmarks. More importantly, it has been used to successfully control a real-world vehicle for hundreds of kilometers, by combining it with a motion planning/control subsystem. The network runs faster than real-time on an embedded GPU, and the system shows good generalization (across sensory modalities and locations) due to the choice of input representation. Furthermore, we demonstrate that by extending the DNN with reinforcement learning (RL), it can better handle rare or unsafe events like aggressive maneuvers and crashes.

ICRA Conference 2022 Conference Paper

Single-Stage Keypoint- Based Category-Level Object Pose Estimation from an RGB Image

  • Yunzhi Lin
  • Jonathan Tremblay
  • Stephen Tyree
  • Patricio A. Vela
  • Stan Birchfield

Prior work on 6-DoF object pose estimation has largely focused on instance-level processing, in which a textured CAD model is available for each object being detected. Category-level 6- DoF pose estimation represents an important step toward developing robotic vision systems that operate in unstructured, real-world scenarios. In this work, we propose a single-stage, keypoint-based approach for category-level object pose estimation that operates on unknown object instances within a known category using a single RGB image as input. The proposed network performs 2D object detection, detects 2D keypoints, estimates 6- DoF pose, and regresses relative bounding cuboid dimensions. These quantities are estimated in a sequential fashion, leveraging the recent idea of convGRU for propagating information from easier tasks to those that are more difficult. We favor simplicity in our design choices: generic cuboid vertex coordinates, single-stage network, and monocular RGB input. We conduct extensive experiments on the challenging Objectron benchmark, outperforming state-of-the-art methods on the 3D IoU metric (27. 6% higher than the MobilePose single-stage approach and 7. 1 % higher than the related two-stage approach).

ICRA Conference 2021 Conference Paper

Fast Uncertainty Quantification for Deep Object Pose Estimation

  • Guanya Shi
  • Yifeng Zhu
  • Jonathan Tremblay
  • Stan Birchfield
  • Fabio Ramos 0001
  • Anima Anandkumar
  • Yuke Zhu

Deep learning-based object pose estimators are often unreliable and overconfident especially when the input image is outside the training domain, for instance, with sim2real transfer. Efficient and robust uncertainty quantification (UQ) in pose estimators is critically needed in many robotic tasks. In this work, we propose a simple, efficient, and plug-and-play UQ method for 6-DoF object pose estimation. We ensemble 2–3 pre-trained models with different neural network architectures and/or training data sources, and compute their average pair-wise disagreement against one another to obtain the uncertainty quantification. We propose four disagreement metrics, including a learned metric, and show that the average distance (ADD) is the best learning-free metric and it is only slightly worse than the learned metric, which requires labeled target data. Our method has several advantages compared to the prior art: 1) our method does not require any modification of the training process or the model inputs; and 2) it needs only one forward pass for each model. We evaluate the proposed UQ method on three tasks where our uncertainty quantification yields much stronger correlations with pose estimation errors than the baselines. Moreover, in a real robot grasping task, our method increases the grasping success rate from 35% to 90%. Video and code are available at https://sites.google.com/view/fastuq.

ICRA Conference 2021 Conference Paper

Hierarchical Planning for Long-Horizon Manipulation with Geometric and Symbolic Scene Graphs

  • Yifeng Zhu
  • Jonathan Tremblay
  • Stan Birchfield
  • Yuke Zhu

We present a visually grounded hierarchical planning algorithm for long-horizon manipulation tasks. Our algorithm offers a joint framework of neuro-symbolic task planning and low-level motion generation conditioned on the specified goal. At the core of our approach is a two-level scene graph representation, namely geometric scene graph and symbolic scene graph. This hierarchical representation serves as a structured, object-centric abstraction of manipulation scenes. Our model uses graph neural networks to process these scene graphs for predicting high-level task plans and low-level motions. We demonstrate that our method scales to long-horizon tasks and generalizes well to novel task goals. We validate our method in a kitchen storage task in both physical simulation and the real world. Experiments show that our method achieves over 70% success rate and nearly 90% of subgoal completion rate on the real robot while being four orders of magnitude faster in computation time compared to standard search-based task-and-motion planner. 1

IROS Conference 2021 Conference Paper

Joint Space Control via Deep Reinforcement Learning

  • Visak Kumar
  • David Hoeller
  • Balakumar Sundaralingam
  • Jonathan Tremblay
  • Stan Birchfield

The dominant way to control a robot manipulator uses hand-crafted differential equations leveraging some form of inverse kinematics / dynamics. We propose a simple, versatile joint-level controller that dispenses with differential equations entirely. A deep neural network, trained via model-free reinforcement learning, is used to map from task space to joint space. Experiments show the method capable of achieving similar error to traditional methods, while greatly simplifying the process by automatically handling redundancy, joint limits, and acceleration / deceleration profiles. The basic technique is extended to avoid obstacles by augmenting the input to the network with information about the nearest obstacles. Results are shown both in simulation and on a real robot via sim-to-real transfer of the learned policy. We show that it is possible to achieve sub-centimeter accuracy, both in simulation and the real world, with a moderate amount of training.

IROS Conference 2021 Conference Paper

Multi-view Fusion for Multi-level Robotic Scene Understanding

  • Yunzhi Lin
  • Jonathan Tremblay
  • Stephen Tyree
  • Patricio A. Vela
  • Stan Birchfield

We present a system for multi-level scene awareness for robotic manipulation. Given a sequence of camera-inhand RGB images, the system calculates three types of information: 1) a point cloud representation of all the surfaces in the scene, for the purpose of obstacle avoidance. 2) the rough pose of unknown objects from categories corresponding to primitive shapes (e. g. , cuboids and cylinders), and 3) full 6-DoF pose of known objects. By developing and fusing recent techniques in these domains, we provide a rich scene representation for robot awareness. We demonstrate the importance of each of these modules, their complementary nature, and the potential benefits of the system in the context of robotic manipulation.

ICRA Conference 2020 Conference Paper

Camera-to-Robot Pose Estimation from a Single Image

  • Timothy E. Lee
  • Jonathan Tremblay
  • Thang To
  • Jia Cheng
  • Terry Mosier
  • Oliver Kroemer
  • Dieter Fox
  • Stan Birchfield

We present an approach for estimating the pose of an external camera with respect to a robot using a single RGB image of the robot. The image is processed by a deep neural network to detect 2D projections of keypoints (such as joints) associated with the robot. The network is trained entirely on simulated data using domain randomization to bridge the reality gap. Perspective-n-point (PnP) is then used to recover the camera extrinsics, assuming that the camera intrinsics and joint configuration of the robot manipulator are known. Unlike classic hand-eye calibration systems, our method does not require an off-line calibration step. Rather, it is capable of computing the camera extrinsics from a single frame, thus opening the possibility of on-line calibration. We show experimental results for three different robots and camera sensors, demonstrating that our approach is able to achieve accuracy with a single frame that is comparable to that of classic off-line hand-eye calibration using multiple frames. With additional frames from a static pose, accuracy improves even further. Code, datasets, and pretrained models for three widely-used robot manipulators are made available.

ICRA Conference 2020 Conference Paper

DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm System

  • Ankur Handa
  • Karl Van Wyk
  • Wei Yang 0019
  • Jacky Liang
  • Yu-Wei Chao
  • Qian Wan
  • Stan Birchfield
  • Nathan D. Ratliff

Teleoperation offers the possibility of imparting robotic systems with sophisticated reasoning skills, intuition, and creativity to perform tasks. However, teleoperation solutions for high degree-of-actuation (DoA), multi-fingered robots are generally cost-prohibitive, while low-cost offerings usually offer reduced degrees of control. Herein, a low-cost, depth-based teleoperation system, DexPilot, was developed that allows for complete control over the full 23 DoA robotic system by merely observing the bare human hand. DexPilot enabled operators to solve a variety of complex manipulation tasks that go beyond simple pick-and-place operations and performance was measured through speed and reliability metrics. DexPilot cost-effectively enables the production of high dimensional, multi-modality, state-action data that can be leveraged in the future to learn sensorimotor policies for challenging manipulation tasks. The videos of the experiments can be found at https://sites.google.com/view/dex-pilot.

IROS Conference 2020 Conference Paper

Indirect Object-to-Robot Pose Estimation from an External Monocular RGB Camera

  • Jonathan Tremblay
  • Stephen Tyree
  • Terry Mosier
  • Stan Birchfield

We present a robotic grasping system that uses a single external monocular RGB camera as input. The object-to-robot pose is computed indirectly by combining the output of two neural networks: one that estimates the object-to-camera pose, and another that estimates the robot-to-camera pose. Both networks are trained entirely on synthetic data, relying on domain randomization to bridge the sim-to-real gap. Because the latter network performs online camera calibration, the camera can be moved freely during execution without affecting the quality of the grasp. Experimental results analyze the effect of camera placement, image resolution, and pose refinement in the context of grasping several household objects. We also present results on a new set of 28 textured household toy grocery objects, which have been selected to be accessible to other researchers. To aid reproducibility of the research, we offer 3D scanned textured models, along with pre-trained weights for pose estimation.

IROS Conference 2020 Conference Paper

MVLidarNet: Real-Time Multi-Class Scene Understanding for Autonomous Driving Using Multiple Views

  • Ke Chen
  • Ryan Oldja
  • Nikolai Smolyanskiy
  • Stan Birchfield
  • Alexander Popov
  • David Wehr
  • Ibrahim Eden
  • Joachim Pehserl

Autonomous driving requires the inference of actionable information such as detecting and classifying objects, and determining the drivable space. To this end, we present Multi-View LidarNet (MVLidarNet), a two-stage deep neural network for multi-class object detection and drivable space segmentation using multiple views of a single LiDAR point cloud. The first stage processes the point cloud projected onto a perspective view in order to semantically segment the scene. The second stage then processes the point cloud (along with semantic labels from the first stage) projected onto a bird's eye view, to detect and classify objects. Both stages use an encoder-decoder architecture. We show that our multi-view, multi-stage, multi-class approach is able to detect and classify objects while simultaneously determining the drivable space using a single LiDAR scan as input, in challenging scenes with more than one hundred vehicles and pedestrians at a time. The system operates efficiently at 150 fps on an embedded GPU designed for a self-driving car, including a postprocessing step to maintain identities over time. We show results on both KITTI and a much larger internal dataset, thus demonstrating the method's ability to scale by an order of magnitude.

ICRA Conference 2020 Conference Paper

Toward Sim-to-Real Directional Semantic Grasping

  • Shariq Iqbal
  • Jonathan Tremblay
  • Andy Campbell
  • Kirby Leung
  • Thang To
  • Jia Cheng
  • Erik Leitch
  • Duncan McKay

We address the problem of directional semantic grasping, that is, grasping a specific object from a specific direction. We approach the problem using deep reinforcement learning via a double deep Q-network (DDQN) that learns to map downsampled RGB input images from a wrist-mounted camera to Q-values, which are then translated into Cartesian robot control commands via the cross-entropy method (CEM). The network is learned entirely on simulated data generated by a custom robot simulator that models both physical reality (contacts) and perceptual quality (high-quality rendering). The reality gap is bridged using domain randomization. The system is an example of end-to-end (mapping input monocular RGB images to output Cartesian motor commands) grasping of objects from multiple pre-defined object-centric orientations, such as from the side or top. We show promising results in both simulation and the real world, along with some challenges faced and the need for future research in this area.

ICRA Conference 2019 Conference Paper

Robust Learning of Tactile Force Estimation through Robot Interaction

  • Balakumar Sundaralingam
  • Alexander Lambert
  • Ankur Handa
  • Byron Boots
  • Tucker Hermans
  • Stan Birchfield
  • Nathan D. Ratliff
  • Dieter Fox

Current methods for estimating force from tactile sensor signals are either inaccurate analytic models or task-specific learned models. In this paper, we explore learning a robust model that maps tactile sensor signals to force. We specifically explore learning a mapping for the SynTouch BioTac sensor via neural networks. We propose a voxelized input feature layer for spatial signals and leverage information about the sensor surface to regularize the loss function. To learn a robust tactile force model that transfers across tasks, we generate ground truth data from three different sources: (1) the BioTac rigidly mounted to a force torque (FT) sensor, (2) a robot interacting with a ball rigidly attached to the same FT sensor, and (3) through force inference on a planar pushing task by formalizing the mechanics as a system of particles and optimizing over the object motion. A total of 140k samples were collected from the three sources. We achieve a median angular accuracy of 3. 5 degrees in predicting force direction (66% improvement over the current state of the art) and a median magnitude accuracy of 0. 06 N (93% improvement) on a test dataset. Additionally, we evaluate the learned force model in a force feedback grasp controller performing object lifting and gentle placement. Our results can be found on https://sites.google.com/view/tactile-force.

ICRA Conference 2019 Conference Paper

Structured Domain Randomization: Bridging the Reality Gap by Context-Aware Synthetic Data

  • Aayush Prakash
  • Shaad Boochoon
  • Mark Brophy
  • David Acuna
  • Eric Cameracci
  • Gavriel State
  • Omer Shapira
  • Stan Birchfield

We present structured domain randomization (SDR), a variant of domain randomization (DR) that takes into account the structure of the scene in order to add context to the generated data. In contrast to DR, which places objects and distractors randomly according to a uniform probability distribution, SDR places objects and distractors randomly according to probability distributions that arise from the specific problem at hand. In this manner, SDR-generated imagery enables the neural network to take the context around an object into consideration during detection. We demonstrate the power of SDR for the problem of 2D bounding box car detection, achieving competitive results on real data after training only on synthetic data. On the KITTI easy, moderate, and hard tasks, we show that SDR outperforms other approaches to generating synthetic data (VKITTI, Sim 200k, or DR), as well as real data collected in a different domain (BDD100K). Moreover, synthetic SDR data combined with real KITTI data outperforms real KITTI data alone.

ICRA Conference 2018 Conference Paper

Synthetically Trained Neural Networks for Learning Human-Readable Plans from Real-World Demonstrations

  • Jonathan Tremblay
  • Thang To
  • Artem Molchanov
  • Stephen Tyree
  • Jan Kautz
  • Stan Birchfield

We present a system to infer and execute a human-readable program from a real-world demonstration. The system consists of a series of neural networks to perform perception, program generation, and program execution. Leveraging convolutional pose machines, the perception network reliably detects the bounding cuboids of objects in real images even when severely occluded, after training only on synthetic images using domain randomization. To increase the applicability of the perception network to new scenarios, the network is formulated to predict in image space rather than in world space. Additional networks detect relationships between objects, generate plans, and determine actions to reproduce a real-world demonstration. The networks are trained entirely in simulation, and the system is tested in the real world on the pick-and-place problem of stacking colored cubes using a Baxter robot.

IROS Conference 2017 Conference Paper

Toward low-flying autonomous MAV trail navigation using deep neural networks for environmental awareness

  • Nikolai Smolyanskiy
  • Alexey Kamenev
  • Jeffrey Smith 0002
  • Stan Birchfield

We present a micro aerial vehicle (MAV) system, built with inexpensive off-the-shelf hardware, for autonomously following trails in unstructured, outdoor environments such as forests. The system introduces a deep neural network (DNN) called TrailNet for estimating the view orientation and lateral offset of the MAV with respect to the trail center. The DNN-based controller achieves stable flight without oscillations by avoiding overconfident behavior through a loss function that includes both label smoothing and entropy reward. In addition to the TrailNet DNN, the system also utilizes vision modules for environmental awareness, including another DNN for object detection and a visual odometry component for estimating depth for the purpose of low-level obstacle detection. All vision systems run in real time on board the MAV via a Jetson TX1. We provide details on the hardware and software used, as well as implementation details. We present experiments showing the ability of our system to navigate forest trails more robustly than previous techniques, including autonomous flights of 1 km.

ICRA Conference 2014 Conference Paper

An inexpensive method for evaluating the localization performance of a mobile robot navigation system

  • Harsha Kikkeri
  • Gershon Parent
  • Mihai Jalobeanu
  • Stan Birchfield

We propose a method for evaluating the localization accuracy of an indoor navigation system in arbitrarily large environments. Instead of using externally mounted sensors, as required by most ground-truth systems, our approach involves mounting only landmarks consisting of distinct patterns printed on inexpensive foam boards. A pose estimation algorithm computes the pose of the robot with respect to the landmark using the image obtained by an on-board camera. We demonstrate that such an approach is capable of providing accurate estimates of a mobile robot's position and orientation with respect to the landmarks in arbitrarily-sized environments over arbitrarily-long trials. Furthermore, because the approach involves minimal outfitting of the environment, we show that only a small amount of setup time is needed to apply the method to a new environment. Experiments involving a state-of-the-art navigation system demonstrate the ability of the method to facilitate accurate localization measurements over arbitrarily long periods of time.

ICRA Conference 2014 Conference Paper

Fast and accurate PoseSLAM by combining relative and global state spaces

  • Brian Peasley
  • Stan Birchfield

We revisit the question of state space in the context of performing loop closure. Although a relative state space has been previously discounted, we show that such a state space is actually extremely powerful, able to achieve recognizable results after just one iteration. The power behind the technique (called POReSS) is the coupling between parameters that causes the orientation of one node to affect the position and orientation of other nodes. At the same time, the approach is fast because, like the more popular incremental state space, the Jacobian never needs to be explicitly computed. Furthermore, we show that while POReSS is able to quickly compute a solution near the global optimum, it is not precise enough to perform the fine adjustments necessary to reach the global minimum. As a result, we augment POReSS with a fast variant of Gauss-Seidel (called Graph-Seidel) on a global state space to allow the solution to settle closer to the global minimum. We show that this combination of POReSS and Graph-Seidel converges more quickly and scales to very large graphs better than other techniques while at the same time computing a competitive residual.

IROS Conference 2014 Conference Paper

Program synthesis by examples for object repositioning tasks

  • Ashley Feniello
  • Hao Dang
  • Stan Birchfield

We address the problem of synthesizing human-readable computer programs for robotic object repositioning tasks based on human demonstrations. A stack-based domain specific language (DSL) is introduced for object repositioning tasks, and a learning algorithm is proposed to synthesize a program in this DSL based on human demonstrations. Once the synthesized program has been learned, it can be rapidly verified and refined in the simulator via further demonstrations if necessary, then finally executed on an actual robot to accomplish the corresponding learned tasks in the physical world. By performing demonstrations on a novel tablet interface, the time required for teaching is greatly reduced compared with using a real robot. Experiments show a variety of object repositioning tasks such as sorting, kitting, and packaging can be programmed using this approach.

ICRA Conference 2013 Conference Paper

3D non-rigid deformable surface estimation without feature correspondence

  • Bryan Willimon
  • Ian D. Walker
  • Stan Birchfield

We propose an algorithm, that extends our previous work, to estimate the current configuration of a non-rigid object using energy minimization and graph cuts. Our approach removes the need for feature correspondence or texture information and extends the boundary energy term. The object segmentation process is improved by using graph cuts along with a skin detector. We introduce an automatic mesh generator that provides a triangular mesh encapsulating the entire non-rigid object without predefined values. Our approach also handles in-plane rotation by reinitializing the mesh after data has been lost in the image sequence. Results display the proposed algorithm over a dataset consisting of seven shirts, two pairs of shorts, two posters, and a pair of pants.

ICRA Conference 2013 Conference Paper

A new approach to clothing classification using mid-level layers

  • Bryan Willimon
  • Ian D. Walker
  • Stan Birchfield

We present a novel approach for classifying items from a pile of laundry. The classification procedure exploits color, texture, shape, and edge information from 2D and 3D local and global information for each article of clothing using a Kinect sensor. The key contribution of this paper is a novel method of classifying clothing which we term L-M-H, more specifically L-C-S-H using characteristics and selection masks. Essentially, the method decomposes the problem into high (H), low (L) and multiple mid-level (characteristics(C), selection masks(S)) layers and produces “local” solutions to solve the global classification problem. Experiments demonstrate the ability of the system to efficiently classify and label into one of three categories (shirts, socks, or dresses). These results show that, on average, the classification rates, using this new approach with mid-level layers, achieve a true positive rate of 90%.

ICRA Conference 2013 Conference Paper

Replacing Projective Data Association with Lucas-Kanade for KinectFusion

  • Brian Peasley
  • Stan Birchfield

We propose to overcome a significant limitation of the KinectFusion algorithm, namely, its sole reliance upon geometric information to estimate camera pose. Our approach uses both geometric and color information in a direct manner that uses all the data in order to perform the association of data between two RGBD point clouds. Data association is performed by aligning the two color images associated with the two point clouds by estimating a projective warp using the Lucas-Kanade algorithm. This projective warp is then used to create a correspondence map between the two point clouds, which is then used as the data association for a point-to-plane error minimization. This approach to correspondence allows camera tracking to be maintained through areas of low geometric features. We show that our proposed LKDA data association technique enables accurate scene reconstruction in environments in which low geometric texture causes the existing approach to fail, while at the same time demonstrating that the new technique does not adversely affect results in environments in which the existing technique succeeds.

IROS Conference 2012 Conference Paper

Accurate on-line 3D occupancy grids using Manhattan world constraints

  • Brian Peasley
  • Stan Birchfield
  • Alexander Cunningham
  • Frank Dellaert

In this paper we present an algorithm for constructing nearly drift-free 3D occupancy grids of large indoor environments in an online manner. Our approach combines data from an odometry sensor with output from a visual registration algorithm, and it enforces a Manhattan world constraint by utilizing factor graphs to produce an accurate online estimate of the trajectory of a mobile robotic platform. We also examine the advantages and limitations of the octree data structure representation of a 3D environment. Through several experiments in environments with varying sizes and construction we show that our method reduces rotational and translational drift significantly without performing any loop closing techniques.

IROS Conference 2012 Conference Paper

An energy minimization approach to 3D non-rigid deformable surface estimation using RGBD data

  • Bryan Willimon
  • Steven Hickson
  • Ian D. Walker
  • Stan Birchfield

We propose an algorithm that uses energy minimization to estimate the current configuration of a non-rigid object. Our approach utilizes an RGBD image to calculate corresponding SURF features, depth, and boundary information. We do not use predetermined features, thus enabling our system to operate on unmodified objects. Our approach relies on a 3D nonlinear energy minimization framework to solve for the configuration using a semi-implicit scheme. Results show various scenarios of dynamic posters and shirts in different configurations to illustrate the performance of the method. In particular, we show that our method is able to estimate the configuration of a textureless nonrigid object with no correspondences available.

ICRA Conference 2012 Conference Paper

Occlusion-aware reconstruction and manipulation of 3D articulated objects

  • Xiaoxia Huang
  • Ian D. Walker
  • Stan Birchfield

We present a method to recover complete 3D models of articulated objects. Structure-from-motion techniques are used to capture 3D point cloud models of the object in two different configurations. A novel combination of Procrustes analysis and RANSAC facilitates a straightforward geometric approach to recovering the joint axes, as well as classifying them automatically as either revolute or prismatic. With the resulting articulated model, a robotic system is able to manipulate the object along its joint axes at a specified grasp point in order to exercise its degrees of freedom. Because the models capture all sides of the object, they are occluded-aware, enabling the robotic system to plan paths to parts of the object that are not visible in the current view. Our algorithm does not require prior knowledge of the object, nor does it make any assumptions about the planarity of the object or scene. Experiments with a PUMA 500 robotic arm demonstrate the effectiveness of the approach on a variety of objects with both revolute and prismatic joints.

ICRA Conference 2011 Conference Paper

Classification of clothing using interactive perception

  • Bryan Willimon
  • Stan Birchfield
  • Ian D. Walker

We present a system for automatically extracting and classifying items in a pile of laundry. Using only visual sensors, the robot identifies and extracts items sequentially from the pile. When an item has been removed and isolated, a model is captured of the shape and appearance of the object, which is then compared against a database of known items. The classification procedure relies upon silhouettes, edges, and other low-level image measurements of the articles of clothing. The contributions of this paper are a novel method for extracting articles of clothing from a pile of laundry and a novel method of classifying clothing using interactive perception. Experiments demonstrate the ability of the system to efficiently classify and label into one of six categories (pants, shorts, short-sleeve shirt, long-sleeve shirt, socks, or underwear). These results show that, on average, classification rates using robot interaction are 59% higher than those that do not use interaction.

IROS Conference 2011 Conference Paper

Model for unfolding laundry using interactive perception

  • Bryan Willimon
  • Stan Birchfield
  • Ian D. Walker

We present an algorithm for automatically unfolding a piece of clothing. A piece of laundry is pulled in different directions at various points of the cloth in order to flatten the laundry. The features of the cloth are extracted and calculated to determine a valid location and orientation in which to interact with it. The features include the peak region, corner locations, and continuity / discontinuity of the cloth. In this paper we present a two-stage algorithm, introducing a novel solution to the unfolding / flattening problem using interactive perception. Simulations using 3D simulation software, and experiments with robot hardware demonstrate the ability of the algorithm to flatten pieces of laundry using different starting configurations. These results show that, at most, the algorithm flattens out a piece of cloth from 11. 1% to 95. 6% of the canonical configuration.

IROS Conference 2010 Conference Paper

Image-based segmentation of indoor corridor floors for a mobile robot

  • Yinxiao Li
  • Stan Birchfield

We present a novel method for image-based floor detection from a single image. In contrast with previous approaches that rely upon homographies, our approach does not require multiple images (either stereo or optical flow). It also does not require the camera to be calibrated, even for lens distortion. The technique combines three visual cues for evaluating the likelihood of horizontal intensity edge line segments belonging to the wall-floor boundary. The combination of these cues yields a robust system that works even in the presence of severe specular reflections, which are common in indoor environments. The nearly real-time algorithm is tested on a large database of images collected in a wide variety of conditions, on which it achieves nearly 90% detection accuracy.

IROS Conference 2010 Conference Paper

Rigid and non-rigid classification using interactive perception

  • Bryan Willimon
  • Stan Birchfield
  • Ian D. Walker

Robotics research tends to focus upon either non-contact sensing or machine manipulation, but not both. This paper explores the benefits of combining the two by addressing the problem of classifying unknown objects, such as found in service robot applications. In the proposed approach, an object lies on a flat background, and the goal of the robot is to interact with and classify each object so that it can be studied further. The algorithm considers each object to be classified using color, shape, and flexibility. Experiments on a number of different objects demonstrate the ability of efficiently classifying and labeling each item through interaction.

IROS Conference 2007 Conference Paper

Person following with a mobile robot using binocular feature-based tracking

  • Zhichao Chen
  • Stan Birchfield

We present the Binocular Sparse Feature Segmentation (BSFS) algorithm for vision-based person following with a mobile robot. BSFS uses Lucas-Kanade feature detection and matching in order to determine the location of the person in the image and thereby control the robot. Matching is performed between two images of a stereo pair, as well as between successive video frames. We use the Random Sample Consensus (RANSAC) scheme for segmenting the sparse disparity map and estimating the motion models of the person and background. By fusing motion and stereo information, BSFS handles difficult situations such as dynamic backgrounds, out-of-plane rotation, and similar disparity and/or motion between the person and background. Unlike color-based approaches, the person is not required to wear clothing with a different color from the environment. Our system is able to reliably follow a person in complex dynamic, cluttered environments in real time.

ICRA Conference 2006 Conference Paper

Qualitative Vision-based Mobile Robot Navigation

  • Zhichao Chen
  • Stan Birchfield

We present a novel, simple algorithm for mobile robot navigation. Using a teach-replay approach, the robot is manually led along a desired path in a teaching phase, then the robot autonomously follows that path in a replay phase. The technique requires a single off-the-shelf, forward-looking camera with no calibration (including no calibration for lens distortion). Feature points are automatically detected and tracked throughout the image sequence, and the feature coordinates in the replay phase are compared with those computed previously in the teaching phase to determine the turning commands for the robot. The algorithm is entirely qualitative in nature, requiring no map of the environment, no image Jacobian, no homography, no fundamental matrix, and no assumption about a flat ground plane. Experimental results demonstrate the capability of autonomous navigation in both indoor and outdoor environments, on both flat and slanted surfaces, with dynamic occluding objects, for distances over 100 m

v2026.09.13