Arrow Research search

Author name cluster

Dezhen Song

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

62 papers
2 author rows

Possible papers

62

ICRA Conference 2025 Conference Paper

A Full-Optical Pretouch Dual-Modal and Dual-Mechanism (PDM 2 ) Sensor for Robotic Grasping

  • Cheng Fang
  • Zhiyu Yan
  • Fengzhi Guo
  • Shuangliang Li
  • Dezhen Song
  • Jun Zou

We report a new full-optical pretouch dual-modal and dual-mechanism (PDM2) sensor based on an air-coupled fiber-tip surface micromachined optical ultrasound transducer (SMOUT). Compared to ring-shaped piezoelectric acoustic receivers in previous PDM 2 sensors, the acoustic signal received by the new fiber-tip SMOUT is readout optically, which is naturally resistant to surrounding electromagnetic interference (EMI) and makes the complex grounding and shielding unnecessary. In addition, the new fiber-tip SMOUT receiver has a much smaller size, which makes it possible to further miniaturize the sensor package into a more compact structure. For verification, a prototype of the full-optical PDM 2 sensor has been designed, fabricated, and characterized. The experimental results show that even with the much smaller acoustic receiver, the new sensor can still achieve ranging and material/structure sensing performances comparable with the previous ones. Therefore, the new full optical PDM 2 sensor design is promising in providing a practical and miniaturized solution for ranging and material/structure sensing to assist robotic grasping of unknown objects.

ICRA Conference 2025 Conference Paper

Energy Efficient Planning for Repetitive Heterogeneous Tasks in Precision Agriculture

  • Shuangyu Xie
  • Ken Goldberg
  • Dezhen Song

Robotic weed removal in precision agriculture introduces a repetitive heterogeneous task planning (RHTP) challenge for a mobile manipulator. RHTP has two unique characteristics: 1) an observe-first-and-manipulate-later (OFML) temporal constraint that forces a unique ordering of two different tasks for each target and 2) energy savings from efficient task collocation to minimize unnecessary movements. RHTP can be framed as a stochastic renewal process. According to the Renewal Reward Theorem, the expected energy usage per task cycle is the long-run average. Traditional task and motion planning focuses on feasibility rather than optimality due to the unknown object and obstacle position prior to execution. However, the known target/obstacle distribution in precision agriculture allows minimizing the expected energy usage. For each instance in this renewal process, we first compute task space partition, a novel data structure that computes all possibilities of task multiplexing and its probabilities with robot reachability. Then we propose a region-based setcoverage problem to formulate the RHTP as a mixed-integer nonlinear programming. We have implemented and solved RHTP using Branch-and-Bound solver. Compared to a baseline in simulations based on real field data, the results suggest a significant improvement in path length, number of robot stops, overall energy usage, and number of replans.

ICRA Conference 2025 Conference Paper

Heterogeneous Sensor Fusion and Active Perception for Transparent Object Reconstruction with a PDM 2 Sensor and a Camera

  • Fengzhi Guo
  • Shuangyu Xie
  • Di Wang 0020
  • Cheng Fang
  • Jun Zou
  • Dezhen Song

Transparent household objects present a challenge for domestic service robots, since neither regular cameras nor RGB-D cameras can provide accurate points for shape reconstruction. The new type of pretouch dual-modality distance and material sensor (PDM 2 ) can provide reliable and accurate depth readings, but it is a point sensor and scanning the object exclusively with the sensor is too inefficient. Hence, we present a sensor fusion approach by combining a regular camera with the PDM 2 sensor. The approach is based on a data fusion algorithm for shape reconstruction and an active perception algorithm for scan planning for the PDM 2 sensor. The data fusion algorithm is a distributed Gaussian process (GP)-based shape reconstruction method that allows for incremental local update to reduce computational time. The active perception algorithm is an optimization-based approach by increasing the information gain (IG) and prioritizing the boundary points under a preset travel distance constraint. We have implemented and tested the algorithms with six different transparent household items. The results show satisfactory shape reconstruction results in all test cases with an average increase in intersection over union (IoU) from 0. 73 to 0. 96.

IROS Conference 2025 Conference Paper

Simulating Automotive Radar with Lidar and Camera Inputs

  • Peili Song
  • Dezhen Song
  • Yifan Yang
  • Enfan Lan
  • Jingtai Liu

Low-cost millimeter automotive radar has received more and more attention due to its ability to handle adverse weather and lighting conditions in autonomous driving. However, the lack of quality datasets hinders research and development. We report a new method that is able to simulate 4D millimeter wave radar signals including pitch, yaw, range, and Doppler velocity along with radar signal strength (RSS) using camera image, light detection and ranging (lidar) point cloud, and ego-velocity. The method is based on two new neural networks: 1) DIS-Net, which estimates the spatial distribution and number of radar signals, and 2) RSS-Net, which predicts the RSS of the signal based on appearance and geometric information. We have implemented and tested our method using open datasets from 3 different models of commercial automotive radar. The experimental results show that our method can successfully generate high-fidelity radar signals. Moreover, we have trained a popular object detection neural network with data augmented by our synthesized radar. The network outperforms the counterpart trained only on raw radar data, a promising result to facilitate future radar-based research and development.

TMLR Journal 2025 Journal Article

TESGNN: Temporal Equivariant Scene Graph Neural Networks for Efficient and Robust Multi-View 3D Scene Understanding

  • Pham Phuoc Minh Quang
  • Nguyen Tiet Nguyen Khoi
  • Ngo Chi Lan
  • Do Tho Truong
  • Dezhen Song
  • Truong-Son Hy

Scene graphs have proven to be highly effective for various scene understanding tasks due to their compact and explicit representation of relational information. However, current methods often overlook the critical importance of preserving symmetry when generating scene graphs from 3D point clouds, which can lead to reduced accuracy and robustness, particularly when dealing with noisy, multi-view data. Furthermore, a major limitation of prior approaches is the lack of temporal modeling to capture time-dependent relationships among dynamically evolving entities in a scene. To address these challenges, we propose Temporal Equivariant Scene Graph Neural Network (TESGNN), consisting of two key components: (1) an Equivariant Scene Graph Neural Network (ESGNN), which extracts information from 3D point clouds to generate scene graph while preserving crucial symmetry properties, and (2) a Temporal Graph Matching Network, which fuses scene graphs generated by ESGNN across multiple time sequences into a unified global representation using an approximate graph-matching algorithm. Our combined architecture TESGNN shown to be effective compared to existing methods in scene graph generation, achieving higher accuracy and faster training convergence. Moreover, we show that leveraging the symmetry-preserving property produces a more stable and accurate global scene representation compared to existing approaches. Finally, it is computationally efficient and easily implementable using existing frameworks, making it well-suited for real-time applications in robotics and computer vision. This approach paves the way for more robust and scalable solutions to complex multi-view scene understanding challenges.

IROS Conference 2025 Conference Paper

Towards Safe Imitation Learning via Potential Field-Guided Flow Matching

  • Haoran Ding
  • Anqing Duan
  • Zezhou Sun
  • Leonel Rozo
  • Noémie Jaquier
  • Dezhen Song
  • Yoshihiko Nakamura

Deep generative models, particularly diffusion and flow matching models, have recently shown remarkable potential in learning complex policies through imitation learning. However, the safety of generated motions remains overlooked, particularly in complex environments with inherent obstacles. In this work, we address this critical gap by proposing Potential Field-Guided Flow Matching Policy (PF2MP), a novel approach that simultaneously learns task policies and extracts obstacle-related information, represented as a potential field, from the same set of successful demonstrations. During inference, PF2MP modulates the flow matching vector field via the learned potential field, enabling safe motion generation. By leveraging these complementary fields, our approach achieves improved safety without compromising task success across diverse environments, such as navigation tasks and robotic manipulation scenarios. We evaluate PF2MP in both simulation and real-world settings, demonstrating its effectiveness in task space and joint space control. Experimental results demonstrate that PF2MP enhances safety, achieving a significant reduction of collisions compared to baseline policies. This work paves the way for safer motion generation in unstructured and obstacle-rich environments.

ICRA Conference 2024 Conference Paper

Coupled Active Perception and Manipulation Planning for a Mobile Manipulator in Precision Agriculture Applications

  • Shuangyu Xie
  • Chengsong Hu
  • Di Wang 0020
  • Joe Johnson
  • Muthukumar Bagavathiannan
  • Dezhen Song

A mobile manipulator often finds itself in an application where it needs to take a close-up view before performing a manipulation task. Named this as a coupled active perception and manipulation (CAPM) problem, we model the uncertainty in the perception process and devise a key state/task planning algorithm that considers reachability conditions jointly established from perception and manipulation task constraints. By minimizing expected energy usage in body key state planning while satisfying task constraints, our algorithm is able to find an energy-efficient trajectory with less body repositioning motion while ensuring the success of the task. We have implemented the algorithm and tested it in both simulation and physical experiments. The results have confirmed that our algorithm has a lower energy consumption compared to a two-stage decoupled approach, while still maintaining a success rate of 100% for the task.

IROS Conference 2024 Conference Paper

Road Boundary Estimation Using Sparse Automotive Radar Inputs

  • Aaron Kingery
  • Dezhen Song

Low-cost millimeter wavelength automotive radar can work effectively under low visibility or low reflection conditions caused by lighting, weather, pollution, or object surface properties when a camera or a lidar may fail. It can serve as a fallback solution to improve safety in autonomous driving. However, after filtering, radar signals tend to be sparse and noisy which poses new challenges in scene understanding. This paper presents a new approach to detecting road boundaries based on sparse radar signals. We model the roadway using a homogeneous model and derive its conditional predictive model under known radar motion. Using this predictive model and modeling radar points using a Dirichlet Process Mixture Model, we employ Mean Field Variational Inference (MFVI) to derive an unconditional road boundary model distribution. To generate initial candidate solutions for the MFVI, we develop a custom Random Sample and Consensus (RANSAC) variant to propose unseen model instances as candidate road boundaries. For each radar point cloud we alternate the MFVI and RANSAC proposal steps until convergence to generate the best estimate of all candidate models. We select the candidate model with the minimum lateral distance to the radar on each side as the estimates of the left and right boundaries. We have implemented the proposed algorithm and it has shown satisfactory results. More specifically, the mean lane boundary estimation error is not more than 11. 0 cm.

ICRA Conference 2024 Conference Paper

Subsurface Feature-based Ground Robot/Vehicle Localization Using a Ground Penetrating Radar

  • Haifeng Li 0008
  • Jiajun Guo
  • Dezhen Song

Robot localization using subsurface features captured by Ground-Penetrating Radar (GPR) complements and improves robustness over existing common sensor modalities, as subsurface features are less sensitive to weather, season and surface scene changes. Here, we propose a novel subsurface feature-based localization method that uses only GPR measurements with a known subsurface map. An efficient feature descriptor, the dominant energy curve (DEC), is designed to identify different locations in cluttered conditions. Specifically, image processing techniques that involve background segmentation, energy point detection, and energy curve refinement are designed to extract DEC features from a 2D radargram. With DECs features obtained, a metric subsurface feature map is constructed. Finally, we perform robot localization by feature matching under a particle swarm optimization framework. We have implemented our method and tested it with the public CMU-GPR dataset. The results show that our algorithm improves accuracy and robustness with real-time performance for robot localization tasks. Specifically, the mean localization error is 0. 50 m for all cases.

IROS Conference 2024 Conference Paper

Toward Precise Robotic Weed Flaming Using a Mobile Manipulator with a Blowtorch

  • Di Wang 0020
  • Chengsong Hu
  • Shuangyu Xie
  • Joe Johnson
  • Hojun Ji
  • Yingtao Jiang
  • Muthukumar Bagavathiannan
  • Dezhen Song

Robotic weed flaming is a new and environmentally friendly approach to weed removal in the agricultural field. Using a mobile manipulator equipped with a blowtorch, we design a new system and algorithm to enable effective weed flaming, which requires robotic manipulation with a soft and deformable end effector, as the thermal coverage of the flame is affected by dynamic or unknown environmental factors such as gravity, wind, atmospheric pressure, fuel tank pressure, and pose of the nozzle. System development includes overall design, hardware integration, and software pipeline. To enable precise weed removal, the greatest challenge is to detect and predict dynamic flame coverage in real time before motion planning, which is quite different from a conventional rigid gripper in grasping or a spray gun in painting. Based on the images from two onboard infrared cameras and the pose information of the blowtorch nozzle on a mobile manipulator, we propose a new dynamic flame coverage model. The flame model uses a center-arc curve with a Gaussian cross-section model to describe the flame coverage in real time. The experiments have demonstrated the working system and shown that our model and algorithm can achieve a mean average precision (mAP) of more than 76% in the reprojected images during online prediction.

IROS Conference 2023 Conference Paper

A Pretouch Perception Algorithm for Object Material and Structure Mapping to Assist Grasp and Manipulation Using a DMDSM Sensor

  • Fengzhi Guo
  • Shuangyu Xie
  • Di Wang 0020
  • Cheng Fang
  • Jun Zou
  • Dezhen Song

We report a new material and structure mapping (MSM) algorithm to assist robotic grasping and manipulation. Building on our new sensor development, the algorithm has four main components: 1) detection of time-of-flight (ToF) durations for the dual modalities of optoacoustic (OA) and pulse-echo ultrasound (US), 2) contour reconstruction by fusing OA and US signals, 3) local noise filtering by checking local consistency of material and structure label (MSL), and 4) medium boundary searching that identifies class boundaries through two-staged clustering and boundary establishment using support vector machine (SVM) hyperplanes. We have implemented our algorithm and tested it with multiple common household items. The experimental results have successfully validated our algorithm design which shows that the average error of contour reconstruction is 0. 05 mm and the true positive rate of MSL is over 98%.

ICRA Conference 2023 Conference Paper

The Third Generation (G3) Dual-Modal and Dual Sensing Mechanisms (DMDSM) Pretouch Sensor for Robotic Grasping

  • Cheng Fang
  • Shuangliang Li
  • Di Wang 0020
  • Fengzhi Guo
  • Dezhen Song
  • Jun Zou

Fingertip-mounted pretouch sensors are very useful for robotic grasping. In this paper, we report a new (G3) dual-modal and dual sensing mechanisms (DMDSM) pretouch sensor for near-distance ranging and material sensing, which is based on pulse-echo ultrasound (US) and optoacoustics (OA). Different from previously reported versions, the G3 sensor utilizes a self-focused US/OA transceiver, thereby eliminating the need of a bulky parabolic reflective mirror for focusing the ultrasound and laser beams. The self-focused laser and ultrasound beams can be easily steered by a (flat) scanning mirror which expands from single-point ranging and detection to areal mapping or imaging. To verify the new design, a prototype G3 DMDSM sensor with a scanning mirror is fabricated. The US and OA ranging performances are tested in experiments. Together with the scanning mirror, thin wire targets made of same or different materials at different positions are scanned and imaged. The ranging and imaging results show that the G3 DMDSM sensor can provide new and better pretouch mapping and imaging capabilities for robotic grasping than its predecessors.

ICRA Conference 2022 Conference Paper

The Second Generation (G2) Fingertip Sensor for Near-Distance Ranging and Material Sensing in Robotic Grasping

  • Cheng Fang
  • Di Wang 0020
  • Dezhen Song
  • Jun Zou

To continuously improve robotic grasping, we are interested in developing a contactless fingertip-mounted sensor for near-distance ranging and material sensing. Previously, we demonstrated a dual-modal and dual sensing mechanisms (DMDSM) pretouch sensor prototype based on pulse-echo ultrasound and optoacoustics. However, the complex system, the bulky and expensive pulser-receiver, and the omni-directionally sensitive microphone block the sensor from practical applications in real robotic fingers. To address these issues, we report the second generation (G2) DMDSM sensor without the pulser-receiver and microphone, which is made possible by redesigning the ultrasound transmitter and receiver to gain much wider acoustic bandwidth. To verify our design, a prototype of the G2 DMDSM sensor has been fabricated and tested. The testing results show that the G2 DMDSM sensor can achieve better ranging and similar material/structure sensing performance, but with much-simplified configuration and operation. The primary results indicate that the G2 DMDSM sensor could provide a promising solution for fingertip pretouch sensing in robotic grasping.

ICRA Conference 2021 Conference Paper

Device Design and System Integration of a Two-Axis Water-immersible Micro Scanning Mirror (WIMSM) to Enable Dual-modal Optical and Acoustic Communication and Ranging for Underwater Vehicles

  • Xiaoyu Duan
  • Di Wang 0020
  • Dezhen Song
  • Jun Zou

To address the communication and ranging challenges caused by underwater environment, we design dual modal devices for autonomous underwater vehicles (AUVs). The dual-modal design builds upon a co-axial ultrasonic and green laser beams which leverage different signal diverging patterns and different responses in the underwater environment by each modality to achieve robust adaptability. Here we report our recent progress in improving scanning and aiming capabilities for dual-modal beam steering. The core part is our Two-Axis Water-immersible Micro Scanning Mirror (WIMSM). We improve hinge design of WIMSM for larger scanning range. We incorporate high speed Hall effect sensor-based pose feedback channel to enable closed-loop scanning and aiming control. We design ultrasonic-assisted laser handshaking method to help AUVs to acquire optical underwater communication. We have prototyped our devices and tested them in a water tank. The initial results are promising.

ICRA Conference 2021 Conference Paper

Fingertip Pulse-Echo Ultrasound and Optoacoustic Dual-Modal and Dual Sensing Mechanisms Near-Distance Sensor for Ranging and Material Sensing in Robotic Grasping

  • Cheng Fang
  • Di Wang 0020
  • Dezhen Song
  • Jun Zou

To improve robotic grasping, we are interested in developing a new non-contact fingertip-mounted sensor for near-distance ranging and material sensing. Here we report new progress in combining direct pulse-echo ultrasound and optoacoustic effects in sensor design to deal with optically and/or acoustically challenging targets (OACTs). Our dual-modal and dual sensing mechanisms (DMDSM) sensor design is enabled by a novel wideband ultrasound transmitter embedded inside a piezoelectric (lead zirconate titanate - PZT) ring transducer. The new DMDSM sensor is capable of differentiating a variety of OACTs. To verify our design, both distance ranging tests and material sensing tests have been conducted. The ranging tests show the sensor can perform both optoacoustic ranging (for light-absorbing materials) and pulse-echo ultrasound ranging (for reflective or transparent materials). For material sensing, the dual-modal spectra from OACTs are collected to compare the new sensor with previous designs. The overall 100% accuracy from the confusion matrices indicates the initial success of our sensor design in differentiating conventional targets as well as the OACTs with the new DMDSM sensor.

IROS Conference 2020 Conference Paper

Fingertip Non-Contact Optoacoustic Sensor for Near-Distance Ranging and Thickness Differentiation for Robotic Grasping *

  • Cheng Fang
  • Di Wang 0020
  • Dezhen Song
  • Jun Zou

We report the feasibility study of a new optoacoustic sensor for both near-distance ranging and material thickness classification for robotic grasping. It is based on the optoacoustic effect where focused laser pulses are used to generate wideband ultrasound signals in the target. With a much smaller optical focal spot, the optoacoustic sensor achieves a lateral resolution of 93 μm, which is six times higher than ultrasound pulse-echo ranging under the same condition. A new multi-mode wideband PZT (lead zirconate titanate) transducer is built to properly receive the wideband optoacoustic signal. The ability to receive both low- and high-frequency components of the optoacoustic signal enhances the material sensing capability, which makes it promising to determine not only material type but also the sub-surface structures. For demonstration, optoacoustic spectra are collected from hard and soft materials with different thickness. A Bag-of-SFA-Symbols (BOSS) classifier is designed to perform primary material and then thickness classification based on the optoacoustic spectra. The accuracy of material / thickness classification reaches ≥ 99% and ≥ 94%, respectively, which shows the feasibility of differentiating solid materials with different thickness by the optoacoustic sensor.

IROS Conference 2020 Conference Paper

Lane Marking Verification for High Definition Map Maintenance Using Crowdsourced Images

  • Binbin Li 0006
  • Dezhen Song
  • Aaron Kingery
  • Dongfang Zheng
  • Yiliang Xu
  • Huiwen Guo

Autonomous vehicles often rely on high-definition (HD) maps to navigate around. However, lane markings (LMs) are not necessarily static objects due to wear & tear from usage and road reconstruction & maintenance. Therefore, the wrong matching between LMs in the HD map and sensor readings may lead to erroneous localization or even cause traffic accidents. It is imperative to keep LMs up-to-date. However, frequently recollecting data with dedicated hardware and specialists to update HD maps is not only cost-prohibitive but also unviable. Here we propose to utilize crowdsourced images from multiple vehicles at different times to help verify LMs for HD map maintenance. We obtain the LM distribution in the image space by considering the camera pose uncertainty in perspective projection. Both LMs in HD map and LMs in the image are treated as observations of LM distributions which allow us to construct posterior conditional distribution (a. k. a Bayesian belief functions) of LMs from either sources. An LM is consistent if belief functions from the map and the image satisfy statistical hypothesis testing. We further extend the Bayesian belief model into a sequential belief update using crowdsourced images. LMs with a higher probability of existence are kept in the HD map whereas those with a lower probability of existence are removed from the HD map. We verify our approach using real data. Experimental results show that our method is capable of verifying and updating LMs in the HD map.

IROS Conference 2020 Conference Paper

Model Quality Aware RANSAC: A Robust Camera Motion Estimator

  • Shu-Hao Yeh
  • Yan Lu
  • Dezhen Song

Robust estimation of camera motion under the presence of outlier noisevision. Despite existing efforts that focus on detecting motion and scene degeneracies, the best existing approach that builds on Random Consensus Sampling (RANSAC) still has non-negligible failure rate. Since a single failure can lead to the failure of the entire visual simultaneous localization and mapping, it is important to further improve the robust estimation algorithm. We propose a new robust camera motion estimator (RCME) by incorporating two main changes: a model-sample consistency test at the model instantiation step and an inlier set quality test that verifies model-inlier consistency using differential entropy. We have implemented our RCME algorithm and tested it under many public datasets. The results have shown a consistent reduction in failure rate when comparing to the RANSAC-based Gold Standard approach and two recent variations of RANSAC methods.

ICRA Conference 2019 Conference Paper

Gaussian Processes Model-Based Control of Underactuated Balance Robots

  • Kuo Chen
  • Jingang Yi
  • Dezhen Song

Control of underactuated balance robot requires external subsystem trajectory tracking and internal unstable subsystem balancing with limited control authority. We present a learning-based control approach for underactuated balance robots. The tracking and balancing control is designed the controller in fast- and slow-time scales. In the slow-time scale, model predictive control is adopted to plan desired internal state profile to achieve external trajectory tracking task. The internal state is then stabilized around the planned profile in the fast-time scale. The control design is based on a learned Gaussian process (GP) regression model without need of a priori knowledge about the robot dynamics. The controller also incorporates the GP model predicted variance to enhance robustness to modeling errors. Experiments are presented using a Furuta pendulum system.

IROS Conference 2019 Conference Paper

On the Tunable Sparse Graph Solver for Pose Graph Optimization in Visual SLAM Problems

  • Chieh Chou
  • Di Wang 0020
  • Dezhen Song
  • Timothy A. Davis 0001

We report a tunable sparse optimization solver that can trade a slight decrease in accuracy for significant speed improvement in pose graph optimization in visual simultaneous localization and mapping (vSLAM). The solver is designed for devices with significant computation and power constraints such as mobile phones or tablets. Two approaches have been combined in our design. The first is a graph pruning strategy by exploiting objective function structure to reduce the optimization problem size which further sparsifies the optimization problem. The second step is to accelerate each optimization iteration in solving increments for the gradient-based search in Gauss-Newton type optimization solver. We apply a modified Cholesky factorization and reuse the decomposition result from last iteration by using Cholesky update/downdate to accelerate the computation. We have implemented our solver and tested it with open source data. The experimental results show that our solver can be twice as fast as the counterpart while maintaining a loss of less than 5% in accuracy.

ICRA Conference 2019 Conference Paper

Steering Co-centered and Co-directional Optical and Acoustic Beams with a Water-immersible MEMS Scanning Mirror for Underwater Ranging and Communication

  • Xiaoyu Duan
  • Dezhen Song
  • Jun Zou

This paper reports the development of a compact optical-acoustic frontend module for underwater communication and ranging. The module is enabled by a new water-immersible MEMS scanning mirror (WIMSM). It is capable of transmitting, receiving and steering co-centered and co-directional laser and ultrasound beams under water. To monitor its rotating angle in real time, scan position sensors based on Hall effect have been integrated into the WIMSM. The angular alignment of the laser and ultrasound beams in both transmission and reception modes has been examined. The experimental results show that the laser and ultrasound beams can remain aligned with less than 2. 1 degrees under envelope of pan and tilt rotations. This capability is critical for the continuing development of the new bi-modal communication and ranging underwater Vehicles (AUVs).

ICRA Conference 2019 Conference Paper

Toward Fingertip Non-Contact Material Recognition and Near-Distance Ranging for Robotic Grasping

  • Cheng Fang
  • Di Wang 0020
  • Dezhen Song
  • Jun Zou

We report the feasibility study of a new acoustic and optical bi-modal distance & material sensor for robotic grasping. The new sensor is designed to be mounted on the robot fingertip to provide last-moment perception before contact happens. It is based on both pulse-echo ultrasound and optoacoustic effects enabled by single-element air-coupled transducers. In contrast to conventional contact-based and recent pre-touch approaches, this new method overcomes their disadvantages and provides robotic fingers with the capability to detect the distance and material type of the target at a near distance before contact occurs, which is crucial for robust and nimble grasping. The proposed sensor has been tested with different materials, shapes, and porous properties. The experimental results show that this sensor design is functional and practical.

IROS Conference 2019 Conference Paper

Virtual Lane Boundary Generation for Human-Compatible Autonomous Driving: A Tight Coupling between Perception and Planning

  • Binbin Li 0006
  • Dezhen Song
  • Ankit Ramchandani
  • Hsin-Min Cheng
  • Di Wang 0020
  • Yiliang Xu
  • Baifan Chen

Existing autonomous vehicle (AV) navigation algorithms treat lane recognition, obstacle avoidance, local path planning, and lane following as separate functional modules which result in driving behavior that is incompatible with human drivers. It is imperative to design human-compatible navigation algorithms to ensure transportation safety. We develop a new tightly-coupled perception-planning framework that combines all these functionalities to ensure human-compatibility. Using GPS-camera-lidar sensor fusion, we detect actual lane boundaries (ALBs) and propose availability-reasonability-feasibility (ARF) threefold tests to determine if we should generate virtual lane boundaries (VLBs) or follow ALBs. If needed, VLBs are generated using a dynamically adjustable multi-objective optimization framework that considers obstacle avoidance, trajectory smoothness (to satisfy vehicle kinodynamic constraints), trajectory continuity (to avoid sudden movements), GPS following quality (to execute global plan), and lane following or partial direction following (to meeting human expectation). Consequently, vehicle motion is more human compatible than existing approaches. We have implemented our algorithm and tested under open source data with satisfying results.

ICRA Conference 2018 Conference Paper

Encoder-Camera-Ground Penetrating Radar Tri-Sensor Mapping for Surface and Subsurface Transportation Infrastructure Inspection

  • Chieh Chou
  • Aaron Kingery
  • Di Wang 0020
  • Haifeng Li 0008
  • Dezhen Song

We report system and algorithmic development for a sensing suite comprising multiple sensors for both surface and subsurface transportation infrastructure inspection focusing on multi-modal mapping for inspection. The sensing suite contains a camera, a ground penetrating radar (GPR), and a wheel encoder. We design the sensing suite and propose a data collection scheme using customized artificial landmarks (ALs). We use ALs to synchronize two types data streams: camera images that are temporally evenly-spaced and GPR/encoder data that are spatially evenly-spaced. We also employ pose graph optimization with synchronization as penalty functions to further refine synchronization and perform data fusion for 3D reconstruction. We have implemented the system and tested it in physical experiments. The results show that our system successfully fuses three sensory data and product metric 3D reconstruction. The sensor fusion approach reduces the end-to-end distance error from 7. 45cm to 3. 10cm.

IROS Conference 2018 Conference Paper

Lane Marking Quality Assessment for Autonomous Driving

  • Binbin Li 0006
  • Dezhen Song
  • Haifeng Li 0008
  • Adam Pike
  • Paul Carlson

Measuring the quality of roads and ensuring they are ready for autonomous driving is important for future transportation systems. Here we focus on developing metrics and algorithms to assess lane marking (LM)qualities from an egocentric view of an inspection vehicle equipped with a global positioning system (GPS)receiver, a frontal-view camera, and a light detection and ranging (LIDAR)system. We propose three quality metrics for LMs: correctness, shape, and visibility. The correctness metric measures the divergence between the expected LMs based on prior map inputs and the actual sensor inputs. The shape metric evaluates smoothness in road curvature and width range. The visibility metric evaluates the contrast between LMs and background road surfaces. We propose a dual-modal algorithm to compute these metrics. We have implemented the algorithms and tested them under KITTI dataset. The results show that our metrics can successfully detect LM anomalies in all testing scenarios.

IROS Conference 2018 Conference Paper

Robotic Subsurface Pipeline Mapping with a Ground-penetrating Radar and a Camera

  • Haifeng Li 0008
  • Chieh Chou
  • Longfei Fan
  • Binbin Li 0006
  • Di Wang 0020
  • Dezhen Song

We propose a novel subsurface pipeline mapping method by fusing Ground Penetrating Radar (GPR) scans and camera images. To facilitate the simultaneous detection of multiple pipelines, we model the GPR sensing process and prove hyperbola response for general scanning with non-perpendicular angles. Furthermore, we fuse visual simultaneous localization and mapping outputs, encoder readings with GPR scans to classify hyperbolas into different pipeline groups. We extensively apply the J-Linkage method and maximum likelihood estimation to improve algorithm robustness and accuracy. As the result, we optimally estimate the radii and locations of all pipelines. We have implemented our method and tested it in physical experiments with representative pipeline configurations. The results show that our method successfully reconstructs all subsurface pipes. Moreover, the average localization error is 4. 69cm.

IROS Conference 2017 Conference Paper

Mirror-assisted calibration of a multi-modal sensing array with a ground penetrating radar and a camera

  • Chieh Chou
  • Shu-Hao Yeh
  • Dezhen Song

To develop a multi-modal in-traffic bridge deck scanning device, we need to estimate the relative pose between a ground penetrating radar (GPR) and a camera. Unlike camera images, GPR output is in a non-Euclidean coordinate system because it only detects underground objects relative to road surface. When road surface is non-planar, its output cannot be trivially mapped to a 3D Cartesian system which is necessary for sensor fusion. Since there is no joint coverage between two sensors due to mounting requirements, we design an artificial planar bridge assisted by a planar mirror as the calibration rig. We combine the pinhole camera model with mirror reflection transformation and model the GPR imaging process. We estimate the camera and mirror poses and extract readings from hyperbolas generated from metal balls. We employ the maximum likelihood estimator to estimate the rigid body transformation between the two sensors and provide the closed form error analysis. We have conducted physical experiments to validate our calibration process and shown the average error of 6. 67 mm for our calibration model. The result is satisfying considering the GPR signal wave length is 18. 75 cm.

IROS Conference 2016 Conference Paper

Visual programming for mobile robot navigation using high-level landmarks

  • Joseph Lee
  • Yan Lu
  • Yiliang Xu
  • Dezhen Song

We propose a visual programming system that allows users to specify navigation tasks for mobile robots using high-level landmarks in a virtual reality (VR) environment constructed from the output of visual simultaneous localization and mapping (vSLAM). The VR environment provides a Google Street View-like interface for users to familiarize themselves with the robot's working environment, specify high-level landmarks, and determine task-level motion commands related to each landmark. Our system builds a roadmap by using the pose graph from the vSLAM outputs. Based on the roadmap, the high-level landmarks, and task-level motion commands, our system generates an output path for the robot to accomplish the navigation task. We present data structures, architecture, interface, and algorithms for our system and show that, given n s search-type motion commands, our system generates a path in O(n s (n r logn r +m r )) time, where n r and m r are the number of roadmap nodes and edges, respectively. We have implemented our system and tested it on real world data.

ICRA Conference 2015 Conference Paper

A robotic bipedal model for human walking with slips

  • Kuo Chen
  • Mitja Trkov
  • Jingang Yi
  • Yizhai Zhang
  • Tao Liu 0006
  • Dezhen Song

Slip is the major cause of falls in human locomotion. We present a new bipedal modeling approach to capture and predict human walking locomotion with slips. Compared with the existing bipedal models, the proposed slip walking model includes the human foot rolling effects, the existence of the double-stance gait and active ankle joints. One of the major developments is the relaxation of the nonslip assumption that is used in the existing bipedal models. We conduct extensive experiments to optimize the gait profile parameters and to validate the proposed walking model with slips. The experimental results demonstrate that the model successfully predicts the human recovery gaits with slips.

IROS Conference 2015 Conference Paper

Robustness to lighting variations: An RGB-D indoor visual odometry using line segments

  • Yan Lu
  • Dezhen Song

Large lighting variation challenges all visual odometry methods, even with RGB-D cameras. Here we propose a line segment-based RGB-D indoor odometry algorithm robust to lighting variation. We know line segments are abundant indoors and less sensitive to lighting change than point features. However, depth data are often noisy, corrupted or even missing for line segments which are often found on object boundaries where significant depth discontinuities occur. Our algorithm samples depth data along line segments, and uses a random sample consensus approach to identify correct depth and estimate 3D line segments. We analyze 3D line segment uncertainties and estimate camera motion by minimizing the Mahalanobis distance. In experiments we compare our method with two state-of-the-art methods including a keypoint-based approach and a dense visual odometry algorithm, under both constant and varying lighting. Our method demonstrates superior robustness to lighting change by outperforming the competing methods on 6 out of 8 long indoor sequences under varying lighting. Meanwhile our method also achieves improved accuracy even under constant lighting when tested using public data.

ICRA Conference 2014 Conference Paper

High level landmark-based visual navigation using unsupervised geometric constraints in local bundle adjustment

  • Yan Lu
  • Dezhen Song
  • Jingang Yi

We present a high level landmark-based visual navigation approach for a monocular mobile robot. We utilize heterogeneous features, such as points, line segments, lines, planes, and vanishing points, and their inner geometric constraints as the integrated high level landmarks. This is managed through a multilayer feature graph (MFG). Our method extends local bundle adjustment (LBA)-based framework by explicitly exploiting different features and their geometric relationships in an unsupervised manner. The algorithm takes a video stream as input, initializes and incrementally updates MFG based on extracted key frames; it also refines localization and MFG landmarks through the LBA. Physical experiments show that our method can reduce the absolute trajectory error of a traditional point landmark-based LBA method by up to 63. 9%.

IROS Conference 2014 Conference Paper

Planar building facade segmentation and mapping using appearance and geometric constraints

  • Joseph Lee
  • Yan Lu
  • Dezhen Song

Segmentation and mapping of planar building facades (PBFs) can increase a robot's ability of scene understanding and localization in urban environments which are often quasi-rectilinear and GPS-challenged. PBFs are basic components of the quasi-rectilinear environment. We propose a passive vision-based PBF segmentation and mapping algorithm by combining both appearance and geometric constraints. We propose a rectilinear index which allows us to segment out planar regions using appearance data. Then we combine geometric constraints such as reprojection errors, orientation constraints, and coplanarity constraints in an optimization process to improve the mapping of PBFs. We have implemented the algorithm and tested it in comparison with state-of-the-art. The results show that our method can reduce the angular error of scene structure by an average of 82. 82%.

ICRA Conference 2014 Conference Paper

Stationary balance control of a bikebot

  • Yizhai Zhang
  • Pengcheng Wang 0002
  • Jingang Yi
  • Dezhen Song
  • Tao Liu 0006

We present the development of the gyroscopic-balanced control of an autonomous bikebot. The bikebot is an actively controlled bicycle-based robotic platform with a gyro-balancer developed to study human dynamic postural balance motor skills through unstable physical human-robot interactions. We also present a dynamic model and analysis for stationary bikebot. A nonlinear balancing controller is designed to stabilize the underactuated stationary bikebot on an orbital trajectory around the unstable equilibrium point that is coupled with another orbit of the actuated gyro-balancer. We then demonstrate the analysis and control design with experimental validations. Finally, we present a set of human riding experiments to show how the bikebot can be used to perturb and excite human sensorimotor feedback loop for dynamic postural balance motor skills.

ICRA Conference 2014 Conference Paper

Toward featureless visual navigation: Simultaneous localization and planar surface extraction using motion vectors in video streams

  • Wen Li
  • Dezhen Song

Unlike the traditional feature-based methods, we propose using motion vectors (MVs) from video streams as inputs for visual navigation. Although MVs are very noisy and with low spatial resolution, MVs do possess high temporal resolution which means it is possible to merge MVs from different frames to improve signal quality. Homography filtering and MV thresholding are proposed to further improve MV quality so that we can establish plane observations from MVs. We propose an extended Kalman filter (EKF) based approach to simultaneously track robot motion and planes. We formally model error propagation of MVs and derive variance of the merged MVs. We have implemented the proposed method and tested it in physical experiments. Results show that the system is capable of performing robot localization and plane mapping with a relative trajectory error of less than 5. 1%.

ICRA Conference 2013 Conference Paper

Automatic bird species detection using periodicity of salient extremities

  • Wen Li
  • Dezhen Song

To assist nature observation, we develop an automatic bird species filtering method that takes videos from cameras with unknown parameters as input, and outputs likelihood of candidate species. The method recognizes the time series of salient extremities, which is the inter-wing tip distance, performs frequency analysis on periodicity, and provides a species prediction metric using likelihood ratios. To analyze the feasibility of the proposed method, we derive the probability that the salient extremity can be recognized in image for an arbitrary camera perspective. We also prove that the periodicity of the IWTD in the image is the same as the wingbeat frequency in the 3D space regardless of camera parameters with the exception of ignorable degenerated cases. Experiment results validate our analysis and show that the algorithm is very robust to segmentation error and data loss up to 30%.

ICRA Conference 2013 Conference Paper

Decentralized searching of multiple unknown and transient radio sources

  • Chang-Young Kim
  • Dezhen Song
  • Jingang Yi

We develop a decentralized algorithm to coordinate a group of mobile robots to search for unknown and transient radio sources. In addition to limited mobility and ranges of communication and sensing, the robot team has to deal with challenges from signal source anonymity, short transmission duration, and variable transmission power. We propose a two-step approach: first, we decentralize belief functions that robots use to track source locations using checkpoint-based synchronization, and second, we propose a decentralized planning strategy to coordinate robots to ensure the existence of checkpoints. We analyze memory usage, data amount in communication, and searching time for the proposed algorithm. We have implemented the proposed algorithm and compared it with two heuristics. The experiment results show that our algorithm successfully trades a modest amount of memory for the fastest searching time among the three methods.

ICRA Conference 2012 Conference Paper

A two-view based multilayer feature graph for robot navigation

  • Haifeng Li 0008
  • Dezhen Song
  • Yan Lu
  • Jingtai Liu

To facilitate scene understanding and robot navigation in a modern urban area, we design a multilayer feature graph (MFG) based on two views from an on-board camera. The nodes of an MFG are features such as scale invariant feature transformation (SIFT) feature points, line segments, lines, and planes while edges of the MFG represent different geometric relationships such as adjacency, parallelism, collinearity, and coplanarity. MFG also connects the features in two views and the corresponding 3D coordinate system. Building on SIFT feature points and line segments, MFG is constructed using feature fusion which incrementally, iteratively, and extensively verifies the aforementioned geometric relationships using random sample consensus (RANSAC) framework. Physical experiments show that MFG can be successfully constructed in urban area and the construction method is demonstrated to be very robust in identifying feature correspondence.

IROS Conference 2012 Conference Paper

Path planning for clothes climbing robots on deformable clothes surface

  • Yuanyuan Liu
  • Xinyu Wu 0001
  • Dezhen Song
  • Ruiqing Fu
  • Duan Zheng
  • Yangsheng Xu

This paper proposes a novel path planning method for a robot to climb on the deformable clothes surface. Based on the deformable characteristic of the clothes, the tension force of clothes is analyzed and the model of tension degree is established. A clothes climbing robot called Clothbot is composed of a two-wheeled gripper and a 2 Degrees of Freedom (DOF) tail. Based on the locomotion of this robot, the weights of tension degree and the locomotion characteristic are added into the A* algorithm. Combined with the two weights applied, the optimal path to the target for the Clothbot is obtained. The Clothbot has been developed to evaluate the algorithm. The simulation and the experiments have verified the feasibility of this method. In addition, The error state of the movement of the robot which is called side tumbling has been corrected by the motion of the 2-DOF tail.

ICRA Conference 2011 Conference Paper

Balance control and analysis of stationary riderless motorcycles

  • Yizhai Zhang
  • Jingliang Li
  • Jingang Yi
  • Dezhen Song

We present balancing control analysis of a stationary riderless motorcycle. We first present the motorcycle dynamics with an accurate steering mechanism model with consideration of lateral movement of the tire/ground contact point. A nonlinear balance controller is then designed. We estimate the domain of attraction (DOA) of motorcycle dynamics under which the stationary motorcycle can be stabilized by steering. For a typical motorcycle/bicycle configuration, we find that the DOA is relatively small and thus balancing control by only steering at stationary is challenging. The balance control and DOA estimation schemes are validated by experiments conducted on the Rutgers autonomous motorcycle. The attitudes of the motorcycle platform are obtained by a novel estimation scheme that fuses measurements from global positioning systems (GPS) and inertial measurement units (IMU). We also present the experiments of the GPS/IMU-based attitude estimation scheme in the paper.

ICRA Conference 2011 Conference Paper

Localization of multiple unknown transient radio sources using multiple paired mobile robots with limited sensing ranges

  • Chang-Young Kim
  • Dezhen Song
  • Yiliang Xu
  • Jingang Yi

We develop a localization method enabling a team of mobile robots to search for multiple unknown transient radio sources. Due to signal source anonymity, short transmission durations, and dynamic transmission patterns, robots cannot treat the radio sources as continuous radio beacons. Moreover, robots do not know the source transmission power and have limited sensing ranges. To cope with these challenges, we pair up robots and develop a sensing model using the signal strength ratio from the paired robots. We formally prove that the sensed conditional joint posterior probability of source locations for the m-robot team can be obtained by combining the pairwise joint posterior probabilities, which can be derived from signal strength ratios. Moreover, we propose a pairwise ridge walking algorithm (PRWA) to coordinate the robot pairs based on the clustering of high probability regions and the minimization of local Shannon entropy. We have implemented and validated the algorithm under hardware-driven simulation.

ICRA Conference 2011 Conference Paper

Robust recognition of planar mirrored walls using a single view

  • Ali-Akbar Agha-Mohammadi
  • Dezhen Song

We report a method for the detection and recognition of a large planar mirror based on the images captured by a monocular camera. We start with deriving a mirror transformation matrix in a homogeneous coordinate and geometric constraints for corresponding real and virtual feature points in the image. We find that existing feature detection methods are not reflection invariant. We introduce a secondary artificial reflection to virtual features to generate secondary features which are proven to share a rigid body motion relationship with the original feature set. We propose an iterative strategy to adjust the secondary mirror configuration so that existing feature matching methods can be used. The combined method yields a robust mirror detection algorithm which has been verified in physical experiments.

AAAI Conference 2010 Conference Paper

A Low False Negative Filter for Detecting Rare Bird Species from Short Video Segments using a Probable Observation Data Set-based EKF Method

  • Dezhen Song
  • Yiliang Xu

We report a new filter for assisting the search for rare bird species. Since a rare bird only appears in front of the camera with very low occurrence (e. g. less than ten times per year) for very short duration (e. g. less than a fraction of a second), our algorithm must have very low false negative rate. We verify the bird body axis information with the known bird flying dynamics from the short video segment. Since a regular extended Kalman filter (EKF) cannot converge due to high measurement error and limited data, we develop a novel Probable Observation Data Set (PODS)-based EKF method. The new PODS-EKF searches the measurement error range for all probable observation data that ensures the convergence of the corresponding EKF in short time frame. The algorithm has been extensively tested in experiments. The results show that the algorithm achieves 95. 0% area under ROC curve in physical experiment with close to zero false negative rate.

AAAI Conference 2010 Conference Paper

Error Aware Monocular Visual Odometry using Vertical Line Pairs for Small Robots in Urban Areas

  • Ji Zhang
  • Dezhen Song

We report a new error-aware monocular visual odometry method that only uses vertical lines, such as vertical edges of buildings and poles in urban areas as landmarks. Since vertical lines are easy to extract, insensitive to lighting conditions/shadows, and sensitive to robot movements on the ground plane, they are robust features if compared with regular point features or line features. We derive a recursive visual odometry method based on the vertical line pairs. We analyze how errors are propagated and introduced in the continuous odometry process by deriving the closed form representation of covariance matrix. We formulate the minimum variance ego-motion estimation problem and present a method that outputs weights for different vertical line pairs. The resulting visual odometry method is tested in physical experiments and compared with two existing methods that are based on point features and line features, respectively. The experiment results show that our method outperforms its two counterparts in robustness, accuracy, and speed. The relative errors of our method are less than 2% in experiments.

ICRA Conference 2009 Conference Paper

Modeling and motion stability analysis of skid-steered mobile robots

  • Hongpeng Wang
  • Junjie Zhang
  • Jingang Yi
  • Dezhen Song
  • Suhada Jayasuriya
  • Jingtai Liu

Skid-steered mobile robots are widely used because of the simplicity of mechanism and high reliability. However, understanding of the kinematics and dynamics of such a robotic platform is challenging due to the complex wheel/ground interactions and kinematic constraints. In this paper, we attempt to develop a kinematic and dynamic modeling scheme to analyze the skid-steered mobile robot. We model wheel/ground interaction and analyze the robot motion stability. As an application example, we present how to utilize the kinematic and dynamic modeling and analysis for robot localization and slip estimation using only low-cost strapdown inertial measurement units (IMU). The extended Kalman filter (EKF)-based localization scheme incorporates the kinematic constraints. The performance of the EKF-based localization and slip estimation scheme are presented. The estimation methodology is tested and validated on a robotic testbed.

ICRA Conference 2009 Conference Paper

Monte Carlo simultaneous localization of multiple unknown transient radio sources using a mobile robot with a directional antenna

  • Dezhen Song
  • Chang-Young Kim
  • Jingang Yi

We report our system and algorithm developments that enable a single mobile robot equipped with a directional antenna to simultaneously localize multiple unknown transient radio sources. Due to signal source anonymity, short transmission durations, and dynamic transmission patterns, the robot cannot treat the radio sources as continuous radio beacons. We model the radio source behaviors using a novel spatiotemporal probability occupancy grid (SPOG) that captures transient characteristics of radio transmissions and tracks the spatiotemporal posterior probability distribution of the radio transmissions. As a Monte Carlo method, we propose a ridge walking motion planning algorithm that enables the robot to efficiently traverse the high probability regions to accelerate the convergence of the posterior probability distribution. We have implemented the algorithms and the experiment results show that our method consistently outperforms methods such as a random walk or a fixed-route patrol mechanism.

IROS Conference 2009 Conference Paper

On the error analysis of vertical line pair-based monocular visual odometry in urban area

  • Ji Zhang
  • Dezhen Song

When a robot travels in urban area, Global Positional System (GPS) signals might be obstructed by buildings. Hence visual odometry is a choice. We notice that the vertical edges from high buildings and poles of street lights are a very stable set of features that can be easily extracted. Thus, we develop a monocular vision-based odometry system that utilizes the vertical edges from the scene to estimate the robot ego-motion. Since it only takes a single vertical line pair to estimate the robot ego-motion on the road plane, here we model the ego-motion estimation process and analyze how the choice of different vertical line pair impacts the accuracy of the ego-motion estimation process. The resulting closed form error model can assist to choose an appropriate pair of vertical lines to reduce the error in computation. We have implemented the proposed method and validated the error analysis results in physical experiments.

IROS Conference 2009 Conference Paper

Systems and algorithms for autonomously simultaneous observation of multiple objects using robotic PTZ cameras assisted by a wide-angle camera

  • Yiliang Xu
  • Dezhen Song

We report an autonomous observation system with multiple pan-tilt-zoom (PTZ) cameras assisted by a fixed wide-angle camera. The wide-angle camera provides large but low resolution coverage and detects and tracks all moving objects in the scene. Based on the output of the wide-angle camera, the system generates spatiotemporal observation requests for each moving object, which are candidates for close-up views using PTZ cameras. Due to the fact that there are usually much more objects than the number of PTZ cameras, the system first assigns a subset of the requests/objects to each PTZ camera. The PTZ cameras then select the parameter settings that best satisfy the assigned competing requests to provide high resolution views of the moving objects. We solve the request assignment and the camera parameter selection problems in real time. The effectiveness of the proposed system is validated in comparison with an existing work using simulation. The simulation results show that in heavy traffic scenarios, our algorithm increases the number of observed objects by over 200%.

ICRA Conference 2008 Conference Paper

An approximation algorithm for the least overlapping p-Frame problem with non-partial coverage for networked robotic cameras

  • Yiliang Xu
  • Dezhen Song
  • Jingang Yi
  • A. Frank van der Stappen

We report our algorithmic development of the pframe problem that addresses the need of coordinating a set of p networked robotic pan-tilt-zoom cameras for n, (n ≫ p), competing polygonal requests. We assume that the p frames have almost no overlap on the coverage between frames and a request is satisfied only if it is fully covered. We then propose a Resolution Ratio with Non-Partial Coverage (RRNPC) metric to quantify the satisfaction level for a given request with respect to a set of p candidate frames. We propose a latticebased approximation algorithm to search for the solution that maximizes the overall satisfaction. The algorithm builds on an induction-like approach that finds the relationship between the solution to the (p — 1)-frame problem and the solution to the p-frame problem. For a given approximation bound ε, the algorithm runs in O(n/ε 3 +p 2 /ε 6 ) time. We have implemented the algorithm and experimental results are consistent with our complexity analysis.

ICRA Conference 2007 Conference Paper

Adaptive Trajectory Tracking Control of Skid-Steered Mobile Robots

  • Jingang Yi
  • Dezhen Song
  • Junjie Zhang
  • Zane Goodwin

Skid-steered mobile robots have been widely used for terrain exploration and navigation. In this paper, we present an adaptive trajectory control design for a skid-steered wheeled mobile robot. Kinematic and dynamic modeling of the robot is first presented. A pseudo-static friction model is used to capture the interaction between the wheels and the ground. An adaptive control algorithm is designed to simultaneously estimate the wheel/ground contact friction information and control the mobile robot to follow a desired trajectory. A Lyapunov-based convergence analysis of the controller and the estimation of the friction model parameter are presented. Simulation and preliminary experimental results based on a four-wheel robot prototype are demonstrated for the effectiveness and efficiency of the proposed modeling and control scheme

IROS Conference 2007 Conference Paper

IMU-based localization and slip estimation for skid-steered mobile robots

  • Jingang Yi
  • Junjie Zhang
  • Dezhen Song
  • Suhada Jayasuriya

Localization and wheel slip estimation of a skidsteered mobile robot is challenging because of the complex wheel/ground interactions and kinematics constraints. In this paper, we present a localization and slip estimation scheme for a skid-steered mobile robot using low-cost inertial measurement units (IMU). We first analyze the kinematics of the skid-steered mobile robot and present a nonlinear Kalman filter (KF)- based simultaneous localization and slip estimation scheme. The KF-based localization design incorporates the wheel slip estimation and utilizes robot velocity constraints and estimates to overcome the large drift resulting from the integration of the IMU acceleration measurements. The estimation methodology is tested and validated experimentally with a computer visionbased localization system.

IROS Conference 2007 Conference Paper

On-demand sharing of a high-resolution panorama video from networked robotic cameras

  • Ni Qin
  • Dezhen Song

Due to their flexibility in coverage and resolution, networked robotic cameras become more and more popular in applications such as natural observation, security surveillance, and distance learning. Equipped with a high optical zoom lens, a networked robotic camera can generate a giga-pixel-level panorama to cover its viewable region. As new live frames enter the system, this panorama can be updated as a panorama video. User requests are usually not limited to the current camera frame. A user may request a specific region associated with a specific time window. To satisfy different spatiotemporal requests for multiple concurrent users, we present systems and algorithms to allow the on-demand sharing of the high-resolution panorama video. The high-resolution panorama video is encoded into a patch-based representation to allow efficient storage and on-demand content delivery. We present system architecture, user interface, data representation, and encoding/decoding algorithms. In the experiment, we have implemented the system using the MPEG-2 codec. Experimental results show that our system can not only satisfy different spatiotemporal queries but also significantly reduce computation time and communication bandwidth requirement.

ICRA Conference 2007 Conference Paper

Scheduling Analysis of Cluster Tools with Buffer/Process Modules

  • Jingang Yi
  • Shengwei Ding
  • Dezhen Song
  • Mike Tao Zhang

Modeling and scheduling of cluster tools are critical to improving the productivity and to enhancing the design of wafer processing flows and equipment for semiconductor manufacturing. In this paper, we extend the decomposition methods in the work of Dwande et al. (2005) for multi-cluster tools with buffer/process modules (BPMs). The computation of the lower-bound cycle time (fundamental period) is presented. Optimality conditions and robot schedules that realize such lower-bound values are then provided using "pull" and "swap" strategies for single-blade and double-blade robots, respectively. The impact of BPMs on throughput and robot schedules is studied. It is found that such an impact depends on the BPM processing time and the cycle times of the decomposed clusters on both sides of BPMs. A chemical vapor deposition (CVD) tool is used as an example of multi-cluster tools to illustrate the proposed method, analysis, and algorithms. The numerical and experimental results demonstrate the effectiveness and efficiency of the algorithms.

ICRA Conference 2006 Conference Paper

A Minimum Variance Calibration Algorithm for Pan-tilt Robotic Cameras in Natural Environments

  • Dezhen Song
  • Ni Qin
  • Ken Goldberg

A new generation of inexpensive robotic pan-tilt cameras can maintain high-resolution panoramic displays of natural environments. However, the pan-tilt mechanisms are imprecise: small errors can produce large errors in the panoramic display. It is thus important to accurately estimate pan-tilt values. We present a new calibration algorithm that does not rely on calibration markers or fixed orthogonal edges which are rarely available in natural scenes. Our calibration algorithm uses image variance density to optimally estimate camera pan and tilt values by incrementally refining image registration using overlapping images from prior frames. Experiments suggest that the new calibration algorithm can reduce calibration error by 81%. In a companion paper, we present a new image registration algorithm based on spherical projection that optimally aligns the resulting frames

ICRA Conference 2006 Conference Paper

Aligning Windows of Live Video From an Imprecise Pan-tilt-zoom Robotic Camera into a Remote Panoramic Display

  • Ni Qin
  • Dezhen Song
  • Ken Goldberg

A pan-tilt-zoom robotic camera can provide detailed live video of selected areas of interest within a large potential viewing field. To provide spatial context for human observers, it is desirable to insert the resulting live video into a large spherical panoramic display representing the entire viewing field. Accurate alignment of the video stream within the panoramic display is difficult due to small errors in the robot pan-tilt values and image distortion due to nonlinear projection. Existing image alignment algorithms cannot keep up with rapid changes in camera position. In this paper, we present a constant-time image alignment algorithm based on spherical projection and projection-invariant selective sampling that accurately registers paired images at 25 frames per second on a standard PC. Experiments suggest that the new alignment algorithm is faster than previous algorithms by a factor four or more. In a companion paper, we present a new calibration algorithm based on image variance density that optimally estimates camera pan-tilt parameters

ICRA Conference 2006 Conference Paper

Trajectory Tracking and Balance Stabilization Control of Autonomous Motorcycles

  • Jingang Yi
  • Dezhen Song
  • Anthony Levandowski
  • Suhada Jayasuriya

We report a new trajectory tracking and balancing control algorithm for an autonomous motorcycle. Building on the existing modeling work of a bicycle, the new dynamic model of the autonomous motorcycle considers the bicycle caster angle and captures the steering effect on the vehicle tracking and balancing. The trajectory tracking control takes an external/internal model decomposition approach. A nonlinear controller is designed to handle the vehicle balancing. The motorcycle balancing is guaranteed by the system internal equilibria calculation and by the trajectory and system dynamics requirements. The proposed control system is validated by numerical simulations, and is based on a real prototype motorcycle system

IROS Conference 2006 Conference Paper

Vision-based Motion Planning for an Autonomous Motorcycle on Ill-Structured Road

  • Dezhen Song
  • Hyun Nam Lee
  • Jingang Yi
  • Anthony Levandowski

We report our development of a vision-based motion planning system for an autonomous motorcycle designed for desert terrain, where uniform road surface and lane markings are not present. The motion planning is based on a vision vector space (V 2 -Space), which is an unitary vector set that represents local collision-free directions in the image coordinate system, V 2 -Space is constructed by extracting the vectors based on the similarity of adjacent pixels, which captures both the color information and the directional information from prior vehicle tire tracks and pedestrian footsteps. We report how V 2 -Space is constructed to reduce the impact of varying lighting conditions in outdoor environments. We also show how V 2 -Space can be used to incorporate vehicle kinematic, dynamic, and time-delay constraints in motion planning to fit the highly dynamic requirements of the motorcycle. The combined algorithm of the V 2 -Space construction and the motion planning runs in O(n) time, where n is the number of pixels in the captured image. Experiments show that our algorithm outputs correct robot motion commands more than 90% of the time

ICRA Conference 2005 Conference Paper

Steady-State Throughput and Scheduling Analysis of Multi-Cluster Tools for Semiconductor Manufacturing: A Decomposition Approach

  • Jingang Yi
  • Shengwei Ding
  • Dezhen Song

Cluster tools are widely used as semiconductor manufacturing equipment. While throughput analysis and scheduling of single-cluster tools have been well-studied, the corresponding research on multi-cluster tools is still at the early stage. This paper analyzes steady-state throughput and scheduling of multi-cluster tools. A decomposition method is utilized to reduce a multi-cluster tool problem to multiple single-cluster tool problems. Existing research on the throughput and scheduling results is then applied to each single-cluster tool. For an M-cluster tool, an O(M) throughput calculation and robot scheduling algorithm is presented. A chemical-mechanical planarization (CMP) polisher is used as an example of the multi-cluster cluster tools to illustrate the proposed decomposition method and algorithms.

ICRA Conference 2004 Conference Paper

An Exact Algorithm Optimizing Coverage-resolution for Automated Satellite Frame Selection

  • Dezhen Song
  • A. Frank van der Stappen
  • Ken Goldberg

Near real time satellite imaging provides timely images of the earth for weather prediction, disaster response, search and rescue, surveillance, and defense applications. As the satellite passes over the earth, camera imaging parameters are changed during each time window based on demand for images, specified as user requested zones in the reachable field of view during that time window. The satellite frame selection (SFS) problem is to find the camera frame parameters that maximize reward during each time window. To automate satellite management, we formalize the SFS problem based on a new reward metric that incorporates both image resolution and coverage. For a set of n client requests we give a series of algorithms, the fastest computes optimal results in O(n/sup 3/) for satellites with continuously variable resolution. We have implemented the algorithms and compare computation speed for all algorithms.

ICRA Conference 2004 Conference Paper

Unsupervised Scoring for Scalable Internet-based Collaborative Teleoperation

  • Ken Goldberg
  • Dezhen Song
  • In Yong Song
  • Jane McGonigal
  • Wei Zheng
  • Dana Plautz

For applications in education and entertainment, scalable Internet-based collaborative teleoperation allows many users simultaneously to share control of a single device. Automated numerical methods that can assess and record performance provide an incentive for users to participate and a means to evaluate individual and group performance. In this paper we describe "unsupervised scoring": a numerical approach to assessment based on clustering and response time. Like unsupervised learning, this approach is based on identifying regularities in the input rather than comparing input with desired output specified by an external supervisor. We present an algorithm for rapidly computing user scores that scales linearly with the number of users. We describe an implemented Java-based user interface incorporating this metric, an application based on the classic Twister game, and results where individual scores are compared with group performance.

IROS Conference 2003 Conference Paper

ShareCam part 1: interface, system architecture, and implementation of a collaboratively controlled robotic Webcam

  • Dezhen Song
  • Ken Goldberg

ShareCam is a robotic pan, tilt, and zoom web-based camera controlled by simultaneous frame requests from online users. Part II describes algorithms. This paper, part I, focuses on the system. Robotic Webcameras are commercially available but currently restrict control only one user at a time. ShareCam introduces a new interface that allows simultaneous control many users. In this Java-based interface, participating users interact desired frames remotely located browsers where users draw desired frames over a fixed panoramic image. User inputs re transmitted back to pair of PC servers that compute optimal camera back to a pair of PC servers that compute optimal camera parameters, servo the camera, and provide a video stream to all users. We describe the system, online experiments, and compare results with two frame selection models based on user "satisfaction", one memoryless and the second based on satisfaction over multiple motion cycles.

IROS Conference 2003 Conference Paper

ShareCam part II: approximate and distributed algorithms for a collaboratively controlled robotic Webcam

  • Dezhen Song
  • Anatol Pashkevich
  • Ken Goldberg

ShareCam is a robotic pan, tilt, and zoom Web-based camera controlled by simultaneous frame requests from online users. Part I describes the system. This paper, part II, focuses on algorithms. The ShareCam problem is to find a camera frame that optimizes a measure of total user satisfaction. We present a grid-based approximation algorithm: given camera frame requests from n users, and approximation bound /spl epsi/, we analyze the trade of between solution quality and processing speed and prove that the algorithm runs in O(n//spl epsi//sup 3/) time. The algorithm can be distributed to run in O(1//spl epsi//sup 3/) time at each client and in O(n + 1//spl epsi//sup 3/) time at the server. Experiments suggest that performance of the distributed algorithm degrades gracefully as clients fail to complete their part of the computation.

ICRA Conference 2002 Conference Paper

Collaborative Online Teleoperation with Spatial Dynamic Voting and a Human "Tele-Actor"

  • Ken Goldberg
  • Dezhen Song
  • Yoek-Nam Khor
  • David Pescovitz
  • Anthony Levandowski
  • Jesse C. Himmelstein
  • Janice Shih
  • Annamarie Ho

Internet-based "online robots" now provide public access to remote locations such as museums and laboratories. The Tele-Actor is a collaborative online teleoperation system for distance learning that allows many students to simultaneously share control of a single mobile resource. Our goal is to preserve the educational advantages of field trips without the drawbacks of group travel. We propose the "spatial dynamic voting" (SDV) interface for multiple operator single robot (MOSR) teleoperation. The SDV collects, displays, and analyzes a sequence of spatial votes from multiple online operators at their Internet browsers. The votes drive the motion of a single mobile robot or human "Tele-Actor". The paper describes Version 3. 0 of the system architecture, SDV interface, algorithms for automated goal selection, and metrics for collaboration and leadership. We report results from a July 2001 field test with 56 remote users.

v2026.09.13