Arrow Research search

Author name cluster

Yipu Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
1 author row

Possible papers

9

IROS Conference 2025 Conference Paper

OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB

  • Yunzhi Lin
  • Yipu Zhao
  • Fu-Jen Chu
  • Xingyu Chen
  • Weiyao Wang 0001
  • Hao Tang
  • Patricio A. Vela
  • Matt Feiszli

To address the challenge of short-term object pose tracking in dynamic environments with monocular RGB input, we introduce a large-scale synthetic dataset Omni-Pose6D, crafted to mirror the diversity of real-world conditions. We additionally present a benchmarking framework for a comprehensive comparison of pose tracking algorithms. We propose a pipeline featuring an uncertainty-aware keypoint refinement network, employing probabilistic modeling to refine pose estimation. Comparative evaluations demonstrate that our approach achieves performance superior to existing baselines on real datasets, underscoring the effectiveness of our synthetic dataset and refinement technique in enhancing tracking precision in dynamic contexts. Our contributions set a new precedent for the development and assessment of object pose tracking methodologies in complex scenes.

ICRA Conference 2021 Conference Paper

Distributed Client-Server Optimization for SLAM with Limited On-Device Resources

  • Yetong Zhang
  • Ming Hsiao
  • Yipu Zhao
  • Jing Dong 0002
  • Jakob J. Engel

Simultaneous localization and mapping (SLAM) is a crucial functionality for exploration robots and virtual/augmented reality (VR/AR) devices. However, some of such devices with limited resources cannot afford the computational or memory cost to run full SLAM algorithms. We propose a general client-server SLAM optimization framework that achieves accurate real-time state estimation on the device with low requirements of on-board resources. The resource-limited device (the client) only works on a small part of the map, and the rest of the map is processed by the server. By sending the summarized information of the rest of map to the client, the on-device state estimation is more accurate. Further improvement of accuracy is achieved in the presence of on-device early loop closures, which enables reloading useful variables from the server to the client. Experimental results from both synthetic and real-world datasets demonstrate that the proposed optimization framework achieves accurate estimation in real-time with limited computation and memory budget of the device.

ICRA Conference 2020 Conference Paper

Closed-Loop Benchmarking of Stereo Visual-Inertial SLAM Systems: Understanding the Impact of Drift and Latency on Tracking Accuracy

  • Yipu Zhao
  • Justin S. Smith
  • Sambhu H. Karumanchi
  • Patricio A. Vela

Visual-inertial SLAM is essential for robot navigation in GPS-denied environments, e. g. indoor, underground. Conventionally, the performance of visual-inertial SLAM is evaluated with open-loop analysis, with a focus on the drift level of SLAM systems. In this paper, we raise the question on the importance of visual estimation latency in closed-loop navigation tasks, such as accurate trajectory tracking. To understand the impact of both drift and latency on visualinertial SLAM systems, a closed-loop benchmarking simulation is conducted, where a robot is commanded to follow a desired trajectory using the feedback from visual-inertial estimation. By extensively evaluating the trajectory tracking performance of representative state-of-the-art visual-inertial SLAM systems, we reveal the importance of latency reduction in visual estimation module of these systems. The findings suggest directions of future improvements for visual-inertial SLAM.

ICRA Conference 2019 Conference Paper

Low-latency Visual SLAM with Appearance-Enhanced Local Map Building

  • Yipu Zhao
  • Wenkai Ye
  • Patricio A. Vela

A local map module is often implemented in modern VO/VSLAM systems to improve data association and pose estimation. Conventionally, the local map contents are determined by co-visibility. While co-visibility is cheap to establish, it utilizes the relatively-weak temporal prior (i. e. seen before, likely to be seen now), therefore admitting more features into the local map than necessary. This paper describes an enhancement to co-visibility local map building by incorporating a strong appearance prior, which leads to a more compact local map and latency reduction in downstream data association. The appearance prior collected from the current image influences the local map contents: only the map features visually similar to the current measurements are potentially useful for data association. To that end, mapped features are indexed and queried with Multi-index Hashing (MIH). An online hash table selection algorithm is developed to further reduce the query overhead of MIH and the local map size. The proposed appearance-based local map building method is integrated into a state-of-the-art VO/VSLAM system. When evaluated on two public benchmarks, the size of the local map, as well as the latency of real-time pose tracking in VO/VSLAM are significantly reduced. Meanwhile, the VO/VSLAM mean performance is preserved or improves.

IROS Conference 2018 Conference Paper

Good Feature Selection for Least Squares Pose Optimization in VO/VSLAM

  • Yipu Zhao
  • Patricio A. Vela

This paper aims to select features that contribute most to the pose estimation in VO/VSLAM. Unlike existing feature selection works that are focused on efficiency only, our method significantly improves the accuracy of pose tracking, while introducing little overhead. By studying the impact of feature selection towards least squares pose optimization, we demonstrate the applicability of improving accuracy via good feature selection. To that end, we introduce the Max-logDet metric to guide the feature selection, which is connected to the conditioning of least squares pose optimization problem. We then describe an efficient algorithm for approximately solving the NP-hard Max-logDet problem. Integrating Max-logDet feature selection into a state-of-the-art visual SLAM system leads to accuracy improvements with low overhead, as demonstrated via evaluation on a public benchmark.

ICRA Conference 2016 Conference Paper

2D-image to 3D-range registration in urban environments via scene categorization and combination of similarity measurements

  • Yipu Zhao
  • Yuanfang Wang
  • Yichang Tsai 0001

Statistical similarity measurements, such as mutual information (MI) and normalized mutual information (NMI), show potential in the registration of 2D-image to 3D-range scans collected in urban environments. However 2D-3D registration with these measurements are of limited usage in urban sensing applications because: 1) it relies on the diversity and dependency between pre-defined pair of 2D-3D attributes, such as the intensity from images and reflectivity from range scans, in the urban sensing environment, and 2) it requires high-end range sensors with strong abilities to capture reflectivity in urban scenarios. In this paper, we propose a robust way of estimating statistical similarity measurements for 2D-3D data that are collected in various urban scenes with both low-cost and high-end range sensors. Rather than estimate the similarity of 2D-3D data on specific pair of 2D-3D attributes, we compute similarity measurements between a set of 2D-3D attribute-pairs that could be dominant in the category of sensed urban scene and combine them into a reliable similarity measurement. By applying the combined similarity measurement to the common framework of statistical 2D-3D registration, we get superior results when compared with state-of-art similarity measurements (MI and NMI) in terms of registration accuracy and robustness to initial condition, as indicated by experiments conducted on two datasets that are collected in various urban scenes, with low-cost and high-end sensors.

ICRA Conference 2012 Conference Paper

Computing object-based saliency in urban scenes using laser sensing

  • Yipu Zhao
  • Mengwen He
  • Huijing Zhao
  • Franck Davoine
  • Hongbin Zha

It becomes a well-known technology that a low-level map of complex environment containing 3D laser points can be generated using a robot with laser scanners. Given a cloud of 3D laser points of an urban scene, this paper proposes a method for locating the objects of interest, e. g. traffic signs or road lamps, by computing object-based saliency. Our major contributions are: 1) a method for extracting simple geometric features from laser data is developed, where both range images and 3D laser points are analyzed; 2) an object is modeled as a graph used to describe the composition of geometric features; 3) a graph matching based method is developed to locate the objects of interest on laser data. Experimental results on real laser data depicting urban scenes are presented; efficiency as well as limitations of the method are discussed.

ICRA Conference 2010 Conference Paper

Scene understanding in a large dynamic environment through a laser-based sensing

  • Huijing Zhao
  • Yiming Liu
  • Xiaolong Zhu
  • Yipu Zhao
  • Hongbin Zha

It became a well known technology that a map of complex environment containing low-level geometric primitives (such as laser points) can be generated using a robot with laser scanners. This research is motivated by the need of obtaining semantic knowledge of a large urban outdoor environment after the robot explores and generates a low-level sensing data set. An algorithm is developed with the data represented in a range image, while each pixel can be converted into a 3D coordinate. Using an existing segmentation method that models only geometric homogeneities, the data of a single object of complex geometry, such as people, cars, trees etc. , is partitioned into different segments. Such a segmentation result will greatly restrict the capability of object recognition. This research proposes a framework of simultaneous segmentation and classification of range image, where the classification of each segment is conducted based on its geometric properties, and homogeneity of each segment is evaluated conditioned on each object class. Experiments are presented using the data of a large dynamic urban outdoor environment, and performance of the algorithm is evaluated.

IROS Conference 2010 Conference Paper

Segmentation and classification of range image from an intelligent vehicle in urban environment

  • Xiaolong Zhu
  • Huijing Zhao
  • Yiming Liu
  • Yipu Zhao
  • Hongbin Zha

As the rapid development of sensing and mapping techniques, it becomes a well-known technology that a map of complex environment can be generated using a robot carrying sensors. However, most of the existing researches represent environments directly using the integration of point clouds or other low-level geometric primitives. It remains an open problem to automatically convert these low-level map representations to semantic descriptions in order to effectively support high-level decision of a robot. Based on another representation of 3D point clouds, i. e. range image, this paper proposes a framework of segmentation and classification of range image, the objective of which is to annotate class labels to the data clusters that are obtained through a graph-based segmentation. Experimental results are presented and evaluated demonstrating that the proposed algorithm has efficiency in understanding the semantic knowledge of a large dynamic urban outdoor environment.

v2026.09.13