Arrow Research search

Author name cluster

In-So Kweon

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

74 papers
2 author rows

Possible papers

74

ICRA Conference 2024 Conference Paper

Complementary Random Masking for RGB-Thermal Semantic Segmentation

  • Ukcheol Shin
  • Kyunghyun Lee 0004
  • In-So Kweon
  • Jean Oh

RGB-thermal semantic segmentation is one potential solution to achieve reliable semantic scene understanding in adverse weather and lighting conditions. However, the previous studies mostly focus on designing a multi-modal fusion module without consideration of the nature of multi-modality inputs. Therefore, the networks easily become over-reliant on a single modality, making it difficult to learn complementary and meaningful representations for each modality. This paper proposes 1) a complementary random masking strategy of RGB-T images and 2) self-distillation loss between clean and masked input modalities. The proposed masking strategy prevents over-reliance on a single modality. It also improves the accuracy and robustness of the neural network by forcing the network to segment and classify objects even when one modality is partially available. Also, the proposed self-distillation loss encourages the network to extract complementary and meaningful representations from a single modality or complementary masked modalities. We achieve state-of-the-art performance over three RGB-T semantic segmentation benchmarks. Our source code is available at https://github.com/UkcheolShin/CRM_RGBTSeg.

ICLR Conference 2022 Conference Paper

Deep Point Cloud Reconstruction

  • Jaesung Choe
  • Byeongin Joung
  • François Rameau
  • Jaesik Park
  • In-So Kweon

Point cloud obtained from 3D scanning is often sparse, noisy, and irregular. To cope with these issues, recent studies have been separately conducted to densify, denoise, and complete inaccurate point cloud. In this paper, we advocate that jointly solving these tasks leads to significant improvement for point cloud reconstruction. To this end, we propose a deep point cloud reconstruction network consisting of two stages: 1) a 3D sparse stacked-hourglass network as for the initial densification and denoising, 2) a refinement via transformers converting the discrete voxels into continuous 3D points. In particular, we further improve the performance of the transformers by a newly proposed module called amplified positional encoding. This module has been designed to differently amplify the magnitude of positional encoding vectors based on the points' distances for adaptive refinements. Extensive experiments demonstrate that our network achieves state-of-the-art performance among the recent studies in the ScanNet, ICL-NUIM, and ShapeNet datasets. Moreover, we underline the ability of our network to generalize toward real-world and unmet scenes.

IROS Conference 2022 Conference Paper

DRL-ISP: Multi-Objective Camera ISP with Deep Reinforcement Learning

  • Ukcheol Shin
  • Kyunghyun Lee 0004
  • In-So Kweon

In this paper, we propose a multi-objective camera ISP framework that utilizes Deep Reinforcement Learning (DRL) and camera ISP toolbox that consist of network-based and conventional ISP tools. The proposed DRL-based camera ISP framework iteratively selects a proper tool from the toolbox and applies it to the image to maximize a given vision task-specific reward function. For this purpose, we implement total 51 ISP tools that include exposure correction, color-and-tone correction, white balance, sharpening, denoising, and the others. We also propose an efficient DRL network architecture that can extract the various aspects of an image and make a rigid mapping relationship between images and a large number of actions. Our proposed DRL-based ISP framework effectively improves the image quality according to each vision task such as RAW-to-RGB image restoration, 2D object detection, and monocular depth estimation.

ICLR Conference 2022 Conference Paper

How Does SimSiam Avoid Collapse Without Negative Samples? A Unified Understanding with Self-supervised Contrastive Learning

  • Chaoning Zhang
  • Kang Zhang 0008
  • Chenshuang Zhang
  • Trung X. Pham
  • Chang D. Yoo
  • In-So Kweon

To avoid collapse in self-supervised learning (SSL), a contrastive loss is widely used but often requires a large number of negative samples. Without negative samples yet achieving competitive performance, a recent work~\citep{chen2021exploring} has attracted significant attention for providing a minimalist simple Siamese (SimSiam) method to avoid collapse. However, the reason for how it avoids collapse without negative samples remains not fully clear and our investigation starts by revisiting the explanatory claims in the original SimSiam. After refuting their claims, we introduce vector decomposition for analyzing the collapse based on the gradient analysis of the $l_2$-normalized representation vector. This yields a unified perspective on how negative samples and SimSiam alleviate collapse. Such a unified perspective comes timely for understanding the recent progress in SSL.

IROS Conference 2021 Conference Paper

Correlate-and-Excite: Real-Time Stereo Matching via Guided Cost Volume Excitation

  • Antyanta Bangunharcana
  • Jae-Won Cho
  • Seokju Lee
  • In-So Kweon
  • Kyung-Soo Kim 0001
  • Soohyun Kim 0001

Volumetric deep learning approach towards stereo matching aggregates a cost volume computed from input left and right images using 3D convolutions. Recent works showed that utilization of extracted image features and a spatially varying cost volume aggregation complements 3D convolutions. However, existing methods with spatially varying operations are complex, cost considerable computation time, and cause memory consumption to increase. In this work, we construct Guided Cost volume Excitation (GCE) and show that simple channel excitation of cost volume guided by image can improve performance considerably. Moreover, we propose a novel method of using top-k selection prior to soft-argmin disparity regression for computing the final disparity estimate. Combining our novel contributions, we present an end-to-end network that we call Correlate-and-Excite (CoEx). Extensive experiments of our model on the SceneFlow, KITTI 2012, and KITTI 2015 datasets demonstrate the effectiveness and efficiency of our model and show that our model outperforms other speed-based algorithms while also being competitive to other state-of-the-art algorithms. Codes will be made available at https://github.com/antabangun/coex.

ICRA Conference 2021 Conference Paper

Stereo Object Matching Network

  • Jaesung Choe
  • Kyungdon Joo
  • François Rameau
  • In-So Kweon

This paper presents a stereo object matching method that exploits both 2D contextual information from images as well as 3D object-level information. Unlike existing stereo matching methods that exclusively focus on the pixel-level correspondence between stereo images within a volumetric space (i. e. , cost volume), we exploit this volumetric structure in a different manner. The cost volume explicitly encompasses 3D information along its disparity axis, therefore it is a privileged structure that can encapsulate the 3D contextual information from objects. However, it is not straightforward since the disparity values map the 3D metric space in a non-linear fashion. Thus, we present two novel strategies to handle 3D objectness in the cost volume space: selective sampling (RoISelect) and 2D-3D fusion (fusion-by-occupancy), which allow us to seamlessly incorporate 3D object-level information and achieve accurate depth performance near the object boundary regions. Our depth estimation achieves competitive performance in the KITTI dataset and the Virtual-KITTI 2. 0 dataset.

AAAI Conference 2020 Conference Paper

CD-UAP: Class Discriminative Universal Adversarial Perturbation

  • Chaoning Zhang
  • Philipp Benz
  • Tooba Imtiaz
  • In-So Kweon

A single universal adversarial perturbation (UAP) can be added to all natural images to change most of their predicted class labels. It is of high practical relevance for an attacker to have flexible control over the targeted classes to be attacked, however, the existing UAP method attacks samples from all classes. In this work, we propose a new universal attack method to generate a single perturbation that fools a target network to misclassify only a chosen group of classes, while having limited influence on the remaining classes. Since the proposed attack generates a universal adversarial perturbation that is discriminative to targeted and non-targeted classes, we term it class discriminative universal adversarial perturbation (CD-UAP). We propose one simple yet effective algorithm framework, under which we design and compare various loss function configurations tailored for the class discriminative universal attack. The proposed approach has been evaluated with extensive experiments on various benchmark datasets. Additionally, our proposed approach achieves state-of-the-art performance for the original task of UAP attacking all classes, which demonstrates the effectiveness of our approach.

ICRA Conference 2020 Conference Paper

CNN-Based Simultaneous Dehazing and Depth Estimation

  • Byeong-Uk Lee
  • Kyunghyun Lee 0004
  • Jean Oh
  • In-So Kweon

It is difficult for both cameras and depth sensors to obtain reliable information in hazy scenes. Therefore, image dehazing is still one of the most challenging problems to solve in computer vision and robotics. With the development of convolutional neural networks (CNNs), lots of dehazing and depth estimation algorithms using CNNs have emerged. However, very few of those try to solve these two problems at the same time. Focusing on the fact that traditional haze modeling contains depth information in its formula, we propose a CNN-based simultaneous dehazing and depth estimation network. Our network aims to estimate both a dehazed image and a fully scaled depth map from a single hazy RGB input with end-to-end training. The network contains a single dense encoder and four separate decoders; each of them shares the encoded image representation while performing individual tasks. We suggest a novel depth-transmission consistency loss in the training scheme to fully utilize the correlation between the depth information and transmission map. To demonstrate the robustness and effectiveness of our algorithm, we performed various ablation studies and compared our results to those of state-of-the-art algorithms in dehazing and single image depth estimation, both qualitatively and quantitatively. Furthermore, we show the generality of our network by applying it to some real-world examples.

ICRA Conference 2020 Conference Paper

Globally Optimal Relative Pose Estimation for Camera on a Selfie Stick

  • Kyungdon Joo
  • Hongdong Li
  • Tae Hyun Oh
  • Yunsu Bok
  • In-So Kweon

Taking selfies has become a photographic trend nowadays. We envision the emergence of the "video selfie" capturing a short continuous video clip (or burst photography) of the user, themselves. A selfie stick is usually used, whereby a camera is mounted on a stick for taking selfie photos. In this scenario, we observe that the camera typically goes through a special trajectory along a sphere surface. Motivated by this observation, in this work, we propose an efficient and globally optimal relative camera pose estimation between a pair of two images captured by a camera mounted on a selfie stick. We exploit the special geometric structure of the camera motion constrained by a selfie stick and define its motion as spherical joint motion. By the new parametrization and calibration scheme, we show that the pose estimation problem can be reduced to a 3-DoF (degrees of freedom) search problem, instead of a generic 6-DoF problem. This allows us to derive a fast branch-and-bound global optimization, which guarantees a global optimum. Thereby, we achieve efficient and robust estimation even in the presence of outliers. By experiments on both synthetic and real-world data, we validate the performance as well as the guaranteed optimality of the proposed method.

ICRA Conference 2020 Conference Paper

Linear RGB-D SLAM for Atlanta World

  • Kyungdon Joo
  • Tae Hyun Oh
  • François Rameau
  • Jean-Charles Bazin
  • In-So Kweon

We present a new linear method for RGB-D based simultaneous localization and mapping (SLAM). Compared to existing techniques relying on the Manhattan world assumption defined by three orthogonal directions, our approach is designed for the more general scenario of the Atlanta world. It consists of a vertical direction and a set of horizontal directions orthogonal to the vertical direction and thus can represent a wider range of scenes. Our approach leverages the structural regularity of the Atlanta world to decouple the non-linearity of camera pose estimations. This allows us separately to estimate the camera rotation and then the translation, which bypasses the inherent non-linearity of traditional SLAM techniques. To this end, we introduce a novel tracking-by-detection scheme to estimate the underlying scene structure by Atlanta representation. Thereby, we propose an Atlanta frame-aware linear SLAM framework which jointly estimates the camera motion and a planar map supporting the Atlanta structure through a linear Kalman filter. Evaluations on both synthetic and real datasets demonstrate that our approach provides favorable performance compared to existing state-of-the-art methods while extending their working range to the Atlanta world.

IROS Conference 2020 Conference Paper

SideGuide: A Large-scale Sidewalk Dataset for Guiding Impaired People

  • Kibaek Park
  • Youngtaek Oh
  • Soomin Ham
  • Kyungdon Joo
  • Hyokyoung Kim
  • Hyoyoung Kum
  • In-So Kweon

In this paper, we introduce a new large-scale sidewalk dataset called SideGuide that could potentially help impaired people. Unlike most previous datasets, which are focused on road environments, we paid attention to sidewalks, where understanding the environment could provide the potential for improved walking of humans, especially impaired people. Concretely, we interviewed impaired people and carefully selected target objects from the interviewees' feedback (objects they encounter on sidewalks). We then acquired two different types of data: crowd-sourced data and stereo data. We labeled target objects at instance-level (i. e. , bounding box and polygon mask) and generated a ground-truth disparity map for the stereo data. SideGuide consists of 350K images with bounding box annotation, 100K images with a polygon mask, and 180K stereo pairs with the ground-truth disparity. We analyzed our dataset by performing baseline analysis for object detection, instance segmentation, and stereo matching tasks. In addition, we developed a prototype that recognizes the target objects and measures distances, which could potentially assist people with disabilities. The prototype suggests the possibility of practical application of our dataset in real life.

IROS Conference 2019 Conference Paper

Camera Exposure Control for Robust Robot Vision with Noise-Aware Image Quality Assessment

  • Ukcheol Shin
  • Jinsun Park
  • Gyumin Shim
  • François Rameau
  • In-So Kweon

In this paper, we propose a noise-aware exposure control algorithm for robust robot vision. Our method aims to capture best-exposed images, which can boost the performance of various computer vision and robotics tasks. For this purpose, we carefully design an image quality metric that captures complementary quality attributes and ensures light-weight computation. Specifically, our metric consists of a combination of image gradient, entropy, and noise metrics. The synergy of these measures allows the preservation of sharp edges and rich texture in the image while maintaining a low noise level. Using this novel metric, we propose a real-time and fully automatic exposure and gain control technique based on the Nelder-Mead method. To illustrate the effectiveness of our technique, a large set of experimental results demonstrates the higher qualitative and quantitative performance compared with conventional approaches.

ICRA Conference 2019 Conference Paper

Depth Completion with Deep Geometry and Context Guidance

  • Byeong-Uk Lee
  • Hae-Gon Jeon
  • Sunghoon Im 0001
  • In-So Kweon

In this paper, we present an end-to-end convolutional neural network (CNN) for depth completion. Our network consists of a geometry network and a context network. The geometry network, a single encoder-decoder network, learns to optimize a multi-task loss to generate an initial propagated depth map and a surface normal. The complementary outputs allow it to correctly propagate initial sparse depth points in slanted surfaces. The context network extracts a local and a global feature of an image to compute a bilateral weight, which enables it to preserve edges and fine details in the depth maps. At the end, a final output is produced by multiplying the initially propagated depth map with the bilateral weight. In order to validate the effectiveness and the robustness of our network, we performed extensive ablation studies and compared the results against state-of-the-art CNN-based depth completions, where we showed promising results on various scenes.

IROS Conference 2019 Conference Paper

DISC: A Large-scale Virtual Dataset for Simulating Disaster Scenarios

  • Hae-Gon Jeon
  • Sunghoon Im 0001
  • Byeong-Uk Lee
  • Dong-Geol Choi
  • Martial Hebert
  • In-So Kweon

In this paper, we present the first large-scale synthetic dataset for visual perception in disaster scenarios, and analyze state-of-the-art methods for multiple computer vision tasks with reference baselines. We simulated before and after disaster scenarios such as fire and building collapse for fifteen different locations in realistic virtual worlds. The dataset consists of more than 300K high-resolution stereo image pairs, all annotated with ground-truth data for semantic segmentation, depth, optical flow, surface normal estimation and camera pose estimation. To create realistic disaster scenes, we manually augmented the effects with 3D models using physical-based graphics tools. We use our dataset to train state-of-the-art methods and evaluate how well these methods can recognize the disaster situations and produce reliable results on virtual scenes as well as real-world images. The results obtained from each task are then used as inputs to the proposed visual odometry network for generating 3D maps of buildings on fire. Finally, we discuss challenges for future research.

IROS Conference 2019 Conference Paper

Fast Perception, Planning, and Execution for a Robotic Butler: Wheeled Humanoid M-Hubo

  • Moonyoung Lee
  • Yujin Heo
  • Jinyong Park
  • Hyundae Yang
  • Ho-Deok Jang
  • Philipp Benz
  • Hyunsub Park
  • In-So Kweon

As the aging population grows at a rapid rate, there is an ever growing need for service robot platforms that can provide daily assistance at practical speed with reliable performance. In order to assist with daily tasks such as fetching a beverage, a service robot must be able to perceive its environment and generate corresponding motion trajectories. This becomes a challenging and computationally complex problem when the environment is unknown and thus the path planner must sample numerous trajectories that often are sub-optimal, extending the execution time. To address this issue, we propose a unique strategy of integrating a 3D object detection pipeline with a kinematically optimal manipulation planner to significantly increase speed performance at run-time. In addition, we develop a new robotic butler system for a wheeled humanoid that is capable of fetching requested objects at 24% of the speed a human needs to fulfill the same task. The proposed system was evaluated and demonstrated in a real-world environment setup as well as in public exhibition.

IROS Conference 2019 Conference Paper

Learning Residual Flow as Dynamic Motion from Stereo Videos

  • Seokju Lee
  • Sunghoon Im 0001
  • Stephen Lin 0001
  • In-So Kweon

We present a method for decomposing the 3D scene flow observed from a moving stereo rig into stationary scene elements and dynamic object motion. Our unsupervised learning framework jointly reasons about the camera motion, optical flow, and 3D motion of moving objects. Three cooperating networks predict stereo matching, camera motion, and residual flow, which represents the flow component due to object motion and not from camera motion. Based on rigid projective geometry, the estimated stereo depth is used to guide the camera motion estimation, and the depth and camera motion are used to guide the residual flow estimation. We also explicitly estimate the 3D scene flow of dynamic objects based on the residual flow and scene depth. Experiments on the KITTI dataset demonstrate the effectiveness of our approach and show that our method outperforms other state-of-the-art algorithms on the optical flow and visual odometry tasks.

IROS Conference 2019 Conference Paper

Vehicular Multi-Camera Sensor System for Automated Visual Inspection of Electric Power Distribution Equipment

  • Jinsun Park
  • Ukcheol Shin
  • Gyumin Shim
  • Kyungdon Joo
  • François Rameau
  • Junhyeok Kim 0004
  • Dong-Geol Choi
  • In-So Kweon

In this paper, we present a multi-camera sensor system along with its control algorithm for automated visual inspection from a moving vehicle. To accomplish this task, we propose a unique hardware configuration consisting of a frontal stereo vision system, six lateral cameras motorized to tilt, and a GPS/IMU sensor mounted on the roof of a car. From the frontal stereo system, we detect electric poles and estimate their corresponding 3D positions. Based on this 3D estimation, the tilt angles of the motorized lateral cameras are controlled in real-time to capture high resolution images of the equipment - typically installed a few meters above the road surface. In addition, inertial odometry information from the GPS/IMU module is utilized for pose estimation, object localization, and re-identification among cameras. Experimental results demonstrate the efficiency and robustness of our system for automated electric equipment maintenance, which can reduce human effort significantly.

ICRA Conference 2017 Conference Paper

Deep representation of industrial components using simulated images

  • Seong-Heum Kim
  • Gyeongmin Choe
  • Byungtae Ahn
  • In-So Kweon

In this paper, we present a visual learning framework to retrieve a 3D model and estimate its pose from a single image. To increase the quantity and quality of training data, we define our simulation space in the near infrared (NIR) band, and utilize the quasi-Monte Carlo (MC) method for scalable photorealistic rendering of manufactured components. Two types of convolutional neural network (CNN) architectures are trained over these synthetic data and a relatively small amount of real data. The first CNN model seeks the most discriminative information and uses it to classify industrial components with fine-grained shape attributes. Once a 3D model is identified, one of the category-specific CNNs is tested for pose regression in the second phase. The mixed data for learning object categories is useful in domain adaptation and attention mechanism in our system. We validate our data-driven method with 88 component models, and the experimental results are qualitatively demonstrated. Also, the CNNs trained with various conditions of mixed data are quantitatively analyzed.

IROS Conference 2016 Conference Paper

EureCar turbo: A self-driving car that can handle adverse weather conditions

  • Unghui Lee
  • Jiwon Jung
  • Seunghak Shin
  • Yongseop Jeong
  • Kibaek Park
  • David Hyunchul Shim
  • In-So Kweon

Autonomous driving technology has made significant advances in recent years. In order for self-driving cars to become practical, they are required to operate safely and reliably even under adverse driving conditions. However, most current autonomous driving cars have only been shown to be operational under amiable weather conditions, i. e. , on sunny days on dry roads. In order to enable autonomous cars to handle adverse driving conditions such as rain and wet roads, the algorithm must be able to detect roads within a tolerable margin of error using sensors such as cameras and laser scanners. In this paper, we propose a sensor fusion algorithms that is able to operate under a variety of weather conditions, including rain. Our algorithm was validated when a strong shower occurred during the 2014 Hyundai Motor Company's Autonomous Car Competition. In this paper, we present the competition results that were collected on the same course on both sunny and rainy days. Based on the comparison, we propose the future directions to improve the autonomous driving capability under adverse environmental conditions.

IROS Conference 2016 Conference Paper

Object proposal using 3D point cloud for DRC-HUBO+

  • Seunghak Shin
  • Inwook Shim
  • Jiyung Jung
  • Yunsu Bok
  • Jun-Ho Oh
  • In-So Kweon

We present an object proposal method which utilizes the 3D data obtained from a depth sensor as well as the color information of images. Our object proposal method is designed to improve the performance of the object detection for a mobile robot equipped with a camera and a laser scanner. Compared to traditional object proposal methods using only 2D images, the proposed method provides much less number of candidate windows for object detection. We show less than 100 object proposal windows per image using the proposed method result in high recall tested on the public dataset. Our method presents object proposals in 3D space as well as in 2D image thus it can further be applied to following tasks for mobile robots such as 3D location and pose estimation of the target object after successful object detection. We validate our method using the real-world object detection dataset for outdoor mobile robots captured during the DRC Finals 2015 and the public dataset for comparison with the previous methods.

IROS Conference 2016 Conference Paper

Thermal Image Enhancement using Convolutional Neural Network

  • Yukyung Choi
  • Namil Kim
  • Soonmin Hwang
  • In-So Kweon

With the advent of commodity autonomous mobiles, it is becoming increasingly prevalent to recognize under extreme conditions such as night, erratic illumination conditions. This need has caused the approaches using multi-modal sensors, which could be complementary to each other. The choice for the thermal camera provides a rich source of temperature information, less affected by changing illumination or background clutters. However, existing thermal cameras have a relatively smaller resolution than RGB cameras that has trouble for fully utilizing the information in recognition tasks. To mitigate this, we aim to enhance the low-resolution thermal image according to the extensive analysis of existing approaches. To this end, we introduce Thermal Image Enhancement using Convolutional Neural Network (CNN), called in TEN, which directly learns an end-to-end mapping a single low resolution image to the desired high resolution image. In addition, we examine various image domains to find the best representative of the thermal enhancement. Overall, we propose the first thermal image enhancement method based on CNN guided on RGB data. We provide extensive experiments designed to evaluate the quality of image and the performance of several object recognition tasks such as pedestrian detection, visual odometry, and image registration.

ICRA Conference 2016 Conference Paper

Vision system and depth processing for DRC-HUBO+

  • Inwook Shim
  • Seunghak Shin
  • Yunsu Bok
  • Kyungdon Joo
  • Dong-Geol Choi
  • Joon-Young Lee
  • Jaesik Park
  • Jun-Ho Oh

This paper presents a vision system and a depth processing algorithm for DRC-HUBO+, the winner of the DRC finals 2015. Our system is designed to reliably capture 3D information of a scene and objects and to be robust to challenging environment conditions. We also propose a depth-map upsampling method that produces an outliers-free depth map by explicitly handling depth outliers. Our system is suitable for robotic applications in which a robot interacts with the real-world, requiring accurate object detection and pose estimation. We evaluate our depth processing algorithm in comparison with state-of-the-art algorithms on several synthetic and real-world datasets.

IROS Conference 2014 Conference Paper

2D-3D camera fusion for visual odometry in outdoor environments

  • Danda Pani Paudel
  • Cédric Demonceaux
  • Adlane Habed
  • Pascal Vasseur
  • In-So Kweon

Accurate estimation of camera motion is very important for many robotics applications involving SfM and visual SLAM. Such accuracy is attempted by refining the estimated motion through nonlinear optimization. As many modern robots are equipped with both 2D and 3D cameras, it is both highly desirable and challenging to exploit data acquired from both modalities to achieve a better localization. Existing refinement methods, such as Bundle adjustment and loop closing, may be employed only when precise 2D-to-3D correspondences across frames are available. In this paper, we propose a framework for robot localization that benefits from both 2D and 3D information without requiring such accurate correspondences to be established. This is carried out through a 2D-3D based initial motion estimation followed by a constrained nonlinear optimization for motion refinement. The initial motion estimation finds the best possible 2D-to-3D correspondences and localizes the cameras with respect the 3D scene. The refinement step minimizes the projection errors of 3D points while preserving the existing relationships between images. The problems of occlusion and that of missing scene parts are handled by comparing the image-based reconstruction and 3D sensor measurements. The effect of data inaccuracies is minimized using an M-estimator based technique. Our experiments have demonstrated that the proposed framework allows to obtain a good initial motion estimate and a significant improvement through refinement.

IROS Conference 2014 Conference Paper

Auto-adjusting camera exposure for outdoor robotics using gradient information

  • Inwook Shim
  • Joon-Young Lee
  • In-So Kweon

We present a new method to auto-adjust camera exposure for outdoor robotics. In outdoor environments, scene dynamic range may be wider than the dynamic range of the cameras due to sunlight and skylight. This can results in failures of vision-based algorithms because important image features are missing due to under-/over-saturation. To solve the problem, we adjust camera exposure to maximize image features in the gradient domain. By exploiting the gradient domain, our method naturally determines the proper exposure needed to capture important image features in a manner that is robust against illumination conditions. The proposed method is implemented using an off-the-shelf machine vision camera and is evaluated using outdoor robotics applications. Experimental results demonstrate the effectiveness of our method, which improves the performance of robot vision algorithms.

ICRA Conference 2014 Conference Paper

Extrinsic calibration of 2D laser sensors

  • Dong-Geol Choi
  • Yunsu Bok
  • Jun-Sik Kim 0001
  • In-So Kweon

This paper describes a new methodology for estimating a relative pose of two 2D laser sensors. Two dimensional laser scan points do not have enough feature information for motion tracking. For this reason, additional image sensors or artificial landmarks have been used to find a relative pose. We propose the method to estimate a relative pose of 2D laser sensors without any additional sensor or artificial landmark. By scanning two orthogonal planes, we utilize only the coplanarity of the scan points on each plane and the orthogonality of the plane normals. Experiments with both synthetic and real data show the validity of the proposed method. To the best of our knowledge this works provides the first solution for the problem.

IROS Conference 2014 Conference Paper

Extrinsic calibration of non-overlapping camera-laser system using structured environment

  • Yunsu Bok
  • Dong-Geol Choi
  • Pascal Vasseur
  • In-So Kweon

In this paper are presented simple and practical solutions to extrinsic calibration between a camera and a 2D laser sensor, without overlap. Previous methods utilized a plane or an intersecting line of two planes as a geometric constraint with enough common field-of-view. These required additional sensors to calibrate non-overlapping systems. In this paper, we present two methods for solving the problem - one utilizes a plane; the other utilizes an intersecting line of two planes. For each method, an initial solution of the relative positions of a non-overlapping camera and a laser sensor, was computed by adopting a reasonable assumption about geometric structures. Then we refined it via non-linear optimization, even if the assumption was not perfectly satisfied. Both simulation results and experiments using real data showed that the proposed methods provided reliable results compared to ground-truth, and similar or better results than those provided by a conventional method.

ICRA Conference 2014 Conference Paper

Hybrid vision-based SLAM coupled with moving object tracking

  • Jihong Min
  • Jungho Kim 0005
  • Hyeongwoo Kim
  • Kiho Kwak
  • In-So Kweon

In this paper we propose a hybrid vision-based SLAM and moving objects tracking (vSLAMMOT) approach. This approach tightly combines two key methods: a superpixel-based segmentation to detect moving objects and a Rao-Blackwellized Particle Filter to estimate a stereo-vision-based SLAM posterior. Most successful methods perform vision-based SLAM (vSLAM) and track moving objects independently. However, we pose both vSLAM and moving object tracking as a single correlated problem to leverage the performance. Our approach estimates the relative camera motion using the previous tracking result, and then detects moving objects from the estimated camera motion recursively. Moving superpixels are detected by a Markov Random Field (MRF) model which uses spatial and temporal information of the moving objects. We demonstrate the performance of the proposed approach for vSLAMMOT using both synthetic and real datasets and compare the performance with other methods.

ICRA Conference 2013 Conference Paper

Generalized laser three-point algorithm for motion estimation of camera-laser fusion system

  • Yunsu Bok
  • Dong-Geol Choi
  • In-So Kweon

This paper presents a new structure-from-motion (SFM) technique called `generalized laser three-point' algorithm. It is designed to estimate the motion of the camera-laser fusion system which consists of a 2D laser sensor and multiple cameras. The laser points are projected onto the images and tracked to other frames to be used as 3D-2D correspondences. However, the typical three-point algorithms cannot estimate the motion of the system if three points are collinear. Using the laser points as 3D points, this case happens frequently if the laser sensor scans a large plane (e. g. open ground). Even in that case, two frames of the laser data are not collinear if the system is moved while it captures the frames. Among three point correspondences required to estimate the motion, we select two points and the other point from different frames. We estimate the relative pose between the frames by solving an 8-degree polynomial equation. The experimental results show that the proposed algorithm is more appropriate for our fusion system than the previous algorithms.

IROS Conference 2012 Conference Paper

Autonomous homing based on laser-camera fusion system

  • Dong-Geol Choi
  • Inwook Shim
  • Yunsu Bok
  • Tae Hyun Oh
  • In-So Kweon

Building maps of unknown environments is a critical factor for autonomous navigation and homing, and this problem is especially challenging in large-scale environments. Recently, sensor fusion systems such as combinations of cameras and laser sensors have become popular in the effort to ensure a general level of performance in this task. In this paper, we present a new homing method in a large-scale environment using a laser-camera fusion system. Instead of fusing data to form a single map builder, we adaptively select sensor data to handle environments which contain ambiguity. For autonomous homing, we propose a new mapping strategy for building a hybrid map and a return strategy for selecting the next target waypoints efficiently. The experimental results demonstrate that the proposed algorithm enables the autonomous homing of a robot in a large-scale indoor environments in real time.

ICRA Conference 2012 Conference Paper

Efficient Data-Driven MCMC sampling for vision-based 6D SLAM

  • Jihong Min
  • Jungho Kim 0005
  • Seunghak Shin
  • In-So Kweon

In this paper, we propose a Markov Chain Monte Carlo (MCMC) sampling method with the data-driven proposal distribution for six-degree-of-freedom (6-DoF) SLAM. Recently, visual odometry priors have been widely used as the process model in the SLAM formulation to improve the SLAM performance. However, modeling the uncertainties of incremental motions estimated by visual odometry is especially difficult under challenging conditions, such as erratic motion. For a particle-based model representation, it can represent the uncertainty of the camera motion well under erratic motion compared to the constant velocity model or a Gaussian noise model, but the manner of representing the proposal distribution and sampling the particles is extremely important, as we can maintain only a limited number of particles in the high-dimensional state space. Hence, we propose an effective sampling approach by exploiting MCMC sampling and the data-driven proposal distribution to propagate the particles. We demonstrate the performance of the proposed approach for 6-DoF SLAM using both synthetic and real datasets and compare the performance with those of other sampling methods.

IROS Conference 2011 Conference Paper

A novel 2. 5D pattern for extrinsic calibration of ToF and camera fusion system

  • Jiyung Jung
  • Yekeun Jeong
  • Jaesik Park
  • Hyowon Ha
  • James Dokyoon Kim
  • In-So Kweon

Recently, many researchers have made efforts for accurate calibration of a Time-of-Flight camera to fully utilize its provided depth values. Yet most previous works focus mainly on intrinsic calibration by modeling its systematic errors and noises while extrinsic calibration is also an important factor when constructing sensor fusion system. In this paper, we present a calibration process that can correctly transfer the depth measurements onto the color image. We use 2. 5D pattern so that sufficient reprojection error can be considered for both color and ToF cameras. The issues on obtaining the correct correspondences for this pattern are discussed. In the optimization stage, the depth constraint is also employed to ensure the depth measurements to lie on the pattern plane. The strengths of the proposed method over previous approaches are evaluated in several robotic applications which require precise ToF and camera calibration.

IROS Conference 2011 Conference Paper

Capturing city-level scenes with a synchronized camera-laser fusion sensor

  • Yunsu Bok
  • Dong-Geol Choi
  • Yekeun Jeong
  • In-So Kweon

In this paper, we present a sensor fusion system of cameras and 2D laser sensors for 3D reconstruction. The proposed system is designed to capture data on a fast-moving ground vehicle. The system consists of six cameras and one 2D laser sensor. In order to capture data at high speed, we synchronized all sensors by detecting the laser ray at a specific angle and generating a trigger signal for the cameras. Reconstruction of 3D structures is done by estimating frame-by-frame motion and accumulating vertical laser scans. The difference between the proposed system and the previous works using two 2D laser sensors is that we do not assume 2D motion. The motion of the system in 3D space (including absolute scale) is estimated accurately by data-level fusion of images and range data. The problem of error accumulation is solved by loop closing, not by GPS. The moving objects are detected by utilizing the depth information provided by the laser sensor. The experimental results show that the estimated path is successfully overlayed on the satellite images.

ICRA Conference 2011 Conference Paper

Complementation of cameras and lasers for accurate 6D SLAM: From correspondences to bundle adjustment

  • Yekeun Jeong
  • Yunsu Bok
  • Jun-Sik Kim 0001
  • In-So Kweon

In this paper, we present an accurate and robust 6D SLAM method that uses multiple 2D sensors, i. e. perspective cameras and planar laser scanners. We have investigated strengths and weaknesses of those two sensors for 6D SLAM by conducting specifically designed experiments, and found that the sensors can complement each other. In order to take full advantages of each approach, we fuse correspondences of those two sensors, rather than individually estimated motions. Correspondences obtained by the two sensors have different characteristics, but can be expressed in a common 2D-3D relation form. We use the correspondences in a single structure-from-motion framework. In the initial motion estimation step, we propose a RANSAC-based method to generate and test multiple motion hypotheses by using multiple pools of correspondences, aiming to avoid potential bias of each sensor data. In the later motion refinement step, we introduce a variant of bundle adjustment to consider different types of constraints from the two sensors. The performance of the proposed method is demonstrated both quantitatively by experiments on closed-loop sequences and qualitatively by large-scale experiments with DGPS trajectory. The proposed method successfully closes a loop of 320 meters in twenty thousand frames by incremental process only.

IROS Conference 2010 Conference Paper

An original approach for automatic plane extraction by omnidirectional vision

  • Jean-Charles Bazin
  • Pierre-Yves Laffont
  • In-So Kweon
  • Cédric Demonceaux
  • Pascal Vasseur

Whereas some methods for plane extraction have been proposed, this problem still remains an open issue due to the complexity of the task. This paper especially focuses on the extraction of points lying on a plane (such as the ground and buildings walls) in sequences acquired by a central omnidirectional camera. Our approach is based on the epipolar constraint for planar scenes (i. e. homography) on a pair of omnidirectional images to detect some interest points belonging to a plane. Our main contribution is the introduction of a new method, called “2-point algorithm for homography”, that imposes some constraints on the homography using vanishing point (VP) information. Compared to the widely used DLT (4-point) algorithm, experiments on real data demonstrated that the proposed “2-point algorithm for homography” is more robust to noise and false matching, even when the plane to extract is not dominant in the image. Finally, we show that our system provides key clues for ground segmentation by GrabCut.

IROS Conference 2010 Conference Paper

Robust visual lock-on and simultaneous localization for an unmanned aerial vehicle

  • Jihong Min
  • Yekeun Jeong
  • In-So Kweon

We present a method for simultaneously locking on to a ground target and estimating the position of an unmanned aerial vehicle (UAV) under countermeasure (CM) conditions, where sensors are prevented from successfully tracking a target. Owing to the limited payload and power of the UAVs, we employ a monocular camera and a global positioning system (GPS) to carry out vision-based simultaneous localization and mapping (SLAM) using both an unscented Kalman filter and a Kalman filter. Since this approach estimates the state of the UAV and the location of the target, we can estimate the position of the target in the image, even in the presence of CMs. Our experiments show that the proposed method successfully locks on to the target and estimates the state of the UAV.

ICRA Conference 2010 Conference Paper

Vision-based navigation with pose recovery under visual occlusion and kidnapping

  • Jungho Kim 0005
  • In-So Kweon

Vision-based robotic applications such as Simultaneous Localization and Mapping (SLAM), global localization, and autonomous navigation have suffered from problems related to dynamic environments involving moving objects and kidnapping. One of the possible solutions to these problems is to establish robust correspondences when obtaining images from static scenes. Therefore we propose an efficient technique for determining correspondences to recover the current camera pose; in the proposed method, the FAST corner detector and SIFT descriptors are combined because in many methods for vision-based robotic applications, corner features have been adopted since they enable fast computation and simplify the computation of the correspondences between consecutive images. However, to recover the pose of the camera after kidnapping or at an unknown initial position, a robust feature matching algorithm is required because the pose of a camera is unlikely to be the same as the poses in the database images. For this purpose, first, we determine some candidates for correspondences by combining corners with their multiple descriptors computed from previously defined scales, and then we select one of these candidates by optimizing the scale using a variant of the mean-shift algorithm. We apply the proposed matching algorithm to kidnapping and visual occlusion problems in autonomous navigation.

ICRA Conference 2010 Conference Paper

Visual tracking for non-rigid objects using Rao-Blackwellized particle filter

  • Jungho Kim 0005
  • Chaehoon Park
  • In-So Kweon

Particle filters have been used for visual tracking during long periods because they enable effective estimation for non-linear and non-Gaussian distributions. However, particle filter-based tracking approaches suffer from occlusion and deformation of the target objects, which result in the large difference between the current observations and the target model. Thus, we present a Rao-Blackwellized particle filter (RBPF)-based tracking algorithm that effectively estimates the joint distribution for the target state and the target model; in the proposed method, the target object is tracked by using the particle filter while the target model is simultaneously updated on the basis of the on-line approximation of a mixture of Gaussians. To ensure the robustness to occlusion, we represent the target model by 16 orientation histograms that are spatially divided, and individually update each histogram through a video sequence. We demonstrate the robustness of the proposed method under occlusion and deformation of the target objects.

ICRA Conference 2009 Conference Paper

Dynamic programming and skyline extraction in catadioptric infrared images

  • Jean-Charles Bazin
  • In-So Kweon
  • Cédric Demonceaux
  • Pascal Vasseur

Unmanned Aerial Vehicles (UAV) are the subject of an increasing interest in many applications and a key requirement for autonomous navigation is the attitude/position stabilization of the vehicle. Some previous works have suggested using catadioptric vision, instead of traditional perspective cameras, in order to gather much more information from the environment and therefore improve the robustness of the UAV attitude/position estimation. This paper belongs to a series of recent publications of our research group concerning catadioptric vision for UAVs. Currently, we focus on the extraction of skyline in catadioptric images since it provides important information about the attitude/position of the UAV. For example, the DEM-based methods can match the extracted skyline with a Digital Elevation Map (DEM) by process of registration, which permits to estimate the attitude and the position of the camera. Like any standard cameras, catadioptric systems cannot work in low luminosity situations because they are based on visible light. To overcome this important limitation, in this paper, we propose using a catadioptric infrared camera and extending one of our methods of skyline detection towards catadioptric infrared images. The task of extracting the best skyline in images is usually converted in an energy minimization problem that can be solved by dynamic programming. The major contribution of this paper is the extension of dynamic programming for catadioptric images using an adapted neighborhood and an appropriate scanning direction. Finally, we present some experimental results to demonstrate the validity of our approach.

ICRA Conference 2009 Conference Paper

Graph-based robust shape matching for robotic application

  • Hanbyul Joo
  • Yekeun Jeong
  • Olivier Duchenne
  • Seong Young Ko
  • In-So Kweon

Shape is one of the useful information for object detection. The human visual system can often recognize objects based on the 2-D outline shape alone. In this paper, we address the challenging problem of shape matching in the presence of complex background clutter and occlusion. To this end, we propose a graph-based approach for shape matching. Unlike prior methods which measure the shape similarity without considering the relation among edge pixels, our approach uses the connectivity of edge pixels by generating a graph. A group of connected edge pixels, which is represented by an “edge” of the graph, is considered together and their similarity cost is defined for the “edge” weight by explicit comparison with the corresponding template part. This approach provides the key advantage of reducing ambiguity even in the presence of background clutter and occlusion. The optimization is performed by means of a graph-based dynamic algorithm. The robustness of our method is demonstrated for several examples including long video sequences. Finally, we applied our algorithm to our grasping robot system by providing the object information in the form of prompt hand-drawn templates.

IROS Conference 2009 Conference Paper

UAV global pose estimation by matching forward-looking aerial images with satellite images

  • Kilho Son
  • Youngbae Hwang
  • In-So Kweon

A global pose estimation method of an Unmanned Aerial Vehicle (UAV) by matching forward-looking aerial images from the UAV flying at low altitude with down-looking images from a satellite is proposed. To overcome the limitation of significantly different camera viewpoints and characteristics, we use buildings as a cue of matching. We extract buildings from aerial images and construct a 3D model of buildings, using the fundamental matrix. We estimate the global pose of the vehicle by matching 3D structure of buildings with satellite images, using a particle filter. Experimental results show that the proposed approach is a promising method to the global pose estimation of the UAV with forward-looking vision data.

IROS Conference 2008 Conference Paper

A robust top-down approach for rotation estimation and vanishing points extraction by catadioptric vision in urban environment

  • Jean-Charles Bazin
  • In-So Kweon
  • Cédric Demonceaux
  • Pascal Vasseur

A key requirement for unmanned aerial vehicles (UAV) applications is the attitude stabilization of the aircraft, which requires the knowledge of its orientation. It is now well established that traditional navigation equipments, like GPS or INS, suffer from several disadvantages. That is why some works have suggested a vision-based approach of the problem. Especially, catadioptric vision is more and more used since it permits to gather much more information from the environment, compared to traditional perspective cameras, and therefore the robustness of the UAV attitude estimation is improved. Rotation estimation from conventional and catadioptric images has been extensively studied. Whereas interesting results can be obtained, the existing methods have non-negligible limitations such as difficult features matching (e. g. repeated texture, blurring or illumination changing) or a high computational cost (e. g. vanishing point extraction or analyze in frequency domain). In order to overcome these limitations, this paper presents a top-down approach for estimating the rotation and extracting the vanishing points in catadioptric images. This new framework is accurate and can run in real-time. To obtain the ground truth data, we also calibrate our catadioptric camera with a gyroscope. Finally, experimental results on a real video sequence are presented and compared to the ground truth data obtained by the gyroscope.

IROS Conference 2008 Conference Paper

Automatic calibration of catadioptric cameras in urban environment

  • Jean-Charles Bazin
  • In-So Kweon
  • Cédric Demonceaux
  • Pascal Vasseur

Camera calibration is an important step for vision-based stabilization of unmanned aerial vehicles (UAV). The goal of this paper is to develop a method for automatic calibration of a catadioptric camera so that it can be easily run before mounting the camera on the UAV or even during the flight to deal with vibrations or shocks. Whereas existing works can provide interesting results, they suffer from several practical limitations (manual line extraction, inaccurate conic fitting, calibration pattern, camera motion, execution time, etc. ..) and therefore cannot be applied in our application. The proposed algorithm aims to determine the most probable calibration that verifies some geometric constraints induced by catadioptric projection. In order to efficiently maximize this probability, we use a particle filtering approach. Experimental results have demonstrated the effectiveness of the proposed method.

IROS Conference 2008 Conference Paper

Efficient color feature extraction and matching for motion estimation and mapping

  • Hyoseok Hwang
  • In-So Kweon

Feature extraction and matching is one of the most significant research areas in robot vision. In this paper, we present a new method for motion estimation and mapping using color feature extraction and matching. The proposed method reduces computational cost and has good performance. The experimental result shows that the proposed method not only runs faster but provides accurate result.

IROS Conference 2008 Conference Paper

Efficient feature tracking for scene recognition using angular and scale constraints

  • Jungho Kim 0005
  • Ouk Choi
  • In-So Kweon

Recently, many vision-based robotic applications such as visual SLAM (Simultaneous Localization And Mapping) and autonomous navigation have achieved good performance using visual features. In these applications, robust feature tracking plays an important role, e. g. , in scene recognition for autonomous navigation and in data association for visual SLAM. In this paper, we propose a hierarchical outlier detection algorithm for robust feature tracking; the algorithm uses a simple window-based correlation (NCC) and enforces angular and scale constraints. The proposed algorithm maximizes the inter-cluster score and detects outliers that do not satisfy the angular constraints. The remaining outliers are detected by enforcing scale constraints using SIFT descriptors. The proposed algorithm is efficient and particularly useful for scene recognition, in which an image corresponding to a query image is searched among data images. Experimental results demonstrate that the proposed algorithm is robust to outliers and image variations such as scale changes. One of the main applications of the proposed algorithm is global localization due to its low computational complexity and robustness to outliers.

IROS Conference 2008 Conference Paper

Robust vision-based autonomous navigation against environment changes

  • Jungho Kim 0005
  • Yunsu Bok
  • In-So Kweon

Recently, many vision-based navigation methods have been introduced as an intelligent robot application. However, many of these methods mainly focus on finding an image in the database corresponding to a query image. Thus, if the environment changes, for example, objects moving in the environment, a robot is unlikely to find consistent corresponding points with one of the database images. To handle these problems, we propose a novel motion-based navigation method in contrast with appearance-based approaches. This algorithm is based on motion estimation by a camera to plan the next movement of a robot and robust feature matching to recognize home and destination locations. Experimental results demonstrate the capability of the vision-based autonomous navigation against environment changes.

ICRA Conference 2008 Conference Paper

UAV Attitude estimation by vanishing points in catadioptric images

  • Jean-Charles Bazin
  • In-So Kweon
  • Cédric Demonceaux
  • Pascal Vasseur

Unmanned aerial vehicles (UAV) are the subject of an increasing interest in many applications and a key requirement is the stabilization of the vehicle. Some previous works have suggested using catadioptric vision, instead of traditional perspective cameras, in order to gather much more information from the environment and therefore improve the robustness of the UAV attitude estimation. This paper belongs to a series of recent publications of our research group concerning catadioptric vision for UAVs. Currently, we focus on the estimation of the complete attitude of a UAV flying in urban environment. In order to avoid the limitations of horizon-based approaches, the difficulties of traditional epipolar methods (such as rotation-translation ambiguity, lack of features, retrieving motion parameters from matrix decomposition, etc. .) and improve UAV dynamic control, we suggest computing infinite homography. We show how catadioptric vision plays a key role to: first, extract a large number of lines, second robustly estimate the associated vanishing points and third, track them even during long video sequences. Therefore it is not only possible to estimate the relative rotation between consecutive frames but also compute the absolute rotation between two distant frames without error accumulation. Finally, we present some experimental results with ground truth data to demonstrate the accuracy and the robustness of our method.

ICRA Conference 2007 Conference Paper

Accurate Motion Estimation and High-Precision 3D Reconstruction by Sensor Fusion

  • Yunsu Bok
  • Youngbae Hwang
  • In-So Kweon

The CCD camera and the 2D laser range finder are widely used for motion estimation and 3D reconstruction. With their own strengths and weaknesses, low-level fusion of these two sensors complements each other. We combine these two sensors to perform motion estimation and 3D reconstruction simultaneously and precisely. We develop a motion estimation scheme appropriate for this sensor system. In the proposed method, the motion between two frames is estimated using three points among the scan data, and refined by nonlinear optimization. We validate the accuracy of the proposed method using real images. The results show that the proposed system is a practical solution for motion estimation as well as for 3D reconstruction.

IROS Conference 2007 Conference Paper

Robust feature matching for loop closing and localization

  • Jungho Kim 0005
  • In-So Kweon

Recently, many vision-based SLAM methods have achieved good results using visual features. However, most algorithms suffer from the accumulated error that inevitably occurs. In this paper, we propose a robust loop detection method by matching image features between the incoming image and key-frame images saved in SLAM. Loop detection is a task of deciding whether a robot has returned to a previously visited area or not. Because a camera is unlikely to have the same pose when a robot revisits the place where it previously encountered, it is crucial to match the features under the different views of the scene. In contrast with view-invariant features, it is hard to match corner points in that situation due to the large variation of neighboring pixels. So we present the robust corner matching method under the view changes. Experimental results demonstrate the capability of the loop closing and mobile robot localization under the different views using the proposed method.

ICRA Conference 2007 Conference Paper

Visual Categorization Robust to Large Intra-Class Variations using Entropy-guided Codebook

  • Sungho Kim 0003
  • In-So Kweon
  • Chil-Woo Lee

Categorizing visual elements is fundamentally important for autonomous mobile robots to get intelligence such as new object acquisition and topological place classification. The main problem of visual categorization is how to reduce the large intra-class variations, especially surface markings of man-made objects. In this paper, we present a robust method by introducing intermediate blurring and entropy-guided codebook selection in a bag-of-words framework. Intermediate blurring can filter out the high frequency of surface markings and provide dominant shape information. Entropy of a hypothesized codebook can provide the necessary measure for the semantic parts among training exemplars. From the first step, a generative optimal codebook for each category is learned using the MDL (minimum description length) principle guided by entropy information. From the second step, a final set of codebook is learned using the discriminative method guided by the inter-category entropy of the codebook. We select the necessary parameters through various evaluations and validate the effect of the surface marking reduction method using a Caltech-101 DB, which has large intra-class variations. Finally, we briefly introduce the impact of the method to the object categorization and segmentation problem.

IROS Conference 2005 Conference Paper

Metric localization using a single artificial landmark for indoor mobile robots

  • Gi-jeong Jang
  • Sungho Kim 0003
  • Jeongho Kim
  • In-So Kweon

We present an accurate metric localization method using a simple artificial landmark for the navigation of indoor mobile robots. The proposed landmark model is designed to have a three-dimensional, multi-colored structure and the projective distortion of the structure encodes the distance and heading of the robot with respect to the landmark. Catadioptric vision is adopted for the robust and easier acquisition of the bearing measurements for the landmark. We propose a practical EKF based self-localization method that uses a single artificial landmark and runs in real time.

IROS Conference 2005 Conference Paper

Recognition-based indoor topological navigation using robust invariant features

  • Zhe Lin 0001
  • Sungho Kim 0003
  • In-So Kweon

In this paper, we present a recognition-based autonomous navigation system for mobile robots. The system is based on our previously proposed robust invariant feature (RIF) detector. This detector extracts highly robust and repeatable features based on the key idea of tracking multi-scale interest points and selecting unique representative local structures with the strongest response in both spatial and scale domains. Weighted Zernike moments are used as the feature descriptor and applied to the place recognition. The navigation system is composed of on-line and off-line two stages. In the off-line learning stage, we train the robot in its workspace by just taking several images of representative places as landmarks. Then, in the on-line navigation stage, the robot recognizes scenes, obtains robust feature correspondences, and navigates the environment autonomously using the iterative pose converging (IPC) algorithm which is based on the idea of the visual servoing technique. The experimental results and the performance evaluation show that the proposed navigation system can achieve excellent performance in complex indoor environments.

IROS Conference 2004 Conference Paper

Change detection using a statistical model of the noise in color images

  • Youngbae Hwang
  • Jun-Sik Kim 0001
  • In-So Kweon

We present a novel change detection method using a statistical model of the image noise. Most change detection methods are based on gray-level images. However, color images can provide much richer scene information. One major problem to use the color images in change detection is how to combine three components in color space as a detection cue. We use the Euclidean color distance of three channels to measure the difference between two consecutive images. Specifically, we present a new noise model for each color channel. Through this modeling we can estimate the distribution of the Euclidean color distance for unchanged regions. We can find the optimal threshold to detect changes using this estimated distribution. Although we use the optimal threshold, inevitably there may be false classifications. To reject these erroneous cases, we adopt the graph cuts method that efficiently minimizes the global energy, which takes into account the effect of neighboring pixels.

IROS Conference 2003 Conference Paper

Automatic edge detection method for the mobile robot application

  • Wang-Heon Lee
  • Dongsu Kim
  • In-So Kweon

This paper proposes a new edge detection method using a 3/spl times/3 ideal binary pattern and lookup table (LUT) for the mobile robot localization without any parameter adjustments. We take the mean of the pixels within the 3/spl times/3 block as a threshold by which the pixels are divided into two groups. The edge magnitude and orientation are calculated by taking the difference of average intensities of the two groups and by searching directional code in the LUT, respectively. And also the input image is not only partitioned into multiple groups according to their intensity similarities by the histogram, but also the threshold of each group is determined by fuzzy reasoning automatically. Finally, the edges are determined through non-maximum suppression using edge confidence measure and edge linking. Applying this edge detection method to the mobile robot localization using projective invariance of the cross ratio, we demonstrate the robustness of the proposed method to the illumination changes in a corridor environment.

ICRA Conference 2003 Conference Paper

Robust model-based 3D object recognition by combining feature matching with tracking

  • Sungho Kim 0003
  • In-So Kweon
  • Incheol Kim

We propose a vision based 3D object recognition and tracking system, which provides high level scene descriptions such as object identification and 3D pose information. The system is composed of object recognition part and real-time tracking part. In object recognition, we propose a feature which is robust to scale, rotation, illumination change and background clutter. A probabilistic voting scheme maximizes the conditional probability defined by the features in correspondence to recognize an object of interest. As a result of object recognition, we obtain the homography between the model image and the input scene. In tracking, a Lie group formalism is used to cast the motion computation problem into simple geometric terms so that tracking becomes a simple optimization problem. An initial object pose is estimated using correspondences between the model image and the 3D CAD model which are predefined and the homography which relates the model image to the input scene. Results from the experiments show the robustness of the proposed system.

ICRA Conference 2002 Conference Paper

Color Landmark Based Self-Localization for Indoor Mobile Robots

  • Gi-jeong Jang
  • Sungho Kim 0003
  • Wang-Heon Lee
  • In-So Kweon

We present a simple artificial landmark model and a robust tracking algorithm for the navigation of indoor mobile robots. The landmark model is designed to have a three-dimensional structure consisting of a multi-colored planar pattern. A stochastic algorithm based on Condensation [1] tracks the landmark model robustly using the color distribution of the pattern. A new self-localization algorithm computes the location of robot with the tracked single landmark. Experimental results show that the proposed landmark model is eflective. Through extensive navigation experiments in a cluttered indoor environment, we demonstrate the feasibility of the single view based self-localization in real-time.

IROS Conference 2001 Conference Paper

A new camera calibration method for robotic applications

  • Jun-Sik Kim 0001
  • In-So Kweon

In this paper, we present a camera calibration method using two views of three pairs of concentric circles with known sizes. A new invariant property for circle is introduced to determine the position of the center of the projected circle. Given two image ellipses and their corresponding centers, a cross-ratio based method estimates the center of projected circles using the new invariant property. The accurate center of projected concentric circles provides correct correspondences between ellipses in the image plane and circles in 3D plane. We also demonstrate that two views of three pairs of coplanar concentric circles are enough to determine the intrinsic camera parameters, such as the focal length, the aspect ratio, and the principal point. We validate the performance of the method using both synthetic and real images. Our method shows a comparable performance with respect to similar calibration methods using planes. The use of concentric circles, however, provides correct correspondences between 3D target points and their image points, and greatly simplifies the calibration problem.

IROS Conference 2001 Conference Paper

Artificial landmark tracking based on the color histogram

  • Kuk-Jin Yoon
  • In-So Kweon

For the fast and accurate self-localization of mobile robots, landmarks can be used very efficiently in the complex workspace. In this paper, we propose a simple color landmark model for self-localization and a fast landmark detection and tracking algorithm based on the proposed landmark model. We develop a color landmark with symmetric and repetitive structures, which shows invariant color histogram characteristics under some geometric distortions. Detection and tracking of the model are accomplished by a factored sampling technique in which color similarity is estimated by the color histogram intersection. We also use the color similarity to update the color histogram model of the landmark model for robust tracking under illumination change. We demonstrate the feasibility of the proposed technique through experiments in cluttered indoor environments.

ICRA Conference 2001 Conference Paper

Robust Object Tracking Using an Adaptive Color Model

  • Gi-jeong Jang
  • In-So Kweon

In this paper we present a new robust face tracking method based on the condensation algorithm that uses a sampling based density representation. A two-dimensional color model is used to approximate the face color. We modified the condensation algorithm to provide color adaptability to the abrupt change of illumination and to the tracking of differently colored people. According to the face size and location uncertainty, the searching range is automatically determined and it makes the algorithm extremely robust and efficient. The tracker operates at real-time and actively controls a camera pan-tilt in order to locate a person's face in the center of the image. Experimental results show the algorithm's robustness to the agile motion of face and to the dramatic change of illumination in the presence of complex background.

IROS Conference 2000 Conference Paper

A novel image-based control-law for the visual servoing system under large pose error

  • HoWon Kim 0002
  • JaeSeung Cho
  • In-So Kweon

In this paper, we analyze and solve the problem of image-based visual servoing system under large initial pose error. In this case we must consider camera field of view and the violation of the assumption of image-based visual servoing. Control-law is derived from linear approximation of the image error between the gripper and the target object. To solve this problem, we propose a unified image-based control strategy with a new image space named as virtual image plane (VIP). And next we propose an image-based control-law by decoupling the translation and rotation motion of the large pose error The remarkable property of our method is that we only use the combination of the image features in the 2D image space. Finally we demonstrate the feasibility of the proposed system through experiments using a six-DOF robot equipped with a stereo camera.

ICRA Conference 2000 Conference Paper

Self-Calibration Using the Linear Projective Reconstruction

  • Jong-Eun Ha
  • Jin-Young Yang
  • Kuk-Jin Yoon
  • In-So Kweon

Self-calibration algorithms that use only the information in the image have been actively researched. However, most algorithms require bundle adjustment in the projective reconstruction or in the nonlinear minimization. We propose a practical self-calibration algorithm that only requires a linear projective reconstruction. We overcome the sensitivity of the algorithm due to image noises by adding another constraint on the principal point. Also, we propose a variant of linear auto-calibration algorithm which uses the similar assumption of the work of Pollefeys et al. (1998), based on the property of the absolute quadric. Experimental results using real and synthetic images demonstrate the feasibility of the proposed algorithm.

IROS Conference 1999 Conference Paper

Calibration and 3D structure recovery under varying cameras using known angles

  • Jong-Eun Ha
  • In-So Kweon

We present an algorithm for the calibration of a camera and the recovery of 3D scene structure up to a scale from image sequences using known angles between lines in the scene. The proposed method computes the intrinsic parameters of the camera using the invariance of angles under the similarity transformation. Specifically, we recover the matrix that is the homography between the projective structure and the Euclidean structure using angles. Since this matrix is a unique one in the given set of image sequences, we can easily deal with the problem of varying intrinsic parameters of the camera. Experimental results on the synthetic and real images demonstrate the feasibility of the proposed algorithm.

ICRA Conference 1998 Conference Paper

3-D Object Recognition Using Projective invariant Relationship by Single-View

  • Kyoung-Sig Roh
  • Bum-Jae You
  • In-So Kweon

We propose a new method for recognizing three-dimensional objects using a three-dimensional invariant relationship for a special structure and geometric hashing by single-view. We use a special structure consisting of four co-planar points and any two non-coplanar points with respect to the plane. We derive an invariant relationship for the structure, which is represented by a plane equation. For recognition of 3-D objects using geometric hashing, a set of points on the plane is mapped into a set of points intersecting the plane and the unit sphere, thereby satisfying the invariant relationship. Experiments using 3-D polyhedral objects are carried out to demonstrate the feasibility of our method for 3-D object recognition.

IROS Conference 1998 Conference Paper

Fast object recognition using salient line groups

  • Dong Jung Kang
  • In-So Kweon

This paper presents an effective recognition method based on perceptual organization of low level features detected in an image. The method uses a dynamic programming (DP) based formulation to represent various line groups such as convex, concave, and more complex patterns consisting of convex and concave shapes. The essential features of perceptual organization such as endpoint proximity, collinearity, parallelism, and connectivity of lines, are incorporated into the DP based formulation as energy terms. As endpoint proximity, we detect two line junctions from image lines. We then search for junction groups by using collinearity constraint between the junctions. A DP-based search algorithm is used to detect a junction chain similar to the model chain, based on a local comparison. The proposed system is able to find line groups from images with broken lines and strong background clutters. We demonstrate the feasibility of our DP-based matching method based on perceptual organization using real images.

IROS Conference 1997 Conference Paper

Obstacle detection and self-localization without camera calibration using projective invariants

  • Kyoung-Sig Roh
  • Wang-Heon Lee
  • In-So Kweon

In this paper, we propose two new vision-based methods for indoor mobile robot navigation. One is a self-localization algorithm using projective invariant and the other is a method for obstacle detection by simple image difference and relative positioning. For a geometric model of corridor environment, we use natural features formed by floor, walls, and door frames. Using the cross-ratios of the features can be effective and robust in building and updating model-base, and image matching. We predefine a risk zone without obstacles for a robot, and store the image of the risk zone, which will be used to detect obstacles inside the zone by comparing the stored image with the current image of a new risk zone. The position of the robot and obstacles are determined by relative positioning. The robustness and feasibility of our algorithms have been demonstrated through experiments in corridor environments using the KASIRI-II indoor mobile robot.

IROS Conference 1997 Conference Paper

Vehicle segmentation using evidential reasoning

  • Joon Woong Lee
  • In-So Kweon

This paper proposes a segmentation algorithm by means of an evidential reasoning to segment moving vehicles in front of our moving car in a road traffic scene. Generally, an evidential reasoning finds the perceptually known evidences of a target and updates a probabilistic expectation for the target to be in an image. Since a noise image produces unreliable features and degrades the detection and localization, selecting image primitives which are less sensitive to noise and well represent the evidences is important. We carry out this task by the probabilistic integration of image features based on maximum a posteriori (MAP) probability that combines the prior and likelihood probabilities using Bayes' rule.

IROS Conference 1996 Conference Paper

A visual tracking algorithm by integrating rigid model and snakes

  • Dong Jung Kang
  • In-So Kweon

This paper presents a robust vision algorithm for tracking the boundary of an object with an arbitrary shape by using monocular image sequences. This method consists of a curve registration based optimization technique and a deformable contour model ("snakes") for the global and the local motion estimations, respectively. By combining techniques, we overcome, among other problems, inaccurate estimate of motion parameters in the curve registration method (which apparently only occur when a rigid or a flexible object is tracked), and the "local position variation" of the deformable contour model, variations, which are due to noisy images and/or complex backgrounds. The curve registration method uses an iterative algorithm to find the minimum normal distance between two curves, one before motion and the corresponding curve after it. Snakes overcome the limitation of the curve registration method, which suffers from the inaccuracy of motion models. We also propose an internal force, which increases local robustness of the deformable contour to background noise. By using the refined snakes' control points, the global update of the previous curve is performed for the re-location of the registered curve. Additionally, we integrate the geometric invariant value of the boundary contour and the curve registration method to solve the occlusion problem in visual tracking. The proposed method is validated through experiments on real images.

IROS Conference 1995 Conference Paper

A Kalman filter based visual tracking algorithm for an object moving in 3D

  • Joon Woong Lee
  • Mun Sang Kim
  • In-So Kweon

Robust and effective real-time visual tracking is realized by combining the first order differential invariants with stochastic filtering. The Kalman filter as an optimal stochastic filter is used to estimate the motion parameters, namely the plant state vector of the moving object with the unknown dynamics in successive image frames. Using the fact that the relative motion between the moving object and the moving observer causes the deformation, we compute the first differential invariants of the image velocity field. The surface orientation and the depth estimate between the observer and the object are computed based on these first order differential invariants. We demonstrate the robustness and feasibility of the proposed tracking algorithm through real experiments in which an X-Y Cartesian robot tracks a toy vehicle moving along 3D rails.

ICRA Conference 1992 Conference Paper

Architecture of behavior-based mobile robot in dynamic environment

  • Mutsumi Watanabe
  • Kazunori Onoguchi
  • In-So Kweon
  • Yoshinori Kuno

The authors are developing a compact cart-type robot system which moves around an office. They propose a behavior-based architecture with three clustered (reflexive, purposive, and adaptive) agents that realizes efficiency in attaining the mission of the robot, and robustness against the various kinds of failures that may occur in a dynamic environment. The reflexive-level group consists of agents with contact, infrared, and ultrasonic sensors which maintain minimal safety of the robot. The role of the purposive-level group is to achieve the global mission of the robot, such as 'if the robot detects a small fire in the office, find and reach a fire extinguisher as soon as possible'. The adaptive-level group stands by to recover from failure in the purposive-level group or in a deadlock situation. Experimental results showed the effectiveness of the method. >

ICRA Conference 1992 Conference Paper

Behavior-based mobile robot using active sensor fusion

  • In-So Kweon
  • Yoshinori Kuno
  • Mutsumi Watanabe
  • Kazunori Onoguchi

The authors present a navigation system using multiple sensors for unknown and dynamic indoor environments. To achieve robustness and flexibility in the mobile robot, a behavior-based system architecture is developed consisting of multilayered behaviors using multiple sensors that were ultrasonic sensors and a video camera. Basic behaviors required for navigation, such as avoiding obstacles, moving toward free space, and following targets are redundantly developed as agents and combined in a behavior-based system architecture. The capabilities of the system were demonstrated in unstructured real office environments using an indoor mobile robot. >

ICRA Conference 1991 Conference Paper

Extracting topographic features for outdoor mobile robots

  • In-So Kweon
  • Takeo Kanade

Methods are presented for building high-level terrain descriptions, referred as topographic maps, by extracting terrain features like peaks, pits, ridges, and ravines from the contour map. The resulting topographic map contains the location and type of terrain features as well as the ground topography. The authors develop new definitions for those topographic features based on the contour map. They build a contour map from an elevation map and generate the connectivity tree of all regions separated by the contours. The authors use this connectivity tree, called a topographic change tree, to extract the topographic features. Experimental results on a digital elevation model support the definitions for topographic features and the approach. >

IROS Conference 1990 Conference Paper

High resolution terrain map from multiple sensor data

  • In-So Kweon
  • Takeo Kanade

Describes a terrain mapping 3D vision system to build a high resolution terrain map from multiple range images and a digital elevation model (DEM). To build a composite map of the environment from multiple sensor data, the terrain mapping system needs a representation of the terrain that must be appropriate for multiple sensor data. Building a composite terrain map also requires estimating motion between sensor views and merging these views into a composite map. The terrain representation described consists of a grid-based representation, called elevation map. The authors develop the locus method to build elevation maps from range images. The locus method uses a model of the sensor to interpolate at arbitrary resolution without making any assumptions on the terrain shape other than the continuity of the surface. They also present a pixel-based or iconic terrain matching algorithm to estimate the vehicle motion from a sequence of range images. This terrain matching method uses the locus method to solve correspondence and occlusion problems. Comprehensive test results using a long sequence of range images and a DEM for rugged outdoor terrain are given.

ICRA Conference 1989 Conference Paper

Terrain mapping for a roving planetary explorer

  • Martial Hebert
  • Claude Caillas
  • Eric Krotkov
  • In-So Kweon
  • Takeo Kanade

The authors are prototyping a legged vehicle, the Ambler, for an exploratory mission on another planet, conceivably Mars, where it is to traverse uncharted areas and collect material samples. They describe how the rover can construct from range imagery a geometric terrain representation, i. e. , elevation map that includes uncertainty, unknown areas, and local features. First, they present an algorithm for constructing an elevation map from a single range image. By virtue of working in spherical-polar space, the algorithm is independent of the desired map resolution and the orientation of the sensor, unlike algorithms that work in Cartesian space. Secondly, the authors present a two-stage matching technique (feature matching followed by iconic matching) that identifies the transformation T corresponding to the vehicle displacement between two viewing positions. Thirdly, to support legged locomotion over rough terrain, they describe methods for evaluating regions of the constructed elevation maps as footholds. >

v2026.09.13