Arrow Research search

Author name cluster

Kai Chen 0028

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

13 papers
1 author row

Possible papers

13

IROS Conference 2025 Conference Paper

ColaDex: Contact-guided Optimization and VLM-assisted Selection for Task-oriented Dexterous Grasp Generation

  • Yiyao Ma
  • Kai Chen 0028
  • Xuecheng Xu
  • Zhongxiang Zhou
  • Liang Xie
  • Rong Xiong
  • Qi Dou 0001

Task-oriented dexterous grasp generation aims to generate stable and functional grasps that enable a robotic hand to effectively interact with objects to accomplish specific tasks. However, generating high-dimensional hand configurations that seamlessly adapt to diverse task requirements and object geometries remains a significant challenge. In this paper, we propose a novel pipeline called ColaDex to address this challenging problem. The core idea of ColaDex is to leverage a vision-language models (VLMs) to select the dexterous grasp from a set of candidates that aligns well with the task description. To this end, we first introduce a contact-guided optimization method to generate a set of high-quality grasp candidates around the object through analytical optimization. Subsequently, to effectively prompt VLMs with the sampled numerous grasp candidates, we propose an object-centric approach that adaptively represents a group of candidates as prototypical contact maps, learned based on the geometric relationships between the grasping hand and object shape. We then feed the task requirement and the generated prototypical contact maps into the VLM, enabling it to reason about grasp-object interactions and assess their alignment with the given task, ultimately selecting the grasp that best aligns with the task requirement. Extensive experiments demonstrate that our prototypical contact map is a more informative prompting mechanism than conventional RGB images, enabling ColaDex to consistently generate high-quality task-oriented grasps and achieve a high success rate across diverse objects and tasks.

IROS Conference 2025 Conference Paper

Gaussian Splatting with Reflectance Regularization for Endoscopic Scene Reconstruction

  • Chengkun Li
  • Kai Chen 0028
  • Shi Qiu
  • Jason Ying-Kuen Chan
  • Qi Dou 0001

Endoscopic reconstruction plays a crucial role in surgical robotics. The dynamic lighting conditions and integrated camera-light source in endoscopic scenes create a distinct reconstruction challenge: shape ambiguity. To mitigate this, we propose a Gaussian Splatting (GS) based framework for endoscopic scene reconstruction, enhanced with reflectance regularization. We embed every 3D Gaussian point with physical reflective attributes and combine this representation with a physically based inverse rendering framework. By jointly training 3DGS for view synthesis with this reflectance regularization, we are able to attain high-quality geometry without changing the volume rendering pipeline. Our experiments demonstrate the superiority in both geometry representation and rendering performance compared to existing GS approaches, making it a practical solution for endoscopic applications. Project is available at: https://med-air.github.io/GSR2.

ICRA Conference 2024 Conference Paper

Multi-objective Cross-task Learning via Goal-conditioned GPT-based Decision Transformers for Surgical Robot Task Automation

  • Jiawei Fu 0001
  • Yonghao Long 0001
  • Kai Chen 0028
  • Wang Wei
  • Qi Dou 0001

Surgical robot task automation has been a promising research topic for improving surgical efficiency and quality. Learning-based methods have been recognized as an interesting paradigm and been increasingly investigated. However, existing approaches encounter difficulties in long-horizon goal-conditioned tasks due to the intricate compositional structure, which requires decision-making for a sequence of sub-steps and understanding of inherent dynamics of goal-reaching tasks. In this paper, we propose a new learning-based framework by leveraging the strong reasoning capability of the GPT-based architecture to automate surgical robotic tasks. The key to our approach is developing a goal-conditioned decision transformer to achieve sequential representations with goal-aware future indicators in order to enhance temporal reasoning. Moreover, considering to exploit a general understanding of dynamics inherent in manipulations, thus making the model’s reasoning ability to be task-agnostic, we also design a cross-task pretraining paradigm that uses multiple training objectives associated with data from diverse tasks. We have conducted extensive experiments on 10 tasks using the surgical robot learning simulator SurRoL [1]. The results show that our new approach achieves promising performance and task versatility compared to existing methods. The learned trajectories can be deployed on the da Vinci Research Kit (dVRK) for validating its practicality in real surgical robot settings. Our project website is at: https://med-air.github.io/SurRoL.

ICRA Conference 2024 Conference Paper

RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation

  • Boshi An
  • Yiran Geng
  • Kai Chen 0028
  • Xiaoqi Li 0020
  • Qi Dou 0001
  • Hao Dong 0003

Robotic manipulation requires accurate perception of the environment, which poses a significant challenge due to its inherent complexity and constantly changing nature. In this context, RGB image and point-cloud observations are two commonly used modalities in visual-based robotic manipulation, but each of these modalities have their own limitations. Commercial point-cloud observations often suffer from issues like sparse sampling and noisy output due to the limits of the emission-reception imaging principle. On the other hand, RGB images, while rich in texture information, lack essential depth and 3D information crucial for robotic manipulation. To mitigate these challenges, we propose an image-only robotic manipulation framework that leverages an eye-on-hand monocular camera installed on the robot’s parallel gripper. By moving with the robot gripper, this camera gains the ability to actively perceive the object from multiple perspectives during the manipulation process. This enables the estimation of 6D object poses, which can be utilized for manipulation. While, obtaining images from more and diverse viewpoints typically improves pose estimation, it also increases the manipulation time. To address this trade-off, we employ a reinforcement learning policy to synchronize the manipulation strategy with active perception, achieving a balance between 6D pose accuracy and manipulation efficiency. Our experimental results in both simulated and real-world environments showcase the state-of-the-art effectiveness of our approach. We believe that our method will inspire further research on real-world-oriented robotic manipulation. See https://rgbmanip.github.io/for more details.

ICRA Conference 2023 Conference Paper

Demonstration-Guided Reinforcement Learning with Efficient Exploration for Task Automation of Surgical Robot

  • Tao Huang
  • Kai Chen 0028
  • Bin Li 0082
  • Yun-Hui Liu 0001
  • Qi Dou 0001

Task automation of surgical robot has the potentials to improve surgical efficiency. Recent reinforcement learning (RL) based approaches provide scalable solutions to surgical automation, but typically require extensive data collection to solve a task if no prior knowledge is given. This issue is known as the exploration challenge, which can be alleviated by providing expert demonstrations to an RL agent. Yet, how to make effective use of demonstration data to improve exploration efficiency still remains an open challenge. In this work, we introduce Demonstration-guided EXploration (DEX), an efficient reinforcement learning algorithm that aims to overcome the exploration problem with expert demonstrations for surgical automation. To effectively exploit demonstrations, our method estimates expert-like behaviors with higher values to facilitate productive interactions, and adopts non-parametric regression to enable such guidance at states unobserved in demonstration data. Extensive experiments on 10 surgical manipulation tasks from SurRoL, a comprehensive surgical simulation platform, demonstrate significant improvements in the exploration efficiency and task success rates of our method. Moreover, we also deploy the learned policies to the da Vinci Research Kit (dVRK) platform to show the effectiveness on the real robot. Code is available at https://github.com/med-air/DEX.

ICRA Conference 2023 Conference Paper

StereoPose: Category-Level 6D Transparent Object Pose Estimation from Stereo Images via Back-View NOCS

  • Kai Chen 0028
  • Stephen James
  • Congying Sui
  • Yun-Hui Liu 0001
  • Pieter Abbeel
  • Qi Dou 0001

Most existing methods for category-level pose estimation rely on object point clouds. However, when considering transparent objects, depth cameras are usually not able to capture high-quality data, resulting in point clouds with severe artifacts. Without a complete point cloud, existing methods are not applicable to challenging transparent objects. To tackle this problem, we present StereoPose, a novel stereo image based framework for category-level object pose estimation, ideally suited for transparent objects. For a robust estimation from pure stereo images, we develop a pipeline that decouples category-level pose estimation into object size estimation, initial pose estimation, and pose refinement. StereoPose then estimates object pose based on representation in the normalized object coordinate space (NOCS). To address the issue of image content aliasing, we further define a back-view NOCS map for the transparent object. The back-view NOCS aims to reduce the network learning ambiguity caused by content aliasing, and leverage informative cues on the back of the transparent object for more accurate pose estimation. To further improve the performance of the stereo framework, StereoPose is equipped with a parallax attention module for stereo feature fusion and an epipolar loss for improving the stereo-view consistency of network predictions. Extensive experiments on the public TOD dataset demonstrate the superiority of the proposed StereoPose framework for category-level 6D transparent object pose estimation. Code and demos will be available on the project homepage: www. cse. cuhk.edu.hk/~kaichen/stereopose.html.

IROS Conference 2023 Conference Paper

Value-Informed Skill Chaining for Policy Learning of Long-Horizon Tasks with Surgical Robot

  • Tao Huang
  • Kai Chen 0028
  • Wang Wei
  • Jianan Li 0006
  • Yonghao Long 0001
  • Qi Dou 0001

Reinforcement learning is still struggling with solving long-horizon surgical robot tasks which involve multiple steps over an extended duration of time due to the policy exploration challenge. Recent methods try to tackle this problem by skill chaining, in which the long-horizon task is decomposed into multiple subtasks for easing the exploration burden and subtask policies are temporally connected to complete the whole long-horizon task. However, smoothly connecting all subtask policies is difficult for surgical robot scenarios. Not all states are equally suitable for connecting two adjacent subtasks. An undesired terminate state of the previous subtask would make the current subtask policy unstable and result in a failed execution. In this work, we introduce value-informed skill chaining (ViSkill), a novel reinforcement learning framework for long-horizon surgical robot tasks. The core idea is to distinguish which terminal state is suitable for starting all the following subtask policies. To achieve this target, we introduce a state value function that estimates the expected success probability of the entire task given a state. Based on this value function, a chaining policy is learned to instruct subtask policies to terminate at the state with the highest value so that all subsequent policies are more likely to be connected for accomplishing the task. We demonstrate the effectiveness of our method on three complex surgical robot tasks from SurRoL, a comprehensive surgical simulation platform, achieving high task success rates and execution efficiency. Code is available at https://github.com/med-air/ViSkill.

IROS Conference 2023 Conference Paper

Visual-Kinematics Graph Learning for Procedure-Agnostic Instrument Tip Segmentation in Robotic Surgeries

  • Jiaqi Liu
  • Yonghao Long 0001
  • Kai Chen 0028
  • Cheuk Hei Leung
  • Zerui Wang
  • Qi Dou 0001

Accurate segmentation of surgical instrument tip is an important task for enabling downstream applications in robotic surgery, such as surgical skill assessment, tool-tissue interaction and deformation modeling, as well as surgical autonomy. However, this task is very challenging due to the small sizes of surgical instrument tips, and significant variance of surgical scenes across different procedures. Although much effort has been made on visual-based methods, existing segmentation models still suffer from low robustness thus not usable in practice. Fortunately, kinematics data from the robotic system can provide reliable prior for instrument location, which is consistent regardless of different surgery types. To make use of such multi-modal information, we propose a novel visual-kinematics graph learning framework to accurately segment the instrument tip given various surgical procedures. Specifically, a graph learning framework is proposed to encode relational features of instrument parts from both image and kinematics. Next, a cross-modal contrastive loss is designed to incorporate robust geometric prior from kinematics to image for tip segmentation. We have conducted experiments on a private paired visual-kinematics dataset including multiple procedures, i. e. , prostatectomy, total mesorectal excision, fundoplication and distal gastrectomy on cadaver, and distal gastrectomy on porcine. The leave-one-procedure-out cross validation demon-strated that our proposed multi-modal segmentation method significantly outperformed current image-based state-of-the-art approaches, exceeding averagely 11. 2% on Dice.

IROS Conference 2022 Conference Paper

SESR: Self-Ensembling Sim-to-Real Instance Segmentation for Auto-Store Bin Picking

  • Biqi Yang
  • Xiaojie Gao
  • Kai Chen 0028
  • Rui Cao
  • Yidan Feng
  • Xianzhi Li 0001
  • Qi Dou 0001
  • Chi-Wing Fu

Instance segmentation is an important task for supporting robotic grasping in auto-store scenarios. Accurate segmentation usually relies on the quantity and quality of available annotated training data. However, it requires tremendous cost to obtain these labels. In this work, without requiring any human annotations on real data, our proposed self-ensembling sim-to-real network, namely SESR, is able to generate precise instance masks for a wide variety of supermarket goods. We design our SESR with a teacher model and a student model trained with a self-ensembling strategy. We adopt different levels of consistency to bridge the sim-to-real gap and boost the model generalization ability. Also, we compile an auto-store bin-picking dataset covering various goods. Extensive experiments on both unseen scenarios and unseen objects validate the effectiveness and superiority of our method over others, and the robot arm demonstrations further show that our segmentation results can support real-time auto-store bin picking.

ICRA Conference 2022 Conference Paper

Towards Robust Part-aware Instance Segmentation for Industrial Bin Picking

  • Yidan Feng
  • Biqi Yang
  • Xianzhi Li 0001
  • Chi-Wing Fu
  • Rui Cao
  • Kai Chen 0028
  • Qi Dou 0001
  • Mingqiang Wei

Industrial bin picking is a challenging task that requires accurate and robust segmentation of individual object instances. Particularly, industrial objects can have irregular shapes, that is, thin and concave, whereas in bin-picking scenarios, objects are often closely packed with strong occlusion. To address these challenges, we formulate a novel part-aware instance segmentation pipeline. The key idea is to decompose industrial objects into correlated approximate convex parts and enhance the object-level segmentation with part-level segmentation. We design a part-aware network to predict part masks and part-to-part offsets, followed by a part aggregation module to assemble the recognized parts into instances. To guide the network learning, we also propose an automatic label decoupling scheme to generate ground-truth part-level labels from instance-level labels. Finally, we contribute the first instance segmentation dataset, which contains a variety of industrial objects that are thin and have non-trivial shapes. Extensive experimental results on various industrial objects demonstrate that our method can achieve the best segmentation results compared with the state-of-the-art approaches.

IROS Conference 2021 Conference Paper

Category-Level 6D Object Pose Estimation via Cascaded Relation and Recurrent Reconstruction Networks

  • Jiaze Wang
  • Kai Chen 0028
  • Qi Dou 0001

Category-level 6D pose estimation, aiming to predict the location and orientation of unseen object instances, is fundamental to many scenarios such as robotic manipulation and augmented reality, yet still remains unsolved. Precisely recovering instance 3D model in the canonical space and accurately matching it with the observation is an essential point when estimating 6D pose for unseen objects. In this paper, we achieve accurate category-level 6D pose estimation via cascaded relation and recurrent reconstruction networks. Specifically, a novel cascaded relation network is dedicated for advanced representation learning to explore the complex and informative relations among instance RGB image, instance point cloud and category shape prior. Furthermore, we design a recurrent reconstruction network for iterative residual refinement to progressively improve the reconstruction and correspondence estimations from coarse to fine. Finally, the instance 6D pose is obtained leveraging the estimated dense correspondences between the instance point cloud and the reconstructed 3D model in the canonical space. We have conducted extensive experiments on two well-acknowledged benchmarks of category-level 6D pose estimation, with significant performance improvement over existing approaches. On the representatively strict evaluation metrics of 3D 75 and 5°2cm, our method exceeds the latest state-of-the-art SPD [1] by 4. 9% and 17. 7% on the CAMERA25 dataset, and by 2. 7% and 8. 5% on the REAL275 dataset. Codes are avaliable at https://wangjiaze.cn/projects/6DPoseEstimation.html.

IROS Conference 2021 Conference Paper

CLMM-Net: Robust Cascaded LiDAR Map Matching based on Multi-Level Intensity Map

  • Kai Chen 0028
  • Lei He
  • Xiaofeng Wang
  • Yuqian Liu
  • Ming Zhao

LiDAR map matching(LMM) is a critical localization technique in autonomous driving while existing methods have problems in terms of both accuracy and robustness when driving in the scenes with poor structure information (e. g. highways). This paper put forward a multi-level intensity map based cascaded network for LiDAR map matching in autonomous driving. The network uses an effective multi-level intensity map representation to compactly encode the appearance and structure information of point clouds, which effectively reduce the position ambiguity in structure-less scenarios. Besides, this method leverages the multi-scale nature of deep neural networks and matches the online LiDAR observation with the offline map in a coarse-to-fine manner so as to balance the time-consuming and precision. Extensive experiments on diverse autonomous driving environments demonstrate the superiority of our proposed method over other existing state-of-the-art methods.

IROS Conference 2019 Conference Paper

Line-based Absolute and Relative Camera Pose Estimation in Structured Environments

  • Haoang Li
  • Ji Zhao 0001
  • Jean-Charles Bazin
  • Wen Chen 0021
  • Kai Chen 0028
  • Yun-Hui Liu 0001

3D lines in structured environments encode particular regularity like parallelism and orthogonality. We leverage this structural regularity to estimate the absolute and relative camera poses. We decouple the rotation and translation, and propose a novel rotation estimation method. We decompose the absolute and relative rotations and reformulate the problem as computing the rotation from the Manhattan frame to the camera frame. To compute this rotation, we propose an accurate and efficient two-step method. We first estimate its two degrees of freedom (DOF) by two image lines, and then estimate its third DOF by another image line. For these lines, we assume their associated 3D lines are mutually orthogonal, or two 3D lines are parallel to each other and orthogonal to the third. Thanks to our two-step DOF estimation, our absolute and relative pose estimation methods are accurate and efficient. Moreover, our relative pose estimation method relies on weaker assumptions or less correspondences than existing approaches. We also propose a novel strategy to reject outliers and identify dominant directions of the scene. We integrate it into our pose estimation methods, and show that it is more robust than RANSAC. Experiments on synthetic and real-world datasets demonstrated that our methods outperform state-of-the-art approaches.

v2026.09.13