Arrow Research search

Author name cluster

Sicheng Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
2 author rows

Possible papers

16

AAAI Conference 2026 Conference Paper

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

  • Sicheng Yang
  • Yukai Huang
  • Weitong Cai
  • Shitong Sun
  • You He
  • Jiankang Deng
  • Hang Zhang
  • Jifei Song

The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs) struggle to resolve these multimodal ambiguous inputs, often failing silently or hallucinating responses. To address these ambiguities, we introduce the Plug-and-Play Clarifier, a zero-shot and modular framework that decomposes the problem into discrete, solvable sub-tasks. Specifically, our framework consists of three synergistic modules: (1) a text clarifier that uses dialogue-driven reasoning to interactively disambiguate linguistic intent, (2) a vision clarifier that delivers real-time guidance feedback, instructing users to adjust their positioning for improved capture quality, and (3) a cross-modal clarifier with grounding mechanism that robustly interprets 3D pointing gestures and identifies the specific objects users are pointing to. Extensive experiments demonstrate that our framework improves the intent clarification performance of small language models (4-8B) by approximately 30%, making them competitive with significantly larger counterparts. We also observe consistent gains when applying our framework to these larger models. Furthermore, our vision clarifier increases corrective guidance accuracy by over 20%, and our cross-modal clarifier improves semantic answer accuracy for referential grounding by 5%. Overall, our method provides a plug-and-play framework that effectively resolves multimodal ambiguity and significantly enhances user experience in egocentric interaction.

AAAI Conference 2026 Conference Paper

VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling

  • Sicheng Yang
  • Xing Hu
  • Qiang Wu
  • Dawei Yang

Vector quantization (VQ) transforms continuous image features into discrete representations, providing compressed, tokenized inputs for generative models. However, VQ-based frameworks suffer from several issues, such as non-smooth latent spaces, weak alignment between representations before and after quantization, and poor coherence between the continuous and discrete domains. These issues lead to unstable codeword learning and underutilized codebooks, ultimately degrading the performance of both reconstruction and downstream generation tasks. To this end, we propose VAEVQ, which comprises three key components: (1) Variational Latent Quantization (VLQ), replacing the AE with a VAE for quantization to leverage its structured and smooth latent space, thereby facilitating more effective codeword activation; (2) Representation Coherence Strategy (RCS), adaptively modulating the alignment strength between pre- and post-quantization features to enhance consistency and prevent overfitting to noise; and (3) Distribution Consistency Regularization (DCR), aligning the entire codebook distribution with the continuous latent distribution to improve utilization. Extensive experiments on two benchmark datasets demonstrate that VAEVQ outperforms state-of-the-art methods.

IROS Conference 2025 Conference Paper

Robotic Hand Tool Use with Contact-Based Demonstration: The Case of Cucumber Peeling

  • Lingzi Xie
  • Shuai Wang 0007
  • Jingxiang Chen
  • Bidan Huang
  • Yi Zhang
  • Sicheng Yang
  • Yuyuan Chen
  • Wang Wei Lee

Robotic hand tool use has garnered significant attention from robotics researchers, because it enhances dexterity beyond the limitations imposed by manipulators with fixed tool configurations and human-involved manual tool changes. Despite extensive research, current methodologies predominantly focus on imitating human hand trajectories, often neglecting the pivotal role of tool-environment interaction. This study addresses this gap by exploring the task of cucumber peeling as a case study to implement contact-based demonstration strategies in robotic tool use. Our approach concentrates on the subtle tool contact behaviors that manifest through contact dynamics. Specifically, we select appropriate tool stiffness for the peeling tasks, which is captured via a handheld teaching device equipped with optical tactile sensors. Subsequently, object-level stiffness control strategies are employed to emulate these behaviors using a three-fingered robotic hand. Experimental results from real-world cucumber peeling trials substantiate our methodology, illustrating that the robotic hand can adjust contact through finger movements, thereby achieving humanlike peeling efficiency without necessitating alterations to the tool structure. This study not only demonstrates the feasibility of sophisticated tool use by robotic hands, but also highlights the critical importance of integrating tactile feedback to refine interaction with the environment.

NeurIPS Conference 2025 Conference Paper

VQ-Seg: Vector-Quantized Token Perturbation for Semi-Supervised Medical Image Segmentation

  • Sicheng Yang
  • Zhaohu Xing
  • Lei Zhu

Consistency learning with feature perturbation is a widely used strategy in semi-supervised medical image segmentation. However, many existing perturbation methods rely on dropout, and thus require a careful manual tuning of the dropout rate, which is a sensitive hyperparameter and often difficult to optimize and may lead to suboptimal regularization. To overcome this limitation, we propose VQ-Seg, the first approach to employ vector quantization (VQ) to discretize the feature space and introduce a novel and controllable Quantized Perturbation Module (QPM) that replaces dropout. Our QPM perturbs discrete representations by shuffling the spatial locations of codebook indices, enabling effective and controllable regularization. To mitigate potential information loss caused by quantization, we design a dual-branch architecture where the post-quantization feature space is shared by both image reconstruction and segmentation tasks. Moreover, we introduce a Post-VQ Feature Adapter (PFA) to incorporate guidance from a foundation model (FM), supplementing the high-level semantic information lost during quantization. Furthermore, we collect a large-scale Lung Cancer (LC) dataset comprising 828 CT scans annotated for central-type lung carcinoma. Extensive experiments on the LC dataset and other public benchmarks demonstrate the effectiveness of our method, which outperforms state-of-the-art approaches. Codes will be released.

IROS Conference 2024 Conference Paper

A High-Performance Anthropomorphic Robotic Arm for Household Applications

  • Tianliang Liu
  • Sicheng Yang
  • Jingchen Li 0001
  • Xiangchi Chen
  • Shuai Wang 0007
  • Xiao Teng
  • Wang Wei Lee
  • Xiong Li 0001

Anthropomorphic robotic arms, mimicking the structure and function of human arms, show great potential for helping people in various tedious and repetitive household tasks. However, such arms mostly consist of multiple serial links controlled independently by actuators at joints with high reduction ratios, posing challenges in household services in terms of load capacity, responsiveness, and safety. In this paper, we propose a high-performance anthropomorphic arm called TRX-Arm based on differential cable transmission, characterized by features of high dynamics, high load capacity, and inherent compliance. TRX-Arm is composed of three deferential cable-driven coupling joints and one independent roll joint. Thanks to the cable differential transmission, the joints are capable of achieving doubled torque and stiffness without replacing motors. To enhance safety in human-robot interaction, the actuators including motors, reducer, belt, and pulley are mounted at the shoulder near the base and drive the joints remotely using cables, thereby minimizing the inertia of the whole arm. The workspace of TRX-Arm has a volume of 1. 56 m 3, much larger than that of the human arm. Real experiments show its capabilities including high repeatability and load capacity as well as high dynamic behavior of a dual-arm robot platform built with TRX-Arms.

AAAI Conference 2024 Conference Paper

Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

  • Zunnan Xu
  • Yachao Zhang
  • Sicheng Yang
  • Ronghui Li
  • Xiu Li

This study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing during inference. To address this problem, we suggest using speech-derived multimodal priors to improve gesture generation. We introduce a novel method that separates priors from speech and employs multimodal priors as constraints for generating gestures. Our approach utilizes a chain-like modeling method to generate facial blendshapes, body movements, and hand gestures sequentially. Specifically, we incorporate rhythm cues derived from facial deformation and stylization prior based on speech emotions, into the process of generating gestures. By incorporating multimodal priors, our method improves the quality of generated gestures and eliminate the need for expensive setup preparation during inference. Extensive experiments and user studies confirm that our proposed approach achieves state-of-the-art performance.

NeurIPS Conference 2024 Conference Paper

MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models

  • Zunnan Xu
  • Yukang Lin
  • Haonan Han
  • Sicheng Yang
  • Ronghui Li
  • Yachao Zhang
  • Xiu Li

Gesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality. Recent advancements have utilized the diffusion model to improve gesture synthesis. However, the high computational complexity of these techniques limits the application in reality. In this study, we explore the potential of state space models (SSMs). Direct application of SSMs in gesture synthesis encounters difficulties, which stem primarily from the diverse movement dynamics of various body parts. The generated gestures may also exhibit unnatural jittering issues. To address these, we implement a two-stage modeling strategy with discrete motion priors to enhance the quality of gestures. Built upon the selective scan mechanism, we introduce MambaTalk, which integrates hybrid fusion modules, local and global scans to refine latent space representations. Subjective and objective experiments demonstrate that our method surpasses the performance of state-of-the-art models. Our project is publicly available at~\url{https: //kkakkkka. github. io/MambaTalk/}.

ICRA Conference 2024 Conference Paper

Thermoformed electronic skins for conformal tactile sensor arrays

  • Peng Lu
  • Jiaming Liang
  • Bidan Huang
  • Sicheng Yang
  • Wang Wei Lee

Robots and prostheses are increasingly designed with curvilinear surfaces for functional, aesthetic, aerodynamic, and safety reasons. Electronic skins (e-skins) capable of sensing contact location and pressure across complex, non-developable surfaces are essential for empowering next-generation robots with tactile awareness. This will facilitate safe and natural human-machine interactions while enhancing object manipulation capabilities. Despite the evident advantages of conformal e-skins, current fabrication methods face significant challenges in realizing their full potential. In this paper, we introduce thermoforming as a technique to efficiently fabricate tactile sensitive e-skins that conform to curvilinear surfaces. The performance, repeatability and uniformity of the sensors are characterized in detail. We also present a custom calibration pipeline where accurate digital replicas of conformal e-skins are generated for use in simulations. Finally, we demonstrate the benefits of 3D e-skins in a tool manipulation task.

IROS Conference 2024 Conference Paper

TRX-Hand5: An Anthropomorphic Hand with Integrated Tactile Feedback for Grasping and Manipulation in Human Environments

  • Sicheng Yang
  • Wang Wei Lee
  • Zhong Zhang 0015
  • Youda Xiong
  • Jiaming Liang
  • Peng Lu
  • Yonghui Zhu
  • Tianliang Liu

Objects of daily life are designed to suit the human hand. Without major modifications to these objects and our environments, robots will need end-effectors with human hand-like configuration and dexterity to efficiently operate on them. Tight integration of tactile and proprioceptive sensors are also critical to ensure robust execution of manipulation policies without sacrificing range-of-motion. Reliability is also key, and a mechanically robust, easy to repair end-effector is important to minimize downtime. To meet these challenges, we designed a 13 degree-of-freedom anthropomorphic hand with over 1000 tactile sensing elements, named TRX-Hand5. Also embedded within are positional encoders and cable tension sensors to provide proprioceptive perception. TRX-Hand5 has a novel biomimetic topology with six small posture motors in the palm to replicate the function of intrinsic hand muscles and five large power motors in the forearm to play the role of forearm flexor muscles. The whole hand weighs 2. 6 kg with its dimensions comparable to those of an adult male’s hand and is capable of actuating its fingertips at over 200°/s while exerting up to 22 N of force. The system can be disassembled in modules for easy maintenance.

ICML Conference 2024 Conference Paper

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

  • Zhaoliang Wan
  • Yonggen Ling
  • Senlin Yi
  • Lu Qi 0001
  • Wang Wei Lee
  • Minglei Lu
  • Sicheng Yang
  • Xiao Teng

This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the "Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision, touch, and proprioception, to enhance robotic manipulation. VinT-6D comprises 2 million VinT-Sim and 0. 1 million VinT-Real entries, collected via simulations in Mujoco and Blender and a custom-designed real-world platform. This dataset is tailored for robotic hands, offering models with whole-hand tactile perception and high-quality, well-aligned data. To the best of our knowledge, the VinT-Real is the largest considering the collection difficulties in the real-world environment so it can bridge the gap of simulation to real compared to the previous works. Built upon VinT-6D, we present a benchmark method that shows significant improvements in performance by fusing multi-modal information. The project is available at https: //VinT-6D. github. io/.

IROS Conference 2023 Conference Paper

A Unified Trajectory Generation Algorithm for Dynamic Dexterous Manipulation

  • Cheng Zhou
  • Wentao Gao
  • Weifeng Lu
  • Yanbo Long
  • Sicheng Yang
  • Longfei Zhao
  • Bidan Huang
  • Yu Zheng 0001

This paper proposes a novel efficient multi-phase trajectory generation algorithm for dynamic dexterous manipulation tasks, such as throwing, catching, dynamic regrasping, and dynamic handover, which can be decomposed into multiple manipulation primitives, including sticking, rolling, approaching, separating, colliding, and grasping. Each manipulation primitive is formulate as a free-terminal optimal control problem (OCP), aimed at computing the optimal pose (position and orientation) trajectories of the object and the robot subject to the pose and force linkage constraints between them and the expected force maintenance at contact. A single-arm regrasping task and a dual-arm dynamic handover task are conducted to demonstrate the effectiveness of the proposed algorithm.

IJCAI Conference 2023 Conference Paper

DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models

  • Sicheng Yang
  • Zhiyong Wu
  • Minglei Li
  • Zhensong Zhang
  • Lei Hao
  • Weihong Bao
  • Ming Cheng
  • Long Xiao

The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the rhythm and semantics of the gesture to the corresponding speech. To address these problems, we present DiffuseStyleGesture, a diffusion model based speech-driven gesture generation approach. It generates high-quality, speech-matched, stylized, and diverse co-speech gestures based on given speeches of arbitrary length. Specifically, we introduce cross-local attention and self-attention to the gesture diffusion pipeline to generate better speech matched and realistic gestures. We then train our model with classifier-free guidance to control the gesture style by interpolation or extrapolation. Additionally, we improve the diversity of generated gestures with different initial gestures and noise. Extensive experiments show that our method outperforms recent approaches on speech-driven gesture generation. Our code, pre-trained models, and demos are available at https: //github. com/YoungSeng/DiffuseStyleGesture.

ICRA Conference 2022 Conference Paper

Real-time Inertial Parameter Identification of Floating-Base Robots Through Iterative Primitive Shape Division

  • Jiafeng Xu
  • Yu Zheng 0001
  • Xinyang Jiang
  • Sicheng Yang
  • Lingzhu Xiang
  • Zhengyou Zhang

Dynamic models play a key role in robot motion generation and control and the identification of inertial parameters is a critical component for obtaining an accurate dynamic model of a robot. This paper presents a novel iterative primitive shape division method for the inertia parameter identification of floating-base robots. Describing a robot by a set of primitive shapes with uniform mass distributions, the method iteratively divides the primitive shapes into smaller ones and refines their masses, which quickly converges to yielding the true inertia parameters of the robot. This method guarantees the physical consistency of the obtained parameters, possesses a high computational efficiency for online deployment, and works without contact force measurement. Furthermore, it can be used to estimate the position and magnitude of an external load applied to the robot. Simulations and experiments on a quadruped robot have been conducted to verify the effectiveness and efficiency of the proposed method.

IROS Conference 2020 Conference Paper

Gain Scheduled Controller Design for Balancing an Autonomous Bicycle

  • Shuai Wang 0007
  • Leilei Cui 0002
  • Jie Lai
  • Sicheng Yang
  • Xiangyu Chen 0001
  • Yu Zheng 0001
  • Zhengyou Zhang
  • Zhong-Ping Jiang

In this paper, the gain scheduling technique is applied to design a balance controller for an autonomous bicycle with an inertia wheel. Previously, two different balance controllers are needed depending on whether the bicycle is stationary or dynamic. The switch between the two different controllers may cause the instability of the autonomous bicycle. Our proposed gain scheduled controller can balance the autonomous bicycle in both stationary and dynamic cases. A physical system is built and experiments are carried out to demonstrate the effectiveness of the gain scheduled controller.

IROS Conference 2020 Conference Paper

Nonlinear Balance Control of an Unmanned Bicycle: Design and Experiments

  • Leilei Cui 0002
  • Shuai Wang 0007
  • Jie Lai
  • Xiangyu Chen 0001
  • Sicheng Yang
  • Zhengyou Zhang
  • Zhong-Ping Jiang

In this paper, nonlinear control techniques are exploited to balance an unmanned bicycle with enlarged stability domain. We consider two cases. For the first case when the autonomous bicycle is balanced by the flywheel, the steering angle is set to zero, and the torque of the flywheel is used as the control input. The controller is designed based on the Interconnection and Damping Assignment Passivity Based Control (IDA-PBC) method. For the second case when the bicycle is balanced by the handlebar, the bicycle’s velocity is high, and the flywheel is turned off. The angular velocity of the handlebar is used as the control input and the balance controller is designed based on feedback linearization. In these cases, the global stability of the closed-loop unmanned bicycle is theoretically proved based on Lyapunov theory. The experiments are conducted to validate the efficacy of the proposed nonlinear balance controllers.

IROS Conference 2019 Conference Paper

Development of a Continuous Vertical-pulling Automatic Doffing Robot for the Ring Spinning

  • Wenzeng Zhang
  • Sicheng Yang
  • Chao Luo
  • Siyun Liu
  • Hong Fu

Doffing robot is an important part of the spinning process in the textile production. This paper analyzes the doffing process of spinning machines and points out the requirements of the structure and functions of the doffer. The locking two-finger gripper, the three-dimensional circulating operation mechanism, the collaborating locating mechanism with the toothed disc and the pre-loosening mechanism by rotating spindles are designed. On this basis the continuous vertical-pulling automatic doffing robot, named CVP doffing robot, for the ring spinning is developed. The kinematics and dynamics analysis of the CVP doffing robot are carried out. The structural parameters of the CVP doffing robot are optimized by establishing kinematics and dynamics models. The forces of pulling out cops before and after the pre-loosing operation are tested. On this basis, the strength of the key components is designed and checked. Finally, the performance of the CVP doffing robot is verified by the doffing experiment.

v2026.09.13