Arrow Research search

Author name cluster

Yuchen Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

AAAI Conference 2026 Conference Paper

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

  • Linfeng Dong
  • Yuchen Yang
  • Hao Wu
  • Wei Wang
  • Yuenan Hou
  • Zhihang Zhong
  • Xiao Sun

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research into complex human-object interactions. It is designed to tackle three interconnected tasks: fine-grained ball tracking, articulated racket pose estimation, and predictive ball trajectory forecasting. Our evaluation of established baselines reveals a critical insight for multi-modal fusion: while naively concatenating racket pose features degrades performance, a Cross-Attention mechanism is essential to unlock their value, leading to trajectory prediction results that surpass strong unimodal baselines. RacketVision provides a versatile resource and a strong starting point for future research in dynamic object tracking, conditional motion forecasting, and multi-modal analysis in sports.

ICLR Conference 2025 Conference Paper

CL-DiffPhyCon: Closed-loop Diffusion Control of Complex Physical Systems

  • Long Wei
  • Haodong Feng
  • Yuchen Yang
  • Ruiqi Feng
  • Peiyan Hu
  • Xiang Zheng
  • Tao Zhang
  • Dixia Fan

The control problems of complex physical systems have broad applications in science and engineering. Previous studies have shown that generative control methods based on diffusion models offer significant advantages for solving these problems. However, existing generative control approaches face challenges in both performance and efficiency when extended to the closed-loop setting, which is essential for effective control. In this paper, we propose an efficient Closed-Loop Diffusion method for Physical systems Control (CL-DiffPhyCon). By employing an asynchronous denoising framework for different physical time steps, CL-DiffPhyCon generates control signals conditioned on real-time feedback from the system with significantly reduced computational cost during sampling. Additionally, the control process could be further accelerated by incorporating fast sampling techniques, such as DDIM. We evaluate CL-DiffPhyCon on two tasks: 1D Burgers' equation control and 2D incompressible fluid control. The results demonstrate that CL-DiffPhyCon achieves superior control performance with significant improvements in sampling efficiency. The code can be found at https://github.com/AI4Science-WestlakeU/CL_DiffPhyCon.

IROS Conference 2025 Conference Paper

Hierarchical Trajectory Planning Method for Piano-Playing Robot

  • Zirui Wang
  • Jiayu Zhang
  • Wei Jiang
  • Tao Jiang
  • Jingdong Zhao
  • Liangliang Zhao
  • Baoshi Cao
  • Le Qi

Piano-playing tasks, which effectively demonstrate bimanual coordination capabilities in humanoid robots, are increasingly becoming a research focus. However, prior research has predominantly focused on Cartesian space trajectory planning without adequately addressing real-world obstacle avoidance constraints and manipulator acceleration limits. This paper proposes a hierarchical trajectory planning framework that systematically incorporates both obstacle avoidance and acceleration constraints. Firstly, discrete Cartesian path points are generated using a dynamic programming approach; secondly, joint space path points are derived considering obstacle avoidance and joint limit constraints through dynamic programming; thirdly, the joint space trajectory is interpolated using a Jacobian inverse-based method; finally, the trajectory is refined using Model Predictive Control (MPC). Experimental results demonstrate that the proposed method produces trajectories satisfying both obstacle avoidance and acceleration constraints, enabling fluent piano piece execution in real-world environments.

IJCAI Conference 2025 Conference Paper

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

  • Shaojun E
  • Yuchen Yang
  • Jiaheng Wu
  • Yan Zhang
  • Tiejun Zhao
  • Ziyan Chen

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively correlated with the coupling between visual encoders and large language models. Existing approaches often face issues such as vector gaps or semantic disparities, resulting in information loss during the propagation process. To address these issues, we propose MAGE (Multimodal Alignment and Generation Enhancement), a novel framework that bridges the semantic spaces of vision and text through an innovative alignment mechanism. By introducing the Intelligent Alignment Network (IAN), MAGE achieves dimensional and semantic alignment. To reduce the gap between synonymous heterogeneous data, we employ a training strategy that combines cross-entropy and mean squared error, significantly enhancing the alignment effect. Moreover, to enhance MAGE’s “Any-to-Any” capability, we developed a fine-tuning dataset for multimodal tool-calling instructions to expand the model’s output capability boundaries. Finally, our proposed multimodal large model architecture, MAGE, achieved significantly better performance compared to similar works across various evaluation benchmarks, including MME, MMBench, and SEED. Complete code and appendix are available at: https: //github. com/GTCOM-NLP/MAGE

ICRA Conference 2024 Conference Paper

A Combination of a Controllable Clutch and an Oscillating Slider Crank Mechanism for Ease of Direct-Teaching with Various Payloads

  • Muhammad Arifin
  • Yuta Kage
  • Yuchen Yang
  • Alexander Schmitz
  • Shigeki Sugano

Direct teaching is a straightforward way of teaching new motion to robots. Active methods with torque sensors, for example, can be used so that the robot can follow the movements of the human, but such methods introduce delays. Alternatively, series clutch actuators are easily backdrivable without delay. However, vertical joints are subject to gravity torques, which need to be compensated when disengaging the clutch. We implemented passive gravity compensation to counteract the robot’s weight, but this mechanism cannot compensate for varying payloads, as adjustable passive gravity compensation is relatively slow and mechanically complex. The varying payload causes an unintended joint movement, i. e. the arm falls down on its own, which is unacceptable during direct teaching. Therefore, this paper demonstrates how the torque output controlled with series clutch actuators can be used to compensate for varying payloads while maintaining high backdrivability. The proposed method is evaluated on a collaborative robot with a clutch in series for each actuator. Real-world experiments with payloads from 0 to 3 kg are conducted. During the experiments, the operator force is measured to evaluate the proposed method.

ICRA Conference 2024 Conference Paper

CLIPUNetr: Assisting Human-robot Interface for Uncalibrated Visual Servoing Control with CLIP-driven Referring Expression Segmentation

  • Chen Jiang
  • Yuchen Yang
  • Martin Jägersand

The classical human-robot interface in uncalibrated image-based visual servoing (UIBVS) relies on either human annotations or semantic segmentation with categorical labels. Both methods fail to match natural human communication and convey rich semantics in manipulation tasks as effectively as natural language expressions. In this paper, we tackle this problem by using referring expression segmentation, which is a prompt-based approach, to provide more in-depth information for robot perception. To generate high-quality segmentation predictions from referring expressions, we propose CLIPUNetr - a new CLIP-driven referring expression segmentation network. CLIPUNetr leverages CLIP’s strong vision-language representations to segment regions from referring expressions, while utilizing its "U-shaped" encoder-decoder architecture to generate predictions with sharper boundaries and finer structures. Furthermore, we propose a new pipeline to integrate CLIPUNetr into UIBVS and apply it to control robots in real-world environments. In experiments, our method improves boundary and structure measurements by an average of 120% and can successfully assist real-world UIBVS control in an unstructured manipulation environment.

ICRA Conference 2024 Conference Paper

Monocular Localization with Semantics Map for Autonomous Vehicles

  • Jixiang Wan
  • Xudong Zhang
  • Shuzhou Dong
  • Yuwei Zhang
  • Yuchen Yang
  • Ruoxi Wu
  • Ye Jiang
  • Jijunnan Li

Accurate and robust localization remains a significant challenge for autonomous vehicles. The cost of sensors and limitations in local computational efficiency make it difficult to scale to large commercial applications. Traditional vision-based approaches focus on texture features that are susceptible to changes in lighting, season, perspective, and appearance. Additionally, the large storage size of maps with descriptors and complex optimization processes hinder system performance. To balance efficiency and accuracy, we propose a novel lightweight visual semantic localization algorithm that employs stable semantic features instead of low-level texture features. First, semantic maps are constructed offline by detecting semantic objects, such as ground markers, lane lines, and poles, using cameras or LiDAR sensors. Then, online visual localization is performed through data association of semantic features and map objects. We evaluated our proposed localization framework in the publicly available KAIST Urban dataset and in scenarios recorded by ourselves. The experimental results demonstrate that our method is a reliable and practical localization solution in various autonomous driving localization tasks.

ICML Conference 2024 Conference Paper

Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation

  • Yuchen Yang
  • Yingdong Shi
  • Cheems Wang
  • Xiantong Zhen
  • Yuxuan Shi
  • Jun Xu 0019

Fine-tuning pretrained large models to downstream tasks is an important problem, which however suffers from huge memory overhead due to large-scale parameters. This work strives to reduce memory overhead in fine-tuning from perspectives of activation function and layer normalization. To this end, we propose the Approximate Backpropagation (Approx-BP) theory, which provides the theoretical feasibility of decoupling the forward and backward passes. We apply our Approx-BP theory to backpropagation training and derive memory-efficient alternatives of GELU and SiLU activation functions, which use derivative functions of ReLUs in the backward pass while keeping their forward pass unchanged. In addition, we introduce a Memory-Sharing Backpropagation strategy, which enables the activation memory to be shared by two adjacent layers, thereby removing activation memory usage redundancy. Our method neither induces extra computation nor reduces training efficiency. We conduct extensive experiments with pretrained vision and language models, and the results demonstrate that our proposal can reduce up to $\sim$$30%$ of the peak memory usage. Our code is released at github.

AAAI Conference 2024 Conference Paper

Robustness-Guided Image Synthesis for Data-Free Quantization

  • Jianhong Bai
  • Yuchen Yang
  • Huanpeng Chu
  • Hualiang Wang
  • Zuozhu Liu
  • Ruizhe Chen
  • Xiaoxuan He
  • Lianrui Mu

Quantization has emerged as a promising direction for model compression. Recently, data-free quantization has been widely studied as a promising method to avoid privacy concerns, which synthesizes images as an alternative to real training data. Existing methods use classification loss to ensure the reliability of the synthesized images. Unfortunately, even if these images are well-classified by the pre-trained model, they still suffer from low semantics and homogenization issues. Intuitively, these low-semantic images are sensitive to perturbations, and the pre-trained model tends to have inconsistent output when the generator synthesizes an image with low semantics. To this end, we propose Robustness-Guided Image Synthesis (RIS), a simple but effective method to enrich the semantics of synthetic images and improve image diversity, further boosting the performance of data-free compression tasks. Concretely, we first introduce perturbations on input and model weight, then define the inconsistency metrics at feature and prediction levels before and after perturbations. On the basis of inconsistency on two levels, we design a robustness optimization objective to eliminate low-semantic images. Moreover, we also make our approach diversity-aware by forcing the generator to synthesize images with small correlations. With RIS, we achieve state-of-the-art performance for various settings on data-free quantization and can be extended to other data-free compression tasks.

IROS Conference 2022 Conference Paper

Pose Refinement with Joint Optimization of Visual Points and Lines

  • Shuang Gao
  • Jixiang Wan
  • Yishan Ping
  • Xudong Zhang
  • Shuzhou Dong
  • Yuchen Yang
  • Haikuan Ning
  • Jijunnan Li

High-precision camera re-localization technology in a pre-established 3D environment map is the basis for many tasks, such as Augmented Reality, Robotics and Autonomous Driving. The point-based visual re-localization approaches are well-developed in recent decades, but are insufficient in some feature-less cases. In this paper, we design a complete pipeline for camera pose refinement with points and lines, which contains the innovatively designed line extracting CNN named VLSE, the line matching and the pose optimization approaches. We adopt a novel line representation and customize a hybrid convolution block based on the Stacked Hourglass network [1], to detect accurate and stable line features on images. Then we apply a geometric-based strategy to obtain precise 2D-3D line correspondences using epipolar constraint and reprojection filtering. A following point-line joint cost function is constructed to optimize the camera pose with the initial coarse pose from the pure point-based localization. Sufficient experiments are conducted on open datasets, i. e, line extractor on Wireframe and YorkUrban, localization performance on InLoc ducl and duc2, to confirm the effectiveness of our point-line joint pose optimization method.

ICRA Conference 2021 Conference Paper

Retrieval and Localization with Observation Constraints

  • Yuhao Zhou
  • Huanhuan Fan
  • Shuang Gao
  • Yuchen Yang
  • Xudong Zhang
  • Jijunnan Li
  • Yandong Guo

Accurate visual re-localization is very critical to many artificial intelligence applications, such as augmented reality, virtual reality, robotics and autonomous driving. To accomplish this task, we propose an integrated visual re-localization method called RLOCS by combining image retrieval, semantic consistency and geometry verification to achieve accurate estimations. The localization pipeline is designed as a coarse-to-fine paradigm. In the retrieval part, we cascade the architecture of ResNet101-GeM-ArcFace and employ DBSCAN followed by spatial verification to obtain a better initial coarse pose. We design a module called observation constraints, which combines geometry information and semantic consistency for filtering outliers. Comprehensive experiments are conducted on open datasets, including retrieval on R-Oxford5k and R-Paris6k, semantic segmentation on Cityscapes, localization on Aachen Day-Night and InLoc. By creatively modifying separate modules in the total pipeline, our method achieves many performance improvements on the challenging localization benchmarks.

v2026.09.13