Arrow Research search

Author name cluster

Yuanqing Xia

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAAI Conference 2026 Conference Paper

GraphGrasp: Lightweight and Efficient Graph-Guided 6-DoF Robotic Grasp Pose Estimation Network

  • Sheng Yu
  • Di-Hua Zhai
  • Yuanqing Xia

6-DoF object grasping is a crucial skill for embodied intelligent robots. Previous methods often rely on large-scale networks for feature extraction, followed by grasp pose prediction, which increases the network's parameter count and overlooks the geometric and graph features of the point cloud. To address these challenges, we propose GraphGrasp, a graph-guided 6-DoF grasping pose prediction method. It performs graph analysis from the perspectives of scene, object, and grasping graphs. First, we introduce a graph feature embedding method based on local-global features to model the scene graph effectively. Then, we use a graph transformer strategy to represent spatial relationships between objects in the object graph. Finally, we propose a multi-metric, multi-level grasp pose evaluation algorithm to predict and explore graspable points, enabling effective construction of grasp graphs and accurate grasp pose evaluation. We test GraphGrasp on the GraspNet-1Billion dataset, and the results show that, compared to previous methods, it achieves nearly the same performance with about 1/5 of the parameters of state-of-the-art methods, significantly improving grasp pose prediction speed. Additionally, in real-world robot grasping scenarios, GraphGrasp outperforms previous methods in practical grasp pose prediction tasks.

EAAI Journal 2025 Journal Article

Advancements in collision avoidance techniques for internet-connected vehicles: A comprehensive review of methods and challenges

  • Khurrum Jalil
  • Yuanqing Xia
  • Jing Zhao

Collision avoidance (CA) in internet-connected vehicles (ICVs) is critical for ensuring safety and efficiency in smart transportation systems. The ICV control system achieves CA through integrated features, including sensor-based perception, communication technologies, and data-driven artificial intelligence, enabling real-time optimization for smooth cruising. By using these adaptive architectures on one platform, ICVs enhances both individual-level vehicle performance and network-wide traffic efficiency. This review examines a wide range of CA methods and strategies, offering comprehensive insights into how ICV systems detect and avoid obstacles in dynamic environments. We critically assess existing research, evaluating the effectiveness, challenges, and future directions of CA techniques, with particular attention to static and dynamic obstacle handling and interactions with other road users. Furthermore, we systematically explore ICV control systems, emphasizing how integrated technologies improve safe mobility and accident prevention. Our research synthesizes findings from peer-reviewed journals and conference proceedings (primarily from the past decade) to support the development of robust CAsystems. These insights aim to advance reliable CA frameworks for ICVs, fostering safer and more efficient transportation networks.

IS Journal 2025 Journal Article

AI-Based Hate Speech Detection System Using Video URLs for Effective Content Moderation

  • Zohaib Ahmad Khan
  • Yuanqing Xia
  • Fiza Khaliq
  • Weiwei Jiang
  • Muhammad Shahid Anwar

Countering online hate speech is essential for creating a safer digital space where positive interactions can thrive. As central hubs of global communication, platforms like social media platforms require effective moderation through explainable and affective computing approaches. This study introduces a novel artificial intelligence-driven system for detecting misogynstic discourse. We collected 11, 245 YouTube video uniform resource locators using specific keywords, then extracted audio to create Urdu transcripts and transliterated them into Roman Urdu, resulting in two distinct datasets. Various feature sets were explored using classic machine learning and deep learning algorithms. The results showed that classical models achieved 0. 90 accuracy on the Urdu dataset, while deep learning models reached 0. 96 accuracy on Roman Urdu. The corpus is publicly available to promote transparency and further research. Comparative evaluations against existing English hate speech dataset demonstrate the effectiveness of the proposed approach. This work lays the foundation for more ethical and transparent content moderation systems.

AAAI Conference 2025 Conference Paper

KeyPose: Category-Level 6D Object Pose Estimation with Self-Adaptive Keypoints

  • Sheng Yu
  • Di-Hua Zhai
  • Yuanqing Xia

Category-level object pose estimation is an important task in computer vision. Some prior methods based on assumptions often struggle with drastic changes in object appearance. To address this challenge, we propose a new method for object pose estimation based on object-adaptive keypoints. In this paper, we first introduce a transformer-based keypoint prediction method for adaptive forecasting of point cloud keypoints. This method calculates the similarity between keypoint features and point cloud features, allowing keypoints to represent object geometry more effectively. Furthermore, to enhance the geometric feature construction of keypoints, we propose a graph-based keypoint feature aggregation method, which considers both the structural relationships between keypoints and the point cloud, strengthening the network's understanding of geometric structures. At this stage, keypoints remain at the geometric spatial level of the object and have not been predicted in NOCS. To improve the accuracy of keypoint prediction in NOCS, we design a NOCS voxelization method that divides NOCS into multiple voxels and accurately predicts NOCS keypoints within these voxels. Experimental results on multiple benchmark datasets demonstrate that our proposed KeyPose method outperforms all existing methods, achieving over 20% improvement in pose accuracy on some critical datasets.

IROS Conference 2025 Conference Paper

RCGNet: RGB-based Category-Level 6D Object Pose Estimation with Geometric Guidance

  • Sheng Yu 0009
  • Di-Hua Zhai
  • Yuanqing Xia

While most current RGB-D-based category-level object pose estimation methods achieve strong performance, they face significant challenges in scenes lacking depth information. In this paper, we propose a novel category-level object pose estimation approach that relies solely on RGB images. This method enables accurate pose estimation in real-world scenarios without the need for depth data. Specifically, we design a transformer-based neural network for category-level object pose estimation, where the transformer is employed to predict and fuse the geometric features of the target object. To ensure that these predicted geometric features faithfully capture the object’s geometry, we introduce a geometric feature-guided algorithm, which enhances the network’s ability to effectively represent the object’s geometric information. Finally, we utilize the RANSAC-PnP algorithm to compute the object’s pose, addressing the challenges associated with variable object scales in pose estimation. Experimental results on benchmark datasets demonstrate that our approach is not only highly efficient but also achieves superior accuracy compared to previous RGB-based methods. These promising results offer a new perspective for advancing category-level object pose estimation using RGB images.

AAAI Conference 2024 Conference Paper

CatFormer: Category-Level 6D Object Pose Estimation with Transformer

  • Sheng Yu
  • Di-Hua Zhai
  • Yuanqing Xia

Although there has been significant progress in category-level object pose estimation in recent years, there is still considerable room for improvement. In this paper, we propose a novel transformer-based category-level 6D pose estimation method called CatFormer to enhance the accuracy pose estimation. CatFormer comprises three main parts: a coarse deformation part, a fine deformation part, and a recurrent refinement part. In the coarse and fine deformation sections, we introduce a transformer-based deformation module that performs point cloud deformation and completion in the feature space. Additionally, after each deformation, we incorporate a transformer-based graph module to adjust fused features and establish geometric and topological relationships between points based on these features. Furthermore, we present an end-to-end recurrent refinement module that enables the prior point cloud to deform multiple times according to real scene features. We evaluate CatFormer's performance by training and testing it on CAMERA25 and REAL275 datasets. Experimental results demonstrate that CatFormer surpasses state-of-the-art methods. Moreover, we extend the usage of CatFormer to instance-level object pose estimation on the LINEMOD dataset, as well as object pose estimation in real-world scenarios. The experimental results validate the effectiveness and generalization capabilities of CatFormer. Our code and the supplemental materials are avaliable at https://github.com/BIT-robot-group/CatFormer.

ICML Conference 2024 Conference Paper

Probabilistic Time Series Modeling with Decomposable Denoising Diffusion Model

  • Tijin Yan
  • Hengheng Gong
  • Yongping He
  • Yufeng Zhan
  • Yuanqing Xia

Probabilistic time series modeling based on generative models has attracted lots of attention because of its wide applications and excellent performance. However, existing state-of-the-art models, based on stochastic differential equation, not only struggle to determine the drift and diffusion coefficients during the design process but also have slow generation speed. To tackle this challenge, we firstly propose decomposable denoising diffusion model ($\text{D}^3\text{M}$) and prove it is a general framework unifying denoising diffusion models and continuous flow models. Based on the new framework, we propose some simple but efficient probability paths with high generation speed. Furthermore, we design a module that combines a special state space model with linear gated attention modules for sequence modeling. It preserves inductive bias and simultaneously models both local and global dependencies. Experimental results on 8 real-world datasets show that $\text{D}^3\text{M}$ reduces RMSE and CRPS by up to 4. 6% and 4. 3% compared with state-of-the-arts on imputation tasks, and achieves comparable results with state-of-the-arts on forecasting tasks with only 10 steps.

EAAI Journal 2023 Journal Article

AdaDerivative optimizer: Adapting step-sizes by the derivative term in past gradient information

  • Weidong Zou
  • Yuanqing Xia
  • Weipeng Cao

AdaBelief fully utilizes “belief” to iteratively update the parameters of deep neural networks. However, the reliability of the “belief” is determined by the gradient’s prediction accuracy, and the key to this prediction accuracy is the selection of the smoothing parameter β 1. AdaBelief also suffers from the overshoot problem, which occurs when the value of parameters exceeds the value of the target and cannot be changed along the gradient direction. In this paper, we propose AdaDerivative to eliminate the overshoot problem of AdaBelief. The key to AdaDerivative is that the “belief” of AdaBelief is replaced by the derivative term’s exponential moving average (EMA), which can be constructed as ( 1 − β 2 ) ∑ i = 1 t β 2 t − i ( g i − g i − 1 ) 2 based on the past and current gradients. We validate the performance of AdaDerivative on a variety of tasks, including image classification, language modeling, node classification, image generation, and object detection tasks. Extensive experimental results demonstrate that AdaDerivative can achieve state-of-the-art performance.

EAAI Journal 2022 Journal Article

Broad learning system based on driving amount and optimization solution

  • Weidong Zou
  • Yuanqing Xia
  • Weipeng Cao

Broad learning system (BLS) was proposed by C. L. Philip Chen to overcome the time-consuming problem of traditional deep learning. However, the prediction precision of BLS is mainly dependent on its regularized parameter λ. Usually, λ is calculated by the trial and error method, which often suffers from the problem of too much calculation. To alleviate this issue, we propose an improved BLS with the driving amount and optimization solution (i. e. , DA-BLS) in the study. The contributions of this study include: First, we use the iterative least square method to replace the ridge regression calculation of BLS, which avoids the selection of λ. Second, we provide the formulas of the driving amount and optimization solution under specific conditions. Third, the universal approximation property of DA-BLS is given. Last but not the least, extensive experimental results on the 1-D nonlinear function, UCI data-sets, and fault diagnosis of TEP show that DA-BLS outperforms the relevant methods such as BLS and the stochastic configuration network.

v2026.09.13