Arrow Research search

Author name cluster

Xiangyu Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

EAAI Journal 2026 Journal Article

A multi-source domain-invariant acoustic feature extraction network for rotating machinery fault diagnosis under unknown cross-working conditions

  • Xinyuan Zhang
  • Xiangang Cao
  • Hongwei Fan
  • Xin Yang
  • Yong Duan
  • Fuyuan Zhao
  • Xiangyu Li

In mechanical systems under cross-working conditions, acoustic feature extraction for rotating machinery faces numerous challenges, including difficulty in acquiring high-quality training data from multi-source domains and large divergence between inter-domain samples. Existing methods often suffer from insufficient constraints in the embedding space and limited model generalization. To address these issues, this paper proposes a Multi-source Domain-Invariant Acoustic Feature Extraction Network (DIAFENet) for unknown cross-working condition tasks. DIAFENet introduces a multi-granularity adversarial learning framework that jointly optimizes classification loss, domain adversarial loss, and a novel feature association-based class boundary constraint loss, aiming to learn discriminative and operation-condition-invariant acoustic representations. The core innovation lies in a three-level domain-invariant feature association strategy: (1) Global-level alignment minimizes overall domain divergence; (2) Subdomain-level alignment refines local feature distribution consistency; (3) Maximization of inter-class association distance within subdomains explicitly enlarges decision margins between classes. The framework integrates a spectrogram-based feature extractor with a self-attention pooling mechanism, and employs a Gradient Reversal Layer (GRL) to adversarially eliminate domain-specific information and promote domain-invariant representation learning. The effectiveness of DIAFENet is rigorously evaluated on 15 cross-operating-condition tasks across two public datasets and a Self-built dataset. The results showed that the average classification accuracy of DIAFENet was 98. 93 %, 97. 31 %, and 96. 55 %. The ablation experiment further verified that the proposed feature association constraint strategy improved accuracy by 2. 75 %, demonstrating its key role in enhancing the compactness and separability of the embedding space. This study provides a reliable acoustic feature extraction scheme for intelligent diagnosis of mechanical equipment under multi-source domain cross-working condition tasks.

AAAI Conference 2026 Conference Paper

Ambiguity-aware Truncated Flow Matching for Ambiguous Medical Image Segmentation

  • Fanding Li
  • Xiangyu Li
  • Xianghe Su
  • Xingyu Qiu
  • Suyu Dong
  • Wei Wang
  • Kuanquan Wang
  • Gongning Luo

A simultaneous enhancement of accuracy and diversity of predictions remains a challenge in ambiguous medical image segmentation (AMIS) due to the inherent trade-offs. While truncated diffusion probabilistic models (TDPMs) hold strong potential with a paradigm optimization, existing TDPMs suffer from entangled accuracy and diversity of predictions with insufficient fidelity and plausibility. To address the aforementioned challenges, we propose Ambiguity-aware Truncated Flow Matching (ATFM), which introduces a novel inference paradigm and dedicated model components. Firstly, we propose Data-Hierarchical Inference, a redefinition of AMIS-specific inference paradigm, which enhances accuracy and diversity at data-distribution and data-sample level, respectively, for an effective disentanglement. Secondly, Gaussian Truncation Representation (GTR) is introduced to enhance both fidelity of predictions and reliability of truncation distribution, by explicitly modeling it as a Gaussian distribution at Ttrunc instead of using sampling-based approximations. Thirdly, Segmentation Flow Matching (SFM) is proposed to enhance the plausibility of diverse predictions by extending semantic-aware flow transformation in Flow Matching (FM). Comprehensive evaluations on LIDC and ISIC3 datasets demonstrate that ATFM outperforms SOTA methods and simultaneously achieves a more efficient inference. ATFM improves GED and HM-IoU by up to 12% and 7.3% compared to advanced methods.

EAAI Journal 2026 Journal Article

Development and application of a new cotton aphid detection network under complex background

  • Liangliang Liu
  • Fengjie Zhao
  • Jinpu Xie
  • Jing Chang
  • Shixin Qiao
  • Xiangyu Li
  • Hongbo Qiao

Development of aphid detection methods and applications is of great significance to the intelligent management of cotton. However, identifying cotton aphids from natural images faces the challenge of small size, high density, and background noise. In this study, we embed the improved Normalized Attention Module (NAM) and Swin-Transformer (ST) to reconstruct the backbone of You Only Look Once (YOLO) structure, and then constructed an improved YOLO model for cotton aphid detection in natural environment. NAM adopts channel and spatial attention submodules for small target detection. ST provides global pixels for complex background target detection. In addition, a mixed loss function is adopted to address the pixel-level feature imbalance problem of small targets in complex backgrounds. The improved YOLO achieves competitive performance on the collected cotton aphid dataset, with an average detection accuracy (mAP50) of 93. 36% and a precision of 95. 57%. Compared with representative YOLO-based detectors, the proposed model shows consistent improvements under complex background conditions, particularly for dense and small-target scenarios. Furthermore, we integrate the proposed detection network with Artificial Intelligence (AI) audio glasses, and realize the acquisition and detection of cotton aphid images through voice commands. Our study provides an AI application to the research on intelligent detection products for cotton aphids.

EAAI Journal 2026 Journal Article

Intelligent pose correction of shield machines via an integrated convolutional long short-term memory Kolmogorov-Arnold network and model reference adaptive control

  • Xiangyu Li
  • Xuanyu Liu
  • Limin Wang
  • Yudong Wang
  • He Zhang
  • Yueyang Huang
  • Junzhi Lu

In underground tunnel construction, Earth Pressure Balance shield machines are required to advance along a designed alignment. However, complex geological conditions and equipment-related disturbances often lead to pose deviations, which can compromise construction quality. This study proposes an integrated intelligent pose correction framework that combines pose prediction with adaptive control. First, key input variables are selected through Pearson correlation analysis and denoised using a hybrid Complete Ensemble Empirical Mode Decomposition with Adaptive Noise-wavelet transform method. A pose prediction model is then developed based on a Convolutional Long Short-Term Memory Kolmogorov-Arnold Network (CL-KAN), which replaces the fully connected layers of a Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) with KAN layers to enhance nonlinear feature representation. Experimental results show that the CL-KAN model achieves high prediction accuracy, with root mean squared error values ranging from 0. 88 to 1. 68 mm for vertical deviations and coefficients of determination ranging from 0. 90 to 0. 97 for the key pose parameters. Compared with a baseline CNN-LSTM, the CL-KAN model reduces the root mean squared error by 12. 3-18. 6% while requiring fewer trainable parameters. To bridge prediction and control, a context-aware perturbation importance analysis (CA-PIA) method is employed to identify influential control features, which subsequently guide the parameter optimization of a model reference adaptive control (MRAC) strategy. Field validation under complex working conditions demonstrates that the proposed framework confines pose deviations within ±7 mm, showing strong robustness and practical applicability for intelligent pose correction in tunnel engineering based on artificial intelligence techniques.

JBHI Journal 2026 Journal Article

TKRL: Targeted Knowledge Rectification Learning Against Teacher-Originated Defects in Domain Continual Segmentation

  • Zhanshi Zhu
  • Wenjian Gu
  • Xiangyu Li
  • Qince Li
  • Yongfeng Yuan
  • Wei Wang
  • Kuanquan Wang
  • Suyu Dong

Knowledge distillation can mitigate catastrophic forgetting in domain continual segmentation by transferring knowledge from the older model to the newer model. However, existing distillation-based methods primarily emphasize knowledge retention while overlooking inherent defects in the older teacher models. As a result, these teacher-originated defects, such as knowledge gaps or biases, are propagated and exacerbate forgetting. To address this challenge, we propose a Targeted Knowledge Rectification Learning framework (TKRL) to probe and correct teacher-originated defects. TKRL consists of two modules: (1) Probe-augmented Class Distillation, which generates gradient-driven “probes” to uncover underrepresented features in the older model, thereby bridging knowledge gaps by distilling hidden information into the new model; (2) Variance-guided Masked Autoencoder, which selectively masks and reconstructs critical high-uncertainty patches across multi-level semantic regions, thereby correcting biases inherited from the older model. Our experimental results show that TKRL effectively rectifies knowledge gaps and biases, thereby mitigating catastrophic forgetting and enhancing performance in domain continual segmentation. The implementation code is publicly available at: https://github.com/PerceptionComputingLab/TKRL_DCMIS.

AAAI Conference 2025 Conference Paper

Benchmarking and Understanding Compositional Relational Reasoning of LLMs

  • Ruikang Ni
  • Da Xiao
  • Qingye Meng
  • Xiangyu Li
  • Shihui Zheng
  • Hongliang Liang

Compositional relational reasoning (CRR) is a hallmark of human intelligence, but we lack a clear understanding of whether and how existing transformer large language models (LLMs) can solve CRR tasks. To enable systematic exploration of the CRR capability of LLMs, we first propose a new synthetic benchmark called Generalized Associative Recall (GAR) by integrating and generalizing the essence of several tasks in mechanistic interpretability (MI) study in a unified framework. Evaluation shows that GAR is challenging enough for existing LLMs, revealing their fundamental deficiency in CRR. Meanwhile, it is easy enough for systematic MI study. Then, to understand how LLMs solve GAR tasks, we use attribution patching to discover the core circuits reused by Vicuna-33B across different tasks, and a set of vital attention heads. Intervention experiments show that the correct functioning of these heads significantly impacts task performance. Especially, we identify two classes of heads whose activations represent the abstract notion of true and false in GAR tasks respectively. They play fundamental roles in CRR across various models and tasks.

JBHI Journal 2025 Journal Article

MedFILIP: Medical Fine-Grained Language-Image Pre-Training

  • Xinjie Liang
  • Xiangyu Li
  • Fanding Li
  • Jie Jiang
  • Qing Dong
  • Wei Wang
  • Kuanquan Wang
  • Suyu Dong

Medical vision-language pretraining (VLP) that leverages naturally-paired medical image-report data is crucial for medical image analysis. However, existing methods struggle to accurately characterize associations between images and diseases, leading to inaccurate or incomplete diagnostic results. In this work, we propose MedFILIP, a fine-grained VLP model, introduces medical image-specific knowledge through contrastive learning, specifically: 1) An information extractor based on a large language model is proposed to decouple comprehensive disease details from reports, which excels in extracting disease deals through flexible prompt engineering, thereby effectively reducing text complexity while retaining rich information at a tiny cost. 2) A knowledge injector is proposed to construct relationships between categories and visual attributes, which help the model to make judgments based on image features, and fosters knowledge extrapolation to unfamiliar disease categories. 3) A semantic similarity matrix based on fine-grained annotations is proposed, providing smoother, information-richer labels, thus allowing fine-grained image-text alignment. 4) We validate MedFILIP on numerous datasets, e. g. , RSNA-Pneumonia, NIH ChestX-ray14, VinBigData, and COVID-19. For single-label, multi-label, and fine-grained classification, our model achieves state-of-the-art performance, the classification accuracy has increased by a maximum of 6. 69%.

IJCAI Conference 2025 Conference Paper

Q-MiniSAM2: A Quantization-based Benchmark for Resource-Efficient Video Segmentation

  • Xuanxuan Ren
  • Xiangyu Li
  • Kun Wei
  • Xu Yang
  • Yanhua Yang

Segment Anything Model 2 (SAM2) is a new-generation, high-precision model for image and video segmentation, offering extensive application prospects across numerous computer vision fields. However, as a large-scale model, its huge memory demands and expansive computing costs pose challenges for practical deployment. This paper presents Q-MiniSAM2, an efficient Quantization-based segmentation benchmark tailored to optimize SAM2 by Minimizing memory consumption and accelerating computations. We begin with applying Post-Training Quantization (PTQ) to SAM2, requiring only a relatively small dataset for network calibration, thereby eliminating the need for retraining. Building upon PTQ, we further introduce a Hierarchy-based Video Quantization method to enhance the model’s capacity to capture video semantics and temporal correlations across different time scales. Furthermore, we observe that SAM2’s memory overhead is predominantly concentrated on processing historical frames, and the redundant cross-attention computations significantly increase memory and computational costs due to the imperceptible change of the short time intervals between these frames. To tackle this issue, an Adaptive Mutual-KV mechanism is proposed to mitigate excessive cross-attention by leveraging inter-frame similarities. Comprehensive experiments demonstrate that the proposed approach achieves superior performance compared to state-of-the-art methods, underscoring its potential for efficient and scalable video segmentation.

NeurIPS Conference 2025 Conference Paper

RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning

  • Hao Gao
  • Shaoyu Chen
  • Bo Jiang
  • Bencheng Liao
  • Yiang Shi
  • Xiaoyang Guo
  • Yuechuan Pu
  • haoran yin

Existing end-to-end autonomous driving (AD) algorithms typically follow the Imitation Learning (IL) paradigm, which faces challenges such as causal confusion and an open-loop gap. In this work, we propose RAD, a 3DGS-based closed-loop Reinforcement Learning (RL) framework for end-to-end Autonomous Driving. By leveraging 3DGS techniques, we construct a photorealistic digital replica of the real physical world, enabling the AD policy to extensively explore the state space and learn to handle out-of-distribution scenarios through large-scale trial and error. To enhance safety, we design specialized rewards to guide the policy in effectively responding to safety-critical events and understanding real-world causal relationships. To better align with human driving behavior, we incorporate IL into RL training as a regularization term. We introduce a closed-loop evaluation benchmark consisting of diverse, previously unseen 3DGS environments. Compared to IL-based methods, RAD achieves stronger performance in most closed-loop metrics, particularly exhibiting a 3× lower collision rate. Abundant closed-loop results are presented in the supplementary material. Code is available at https: //github. com/hustvl/RAD for facilitating future research.

ECAI Conference 2024 Conference Paper

MC-SORT: A Motion Correction-Based Framework for Long-Term Multiple Object Tracking

  • Xiangyu Li
  • Yunchuan Qin
  • Ruihui Li
  • Guanghua Tan
  • Zhuo Tang
  • Kenli Li 0001

Long-term occlusion is one of the most formidable challenges in Multi-Object Tracking (MOT). The motion models of existing SORT-based trackers are unreliable in estimating the motion states of long-term occluded targets. This is mainly because as the occlusion period increases, the increases speed of estimation errors in the motion model increases faster. In practical applications, we believe that the estimation error of the tracker during long-term occlusion is mainly concentrated in the estimation error of the motion model on the velocity of the occluded target. In this work, we have demonstrated that in the long-term occlusion period, appropriately correcting the estimated values of the motion model on the target motion velocity and fully utilizing the temporal and attribute information of the target’s historical trajectory as calculation indicators of correlation are beneficial for improving the robustness of the tracker in long-term occlusion. We refer to our proposed motion correction-based framework as MC-SORT, which mainly consists of a Momentum Compensation Module (MCM) and a Backtracking Re-association (BRA) module. The former can correct the estimated value of the target’s motion state during long-term occlusion, the latter uses the temporal and attribute information of the target’s historical trajectory during long-term occlusion as correlation indicators to measure the degree of correlation between the target and trajectory. Our proposed MC-SORT has the characteristics of simplicity, online, real-time, and plug-and-play, particularly improving the robustness of the tracker in long-term occlusion. The extensive experimental results on the MOT17 and MOT20 datasets demonstrate the robustness and superiority of our framework.

JBHI Journal 2022 Journal Article

Hematoma Expansion Context Guided Intracranial Hemorrhage Segmentation and Uncertainty Estimation

  • Xiangyu Li
  • Gongning Luo
  • Wei Wang
  • Kuanquan Wang
  • Yue Gao
  • Shuo Li

Accurate segmentation of the Intracranial Hemorrhage (ICH) in non-contrast CT images is significant for computer-aided diagnosis. Although existing methods have achieved remarkable 1 1 The code will be available from https://github.com/JohnleeHIT/SLEX-Net.results, none of them incorporated ICH’s prior information in their methods. In this work, for the first time, we proposed a novel SLice EXpansion Network (SLEX-Net), which incorporated hematoma expansion in the segmentation architecture by directly modeling the hematoma variation among adjacent slices. Firstly, a new module named Slice Expansion Module (SEM) was built, which can effectively transfer contextual information between two adjacent slices by mapping predictions from one slice to another. Secondly, to perceive contextual information from both upper and lower slices, we designed two information transmission paths: forward and backward slice expansion, and aggregated results from those paths with a novel weighing strategy. By further exploiting intra-slice and inter-slice context with the information paths, the network significantly improved the accuracy and continuity of segmentation results. Moreover, the proposed SLEX-Net enables us to conduct an uncertainty estimation with one-time inference, which is much more efficient than existing methods. We evaluated the proposed SLEX-Net and compared it with some state-of-the-art methods. Experimental results demonstrate that our method makes significant improvements in all metrics on segmentation performance and outperforms other existing uncertainty estimation methods in terms of several metrics.

AAAI Conference 2021 Conference Paper

Generalized Zero-Shot Learning via Disentangled Representation

  • Xiangyu Li
  • Zhe Xu
  • Kun Wei
  • Cheng Deng

Zero-Shot Learning (ZSL) aims to recognize images belonging to unseen classes that are unavailable in the training process, while Generalized Zero-Shot Learning (GZSL) is a more realistic variant that both seen and unseen classes appear during testing. Most GZSL approaches achieve knowledge transfer based on the features of samples that inevitably contain information irrelevant to recognition, bringing negative influence for the performance. In this work, we propose a novel method, dubbed Disentangled-VAE, which aims to disentangle category-distilling factors and category-dispersing factors from visual as well as semantic features, respectively. In addition, a batch re-combining strategy on latent features is introduced to guide the disentanglement, encouraging the distilling latent features to be more discriminative for recognition. Extensive experiments demonstrate that our method outperforms the state-of-the-art approaches on four challenging benchmark datasets.

v2026.09.13