Arrow Research search

Author name cluster

Xing Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

JBHI Journal 2026 Journal Article

GPFD-Net: A Geometry-Pose Frequency Decoupling Network for Privacy-Preserving Human Action Recognition in Healthcare

  • Xing Li
  • Jingfan Liang
  • Ge Gao
  • Li Wang
  • Haifeng Wang
  • Shihao Han

Human Action Recognition (HAR) holds significant application value in healthcare informatics, facilitating tasks such as clinical diagnosis and rehabilitation monitoring. Point cloud sequences have emerged as a pivotal modality for balancing privacy preservation with high-fidelity geometric structural representation, ensuring anonymity while retaining critical 3D behavioral information. However, existing point cloud sequence encoding methods struggle to precisely encode micro-geometric details and macro-pose contours within the spatial dimension, as well as the dynamic heterogeneity of actions within the temporal dimension. These limitations impede the realization of high-precision clinical motion analysis. To address these challenges, we propose a Geometry-Pose Frequency Decoupling Network (GPFD-Net) for human action recognition. First, we design a Geometry-Pose Parallel-Collaborative Spatial Encoder (GPCSE). This module designs a parallel dual-stream architecture to explicitly capture and fuse complementary micro-geometric details and macro-pose contours, generating an informative geometry-enhanced pose feature sequence. Second, we introduce a Frequency-Decoupled Temporal Capturer (FDTC). This module adaptively decomposes the geometry-enhanced pose feature sequence into a smooth trend sequence and a transient detail sequence, which are subsequently processed by two parallel expert encoders via differentiated encoding to achieve robust human action recognition. Extensive experiments on four public benchmark datasets demonstrate that GPFD-Net achieves superior performance. The proposed method provides a novel paradigm for high-precision and privacy-preserving motion analysis in healthcare applications.

ICRA Conference 2025 Conference Paper

A Helping (Human) Hand in Kinematic Structure Estimation

  • Adrian Pfisterer
  • Xing Li
  • Vito Mengers
  • Oliver Brock

Visual uncertainties such as occlusions, lack of texture, and noise present significant challenges in obtaining accurate kinematic models for safe robotic manipulation. We introduce a probabilistic real-time approach that leverages the human hand as a prior to mitigate these uncertainties. By tracking the constrained motion of the human hand during manipulation and explicitly modeling uncertainties in visual observations, our method reliably estimates an object's kinematic model online. We validate our approach on a novel dataset featuring challenging objects that are occluded during manipulation and offer limited articulations for perception. The results demonstrate that by incorporating an appropriate prior and explicitly accounting for uncertainties, our method produces accurate estimates, outperforming two recent baselines by 195 % and 140 %, respectively. Furthermore, we demonstrate that our approach's estimates are precise enough to allow a robot to manipulate even small objects safely.

NeurIPS Conference 2025 Conference Paper

Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM Inference

  • Zijie Geng
  • Jie Wang
  • Ziqi Liu
  • Feng Ju
  • Yiming Li
  • Xing Li
  • Mingxuan Yuan
  • Jianye Hao

Key-Value (KV) cache eviction---which retains the KV pairs of the most important tokens while discarding less important ones---is a critical technique for optimizing both memory usage and inference latency in large language models (LLMs). However, existing approaches often rely on simple heuristics---such as attention weights---to measure token importance, overlooking the spatial relationships between token value states in the vector space. This often leads to suboptimal token selections and thus performance degradation. To tackle this problem, we propose a novel method, namely **AnDPro** (**An**chor **D**irection **Pro**jection), which introduces a projection-based scoring function to more accurately measure token importance. Specifically, AnDPro operates in the space of value vectors and leverages the projections of these vectors onto an *``Anchor Direction''*---the direction of the pre-eviction output---to measure token importance and guide more accurate token selection. Experiments on $16$ datasets from the LongBench benchmark demonstrate that AnDPro can maintain $96. 07\\%$ of the full cache accuracy using only $3. 44\\%$ KV cache budget, reducing KV cache budget size by $46. 0\\%$ without compromising quality compared to previous state-of-the-arts.

NeurIPS Conference 2025 Conference Paper

AttentionPredictor: Temporal Patterns Matter for KV Cache Compression

  • Qingyue Yang
  • Jie Wang
  • Xing Li
  • Zhihai Wang
  • Chen Chen
  • Lei Chen
  • Xianzhi Yu
  • Wulong Liu

With the development of large language models (LLMs), efficient inference through Key-Value (KV) cache compression has attracted considerable attention, especially for long-context generation. To compress the KV cache, recent methods identify critical KV tokens through static modeling of attention scores. However, these methods often struggle to accurately determine critical tokens as they neglect the *temporal patterns* in attention scores, resulting in a noticeable degradation in LLM performance. To address this challenge, we propose **AttentionPredictor**, which is the **first learning-based method to directly predict attention patterns for KV cache compression and critical token identification**. Specifically, AttentionPredictor learns a lightweight, unified convolution model to dynamically capture spatiotemporal patterns and predict the next-token attention scores. An appealing feature of AttentionPredictor is that it accurately predicts the attention score and shares the unified prediction model, which consumes negligible memory, among all transformer layers. Moreover, we propose a cross-token critical cache prefetching framework that hides the token estimation time overhead to accelerate the decoding stage. By retaining most of the attention information, AttentionPredictor achieves **13$\times$** KV cache compression and **5. 6$\times$** speedup in a cache offloading scenario with comparable LLM performance, significantly outperforming the state-of-the-arts. The code is available at https: //github. com/MIRALab-USTC/LLM-AttentionPredictor.

JBHI Journal 2025 Journal Article

Efficient Vulnerability Assessment of Multi-Bit Faults in Embedded Healthcare Software

  • Jinting Ren
  • Yu Wu
  • Hengyi Ren
  • Xing Li
  • Zhaoyang Han

With the growing use of wearable and implantable medical devices, ensuring the reliability of embedded software is crucial for patient safety. These devices are susceptible to soft errors, and traditional single-bit fault models often fail to account for multi-bit faults, leading to either insufficient protection or excessive overhead. We introduce V-FAME, a lightweight and efficient vulnerability analysis framework for embedded healthcare systems. V-FAME uses machine learning and a multi-bit fault model to identify fault-sensitive code regions quickly and accurately. Unlike existing tools that focus on instruction-level analysis, V-FAME uses a basic block-level classification approach to significantly prune the fault space. Our experiments show V-FAME achieves a speedup of over 6. 16x while maintaining an accuracy of over 80%. This framework supports reliable and cost-effective fault mitigation in medical applications.

NeurIPS Conference 2025 Conference Paper

Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal Transport

  • Taoran Zheng
  • Yan Yang
  • Xing Li
  • Xiang Gu
  • Jian Sun
  • Zongben Xu

Medical image reconstruction from measurement data is a vital but challenging inverse problem. Deep learning approaches have achieved promising results, but often requires paired measurement and high-quality images, which is typically simulated through a forward model, i. e. , retrospective reconstruction. However, training on simulated pairs commonly leads to performance degradation on real prospective data due to the retrospective-to-prospective gap caused by incomplete imaging knowledge in simulation. To address this challenge, this paper introduces imaging Knowledge-Informed Dynamic Optimal Transport (KIDOT), a novel dynamic optimal transport framework with optimality in the sense of preserving consistency with imaging physics in transport, that conceptualizes reconstruction as finding a dynamic transport path. KIDOT learns from unpaired data by modeling reconstruction as a continuous evolution path from measurements to images, guided by an imaging knowledge-informed cost function and transport equation. This dynamic and knowledge-aware approach enhances robustness and better leverages unpaired data while respecting acquisition physics. Theoretically, we demonstrate that KIDOT naturally generalizes dynamic optimal transport, ensuring its mathematical rationale and solution existence. Extensive experiments on MRI and CT reconstruction demonstrate KIDOT's superior performance. Code is available at https: //github. com/TaoranZheng717/KIDOT.

EAAI Journal 2024 Journal Article

An improved smoking behavior detection algorithm via incorporating an interference information filtering network

  • Yi Li
  • Haojie Zhou
  • Jing Feng
  • Xing Li
  • Xiaobin Xu
  • Pingzhi Hou
  • Xiaomin Hu

Efficient and accurate identification of smoking behavior in public places is crucial for ensuring public health and safety. However, due to various factors like small target size, complex image background, varying cigarette angles and numerous similar objects, current methods still grapple with challenges such as missed detection and false detection when identifying cigarette targets in smoking behavior. This paper presents an enhanced version of the You Only Look Once version5-small algorithm to address these issues effectively. Firstly, to bolster the model's ability for feature extraction of cigarette targets, the Swin Transformer Block structure is integrated into the backbone network to capture long-range dependencies. Secondly, a novel Hybrid Spatial Pyramid Pooling-Fast with Cross Stage Partial Connection module is built based on the foundation of the Spatial Pyramid Pooling-Fast module by integrating both maximum pooling and average pooling to enhance the fusion ability of the multi-scale feature maps. Thirdly, a novel interference information filtering network is introduced to effectively reduce the impact of noise and confusion caused by similar objects, thus enhancing the performance of the model. According to the experimental results, the accuracy of the improved algorithm on the self-made cigarette target image data set reaches 93. 5 %, and the recall reaches 89. 1 %.

EAAI Journal 2024 Journal Article

Discrete artificial bee colony algorithm with fixed neighborhood search for traveling salesman problem

  • Xing Li
  • Shaoping Zhang
  • Peng Shao

For the artificial bee colony algorithm (ABC), it is easy to fall into local optimum and has lower convergence accuracy when solving the traveling salesman problem. For addressing this demerit further, a discrete artificial bee colony algorithm with fixed neighborhood search for traveling salesman problem (TSP), called DABC-FNS, is proposed. In DABC-FNS, the solution obtained by the discrete artificial bee colony algorithm is expressed by positive integer coding method. Meanwhile, the local enhancement strategy and the 2-opt strategy with fixed neighborhood search are introduced to improve the solution accuracy of the ABC algorithm. In order to verify the effectiveness of the DABC-FNS algorithm, more than 30 benchmark TSP instances are simulated by DABC-FNS algorithm and other state-of -the-art competitors. The experimental results show that the DABC-FNS algorithm has achieved better accuracy for most TSP instances, which also demonstrates that it can overcome the premature phenomenon and has certain advantages in solving the traveling salesman problem.

NeurIPS Conference 2024 Conference Paper

Towards Next-Generation Logic Synthesis: A Scalable Neural Circuit Generation Framework

  • Zhihai Wang
  • Jie Wang
  • Qingyue Yang
  • Yinqi Bai
  • Xing Li
  • Lei Chen
  • Jianye Hao
  • Mingxuan Yuan

Logic Synthesis (LS) aims to generate an optimized logic circuit satisfying a given functionality, which generally consists of circuit translation and optimization. It is a challenging and fundamental combinatorial optimization problem in integrated circuit design. Traditional LS approaches rely on manually designed heuristics to tackle the LS task, while machine learning recently offers a promising approach towards next-generation logic synthesis by neural circuit generation and optimization. In this paper, we first revisit the application of differentiable neural architecture search (DNAS) methods to circuit generation and found from extensive experiments that existing DNAS methods struggle to exactly generate circuits, scale poorly to large circuits, and exhibit high sensitivity to hyper-parameters. Then we provide three major insights for these challenges from extensive empirical analysis: 1) DNAS tends to overfit to too many skip-connections, consequently wasting a significant portion of the network's expressive capabilities; 2) DNAS suffers from the structure bias between the network architecture and the circuit inherent structure, leading to inefficient search; 3) the learning difficulty of different input-output examples varies significantly, leading to severely imbalanced learning. To address these challenges in a systematic way, we propose a novel regularized triangle-shaped circuit network generation framework, which leverages our key insights for completely accurate and scalable circuit generation. Furthermore, we propose an evolutionary algorithm assisted by reinforcement learning agent restarting technique for efficient and effective neural circuit optimization. Extensive experiments on four different circuit benchmarks demonstrate that our method can precisely generate circuits with up to 1200 nodes. Moreover, our synthesized circuits significantly outperform the state-of-the-art results from several competitive winners in IWLS 2022 and 2023 competitions.

EAAI Journal 2023 Journal Article

A representation-learning-based approach to predict stock price trend via dynamic spatiotemporal feature embedding

  • Bowen Pang
  • Wei Wei
  • Xing Li
  • Xiangnan Feng
  • Chao Li

Stock price trend prediction is a fascinating but difficult research topic. Recently, GNN-based models have been continuously proposed, which are believed to be more effective since they consider the information about stocks themselves and the information between stocks. However, the graph data are often static, unstructured and not all-inclusive, which cannot dynamically reflect all relationships between stocks. Therefore, we propose a novel model PriceExploration-Network (PE-Net), effectively utilizing both temporal and cross-sectional information contained in price to predict the price trend. PE-Net only requires the price data effectively saving the trouble of fetching alternative data and is able to capture the dynamic implicit relations between stocks by combining clustering techniques and GAT architecture. The effectiveness of PE-Net is examined on real-world S&P 500 constituents and the results demonstrate that PE-Net can outperform state-of-the-art models w. r. t. both accuracy and AUC.

TIST Journal 2023 Journal Article

Representation Learning of Enhanced Graphs Using Random Walk Graph Convolutional Network

  • Xing Li
  • Wei Wei
  • Ruizhi Zhang
  • Zhenyu Shi
  • Zhiming Zheng
  • Xiangnan Feng

Nowadays, graph structure data has played a key role in machine learning because of its simple topological structure, and therefore, the graph representation learning methods have attracted great attention. And it turns out that the low-dimensional embedding representation obtained by graph representation learning is extremely useful in various typical tasks, such as node classification and content recommendation. However, most of the existing methods do not further dig out potential structural information on the original graph structure. Here, we propose wGCN, which utilizes random walk to obtain the node-specific mesoscopic structures (high-order local structure) of the graph and utilizes these mesoscopic structures to enhance the graph and organize the characteristic information of the nodes. Our method can effectively generate node embedding for data of previously unknown categories, which has been proven in a series of experiments conducted on many types of graph networks. And compared to baselines, our method shows the best performance on most datasets and achieves competitive results on others. It is believed that combining the mesoscopic structure to further explore the structural information of the graph will greatly improve the learning efficiency of the graph neural network.

v2026.09.13