Arrow Research search

Author name cluster

Ye Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

46 papers
2 author rows

Possible papers

46

EAAI Journal 2026 Journal Article

A real-time vehicle detection method in unmanned aerial vehicle images with selective contextual features

  • Wanxia Huang
  • Chaojun Dong
  • Xiankun Liu
  • Ye Li
  • Yikui Zhai
  • Kaitong Ou
  • Hao Quan

Detecting small and occluded objects in unmanned aerial vehicle (UAV) images remains a critical challenge. The inferior feature quality of these small and occluded objects leads to incomplete feature extraction, resulting in missed detections. To address this challenge, we propose an innovative detector based on ObjectBox to enhance detection performance and reduce missed detections of small and occluded objects by incorporating the neck called selective fused deformable context feature path aggregation network (SFDCFPAN) and the decoupled head. Firstly, we designed a neck called the selective feature path aggregation network (SFPAN) to fuse features and reduce the loss of spatial information. Subsequently, we provide a feature extraction module named fused deformable context feature extraction module (FDC) to model object shapes and then fuse context features to obtain the object’s semantic and spatial information. We employ the FDC module as the feature extraction module at specific locations within four feature layers of SFPAN, denoted as SFDCFPAN, to enhance the detector’s feature extraction and modeling capabilities. Lastly, we introduce a decoupled head structure to alleviate the mutual interference between classification and localization tasks. We conduct a comparative analysis of our detector with popular detectors on the VisDrone2019 and the UAVDT sub-dataset. Experimental results demonstrate the superior performance of our detector, achieving high accuracy on the two datasets while meeting real-time constraints. Furthermore, we integrate SFPAN and SFDCFPAN into various detectors. Experimental results exhibit the substantial enhancement in detector accuracy achieved by these feature fusion frameworks without compromising real-time performance, demonstrating the applicability to existing detectors.

AAAI Conference 2026 Conference Paper

Connectivity-Guided Sparsification of 2-FWL GNNs: Preserving Full Expressivity with Improved Efficiency

  • Rongqin Chen
  • Fan Mo
  • Pak Lon Ip
  • Shenghui Zhang
  • Dan Wu
  • Ye Li
  • Leong Hou U

Higher-order Graph Neural Networks (HOGNNs) based on the 2-FWL test achieve superior expressivity by modeling 2-node and 3-node interactions, but incur cubic computational cost. Existing efficiency methods typically reduce this burden at the expense of expressivity. We propose Co-Sparsify, a connectivity-aware sparsification framework that eliminates provably redundant computations while preserving full 2-FWL expressive power. Our key insight is that 3-node interactions are expressively necessary only within biconnected components, namely, maximal subgraphs where every node pair lies on a cycle. Outside these components, structural relationships are fully captured via 2-node message passing and graph readouts, rendering higher-order modeling unnecessary. Co-Sparsify restricts 2-node message passing to connected components and 3-node interactions to biconnected components, eliminating redundant computation without approximation or sampling. We prove that Co-Sparsified GNNs match the expressivity of the 2-FWL test. Empirically, when applied to PPGN, Co-Sparsify matches or exceeds accuracy on synthetic substructure counting tasks and achieves state-of-the-art performance on real-world benchmarks (ZINC, QM9 and TUD). This study demonstrates that high expressivity and scalability are not mutually exclusive: principled, topology-guided sparsification enables powerful, efficient GNNs with theoretical guarantees.

AAAI Conference 2026 Conference Paper

DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction

  • Xiao Yu
  • Zhaojie Fang
  • Guanyu Zhou
  • Yin Shen
  • Huoling Luo
  • Ye Li
  • Ahmed Elazab
  • Xiang Wan

Lung cancer continues to be the leading cause of cancer-related deaths globally. Early detection and diagnosis of pulmonary nodules are essential for improving patient survival rates. Although previous research has integrated multimodal and multi-temporal information, outperforming single modality and single time point, the fusion methods are limited to inefficient vector concatenation and simple mutual attention, highlighting the need for more effective multimodal information fusion. To address these challenges, we introduce a Dual-Graph Spatiotemporal Attention Network, which leverages temporal variations and multimodal data to enhance the accuracy of predictions. Our methodology involves developing a Global-Local Feature Encoder to better capture the local, global, and fused characteristics of pulmonary nodules. Additionally, a Dual-Graph Construction method organizes multimodal features into inter-modal and intra-modal graphs. Furthermore, a Hierarchical Cross-Modal Graph Fusion Module is introduced to refine feature integration. We also compiled a novel multimodal dataset named the NLST-cmst dataset as a comprehensive source of support for related research. Our extensive experiments, conducted on both the NLST-cmst and curated CSTL-derived datasets, demonstrate that our DGSAN significantly outperforms state-of-the-art methods in classifying pulmonary nodules with exceptional computational efficiency.

EAAI Journal 2026 Journal Article

Dual-channel feature fusion based image enhancement for low-cost train exterior fault detection

  • Haifeng Song
  • Yu Mei
  • Renxing Yin
  • Ye Li
  • Hairong Dong

Computer vision is a critical sensing modality for real-time train fault monitoring. However, low-light conditions often compromise the detection reliability of trackside systems. In low-cost setups, insufficient illumination degrades signal-to-noise ratios (SNR) and obscures structural features, rendering automated fault detection unreliable. To ensure robust sensing, this paper proposes a dual-cycle mapping framework for low-light image enhancement. The approach establishes bidirectional mapping between low-light and normal-light domains for adaptive brightness correction and color reconstruction. A Dual-Channel Restoration Unit (DCRU), featuring brightness correction and detail recovery branches, is introduced to mitigate uneven illumination and enhance feature discriminability. Furthermore, a Two-Dimensional Discrete Cosine Transform (2D-DCT) with dynamic weighting fuses frequency-domain features to optimize structural and local representations. Experimental results show the proposed method outperforms existing approaches, achieving a Peak Signal-to-Noise Ratio (PSNR) of 28. 08, a Structural Similarity Index Measure (SSIM) of 0. 93, and a Learned Perceptual Image Patch Similarity (LPIPS) of 0. 03. Integration with the You Only Look Once version 8 (YOLOv8) model demonstrates substantial downstream gains: precision and recall increased by 13% and 15%, respectively, while mean Average Precision (mAP) at an Intersection over Union (IoU) of 0. 5, and mAP across the IoU range of 0. 5-0. 95, improved by 10% and 18%. The proposed method provides a feasible solution for train exterior fault detection.

AAAI Conference 2026 Conference Paper

Mass Concept Erasure in Diffusion Models with Concept Hierarchy

  • Jiahang Tu
  • Ye Li
  • Yiming Wu
  • Hanbin Zhao
  • Chao Zhang
  • Hui Qian

The success of diffusion models has raised concerns about the generation of unsafe or harmful content, prompting concept erasure approaches that fine-tune modules to suppress specific concepts while preserving general generative capabilities. However, as the number of erased concepts grows, these methods often become inefficient and ineffective, since each concept requires a separate set of fine-tuned parameters and may degrade the overall generation quality. In this work, we propose a supertype-subtype concept hierarchy that organizes erased concepts into a parent–child structure. Each erased concept is treated as a child node, and semantically related concepts (e.g., macaw, and bald eagle) are grouped under a shared parent node, referred to as a supertype concept (e.g., bird). Rather than erasing concepts individually, we introduce an effective and efficient group-wise suppression method, where semantically similar concepts are grouped and erased jointly by sharing a single set of learnable parameters. During the erasure phase, standard diffusion regularization is applied to preserve denoising process in unmasked regions. To mitigate the degradation of supertype generation caused by excessive erasure of semantically related subtypes, we propose a novel method called Supertype-Preserving Low-Rank Adaptation (SuPLoRA), which encodes the supertype concept information in the frozen down-projection matrix and updates only the up-projection matrix during erasure. Theoretical analysis demonstrates the effectiveness of SuPLoRA in mitigating generation performance degradation. We construct a more challenging benchmark that requires simultaneous erasure of concepts across diverse domains, including celebrities, objects, and pornographic content. Comprehensive experiments demonstrate that our method achieves a superior balance between effective multi-concept erasure and the preservation of desirable generative performance.

JBHI Journal 2026 Journal Article

ULER: An Ultra-Lightweight Joint Attention Model for Emotion Recognition from Multimodal Physiological Signals

  • Haotian Liang
  • Geng-Xin Xu
  • Ting Jiang
  • Xichao Zhang
  • Dan Wu
  • Ye Li

Emotion recognition using multimodal physiological signals has gained widespread attention for providing objective insight into emotional states. However, existing models are often computationally complex and parameter heavy, hindering deployment on resource constrained wearable devices. To overcome this, we propose ULER (Ultra Lightweight Emotion Recognition), an efficient framework that integrates lightweight multi-scale convolution (LMSC) with a novel joint attention mechanism (LJA) based on a dynamic feature fusion network (DFFN) for effective multimodal fusion. Using subject dependent and total dataset training strategies, ULER is evaluated on three benchmarks (DEAP, DREAMER, WESAD) and outperforms recent state-of-the-art (SOTA) methods. It achieves accuracies of 99. 34%, 99. 46%, and 99. 23% on DEAP binary valence, binary arousal, and four class tasks, respectively, with only 0. 60 M parameters, 9. 33 M FLOPs, and 31. 10 ms inference latency. In a wearable oriented reduced channel setup (11 EEG channels), ULER also surpasses most SOTA models using standard 32 channels. Principal contributions emphasize: (1) a systematic comparison against recent SOTA methods on multiple datasets; (2) a wearable design with reduced-channel configuration; and (3) the DFFN as a core component for efficient multimodal fusion. This work demonstrates the potential for practical, real-time emotion monitoring in personal healthcare and other scenarios.

EAAI Journal 2025 Journal Article

A collaborative surface target detection and localization method for an unmanned surface vehicle swarm

  • Bo Wang
  • Chenyu Mao
  • Kaixin Wei
  • Xueyi Wu
  • Ye Li

A single unmanned surface vehicle (USV) designed for marine missions suffers from limited payload, low efficiency and weak intelligence, while a swarm of USVs shows significant advantages in mission flexibility, diverse payload and task efficiency. One of the key issues for an USV swarm is how to achieve highly efficient collaborative perception. To address this issue, a method framework of collaborative surface target detection and localization based on multiple sensors for a swarm including 4 USVs is designed. First, perception systems are constructed, a joint calibration method for different sensors is proposed, and a lightweight target detection method improved with attention mechanism and lightweight adaptive spatial feature fusion is designed. Second, a specialized fusion method using sensor principles based on an extended Kalman filter (EKF) is proposed for a single USV to obtain a target state model. Third, the obtained target models from different USVs are registered with fuzzy matching and integrated into the complete model in a geographic coordinate system. The proposed method is applied to the collaborative perception system on our developed 4 USV swarm and verified in real marine environment and simulation. Experimental results show that our proposed method framework significantly improves the accuracy, efficiency, and reliability of the target detection and localization. The proposed LAF-YOLOv8-s reduces the model size by 5. 1M, while the mean average precision (mAP) reaches 68. 7%, which is significantly superior to other methods. The average collaborative localization error is reduced by 2. 9m. The dataset is available at https: //github. com/maochenyu1/WSLight.

IJCAI Conference 2025 Conference Paper

A Priori Estimation of the Approximation, Optimization and Generalization Errors of Random Neural Networks for Solving Partial Differential Equations

  • Xianliang Xu
  • Ye Li
  • Zhongyi Huang

In recent years, neural networks have achieved remarkable progress in various fields and have also drawn much attention in applying them on scientific problems. A line of methods involving neural networks for solving partial differential equations (PDEs), such as Physics-Informed Neural Networks (PINNs) and the Deep Ritz Method (DRM), has emerged. Although these methods outperform classical numerical methods in certain cases, the optimization problems involving neural networks are typically non-convex and non-smooth, which can result in unsatisfactory solutions for PDEs. In contrast to deterministic neural networks, the hidden weights of random neural networks are sampled from some prior distribution and only the output weights participate in training. This makes training much simpler, but it remains unclear how to select the prior distribution. In this paper, we focus on Barron type functions and approximate them under Sobolev norms by random neural networks with clear prior distribution. In addition to the approximation error, we also derive bounds for the optimization and generalization errors of random neural networks for solving PDEs when the solutions are Barron type functions.

JBHI Journal 2025 Journal Article

BpBLS: A Knowledge-Embedded Bi-Incremental Broad Learning System for Wearable Cuffless Blood Pressure Estimation

  • Zi-Xuan Huang
  • Zeng-Ding Liu
  • Ye Li
  • C. L. Philip Chen
  • Hai-Liang Wang
  • Fen Miao

Cuffless blood pressure (BP) measurement has gained increasing attention due to the global aging population. Data-driven approaches have shown high accuracy for cuffless BP estimation. However, when deployed on wearable devices, they often suffer from being time-consuming because a complete retraining process is required if the training samples or structure need to be expanded. To address these issues, we propose a Knowledge-embedded Bi-incremental Broad Learning System (BpBLS) for cuffless BP estimation using biosignals collected from wearable devices. With a flat structure, BpBLS can be updated flexibly and quickly without retraining the entire model for incremental biosignals and/or training samples. In BpBLS, a novel Pulse Pressure Regularization (PPR) method is proposed to comprehensively capture the BP knowledge, which is then embedded into the system to enhance its accuracy. Two large-scale wearable BP datasets, the CAS-BP dataset and the Aurora-BP dataset, are utilized to validate the proposed BpBLS. Experimental results demonstrate that the proposed BpBLS exhibits superior performance in both estimation accuracy and computational efficiency compared to the state-of-the-art approaches. For the CAS-BP dataset, the estimation error is 0. 60 $\pm$ 8. 44 for systolic BP (SBP) and 0. 29 $\pm$ 6. 45 for diastolic BP (DBP); while for the Aurora-BP dataset, the estimation error is -0. 32 $\pm$ 8. 46 mmHg for SBP and 0. 05 $\pm$ 6. 65 mmHg for DBP. More importantly, the training time for both datasets is less than ten seconds, and the computational efficiency is improved by an order of magnitude compared to traditional machine learning methods. Our work will serve as a novel flexible and lightweight framework for cuffless BP measurement. Our code is available at https://github.com/Huangzx1023/BpBLS-for-BP-Estimation.

JBHI Journal 2025 Journal Article

cVAN: A Novel Sleep Staging Method via Cross-View Alignment Network

  • Zhanjiang Yang
  • Meiyu Qiu
  • Xiaomao Fan
  • Genan Dai
  • Wenjun Ma
  • Xiaojiang Peng
  • Xianghua Fu
  • Ye Li

Sleep staging is imperative for evaluating sleep quality and diagnosing sleep disorders. Extant sleep staging methods with fusing multiple data-views of physiological signals have achieved promising results. However, they remain neglectful of the relationship among different data-views at different feature scales with view position-alignment. To address this, we propose a novel cross-view alignment network, termed cVAN, utilising scale-aware attention for sleep stages classification. Specifically, cVAN principally incorporates two sub-networks of a residual-like network which learn spectral information from time-frequency images and a transformer-like network which learns corresponding temporal information. The prime advantage of cVAN is to adaptively align the learned feature scales among the different data-views of physiological signals with a scale-aware attention by reorganizing feature maps. Extensive experiments on three public sleep datasets demonstrate that cVAN can achieve a new state-of-the-art result, which is superior to existing counterparts.

IROS Conference 2025 Conference Paper

Design and Development of a Deformable Spherical Robot for Amphibious Applications *

  • Le Xu
  • Ruoyu Ren
  • Xiaojie Wei
  • Hao Lee
  • Hank Zhang
  • Kaicheng Yu
  • Ye Li
  • Chenyu Liu

This paper presents a deformable spherical robot with a six-strut topological structure capable of achieving multimodal locomotion in complex amphibious environments. The robot realizes isotropic rolling and asymmetric jumping through its innovative geometric-based configuration while integrating an airbag-driven module for underwater buoyancy control. Based on collision dynamics analysis, we develop a prototype of the deformable spherical robot. Experiments conducted on land, in transition zones, and underwater validate the robot’s multimodal locomotion feasibility in multi-medium environments.

AAAI Conference 2025 Conference Paper

Improving Generalization of Deep Neural Networks by Optimum Shifting

  • Yuyan Zhou
  • Ye Li
  • Lei Feng
  • Sheng-Jun Huang

Recent studies showed that the generalization of neural networks is correlated with the sharpness of the loss landscape and flat minima suggests a better generalization ability than sharp minima. In this paper, we propose a novel method called optimum shifting, which changes the parameters of a neural network from a sharp minimum to a flatter one while maintaining the same training loss value. Our method is based on the observation that when the input and output of a neural network are fixed, the matrix multiplications within the network can be treated as systems of under-determined linear equations, enabling adjustment of parameters in the solution space, which can be simply accomplished by solving a constrained optimization problem. Furthermore, we introduce a practical stochastic optimum shifting technique utilizing the neural collapse theory to reduce computational costs and provide more degrees of freedom for optimum shifting. Extensive experiments with various deep neural network architectures on benchmark datasets demonstrate the effectiveness of our method.

YNIMG Journal 2025 Journal Article

MR-guided graph learning of 18F-florbetapir PET enables accurate and interpretable Alzheimer’s disease staging

  • Xinyi Chen
  • Lijuan Chen
  • Weiheng Yao
  • Qiankun Zuo
  • Ye Li
  • Dong Liang
  • Shuqiang Wang
  • Meiyun Wang

PURPOSE: Subtle structural and molecular brain changes make noninvasive early detection and staging of Alzheimer's disease (AD) challenging and critical for effective intervention. This study develops a novel graph convolutional network (GCN) learning framework that integrates amyloid-β PET imaging and MRI structural features, aiming for improved early detection and accurate staging of AD. METHODS: The retrospective study utilized 18F-florbetapir PET scans from the Alzheimer's Disease Neuroimaging Initiative (ADNI) as the training dataset (323 scans from 196 subjects - 45 normal control, 80 mild cognitive impairment/MCI, 71 AD) and two independent datasets for testing (99 scans from 85 subjects - 31 normal control, 15 MCI, 44 AD). Individual brain graphs were constructed for each PET scan, and graph learning framework was designed to extract molecular features from PET while integrating structural features from MRI. Performance was evaluated using receiver operating characteristic (ROC) analysis, comparing results against cortical SUVR. Additionally, a biomarker GCN_score was defined based on identified salient regions-of-interest, with its effectiveness assessed using the Kruskal-Wallis test and Cohen's effect size. RESULTS: The framework achieved AUCs of 89.8 % (specificity 83.6 %, sensitivity 81.6 %) for distinguishing MCI from normal controls and 88.3 % (specificity 81.6 %, sensitivity 80.6 %) for MCI from AD in the ADNI dataset, with comparable performance in external testing. All results significantly outperformed cortical SUVR (DeLong test p < 0.001). The GCN_score demonstrated superior group differentiation (Cohen's effect sizes 1.744 and 1.32) compared to cortical SUVR (0.309 and 0.641). CONCLUSION: The proposed graph-based learning framework effectively integrates PET and MRI features for accurate AD stage distinction, showing significant promise for early detection and facilitating timely intervention.

JBHI Journal 2025 Journal Article

Multi-Scale Spatiotemporal Dynamic Graph Neural Network for Early Prediction of Mortality Risks in Heart Failure Patients

  • Longfei Liu
  • Rongqin Chen
  • Jifu Qu
  • Chunli Liu
  • Ye Li
  • Dan Wu

Heart Failure (HF) stands as a principal public health issue worldwide, imposing a significant burden on healthcare systems. While existing prognostic methods have achieved certain milestones in predicting the early mortality risk of HF patients, they have not fully considered the dynamic interdependencies among physiological parameters. This paper introduces a novel Multi-scale Spatiotemporal Dynamic Graph Neural Network, MSTD-GNN, which enhances the prediction capability for early mortality in HF patients by dynamically extracting spatio-temporal information of physiological parameters from ICU patient Electronic Health Records (EHRs). Our model constructs dynamic graphs to model multivariate time series data, revealing the implicit dependencies between physiological parameters and capturing the inherent dynamics of the data. We conducted experiments using the MIMIC-III and MIMIC-IV datasets. The experimental results show that, compared to existing methods, MSTD-GNN demonstrates superior performance in predicting the early mortality risk of HF patients. On the MIMIC-III and MIMIC-IV datasets, the AUC scores of MSTD-GNN reached 83. 93% and 81. 74%, respectively. Furthermore, through dynamic graphs, our model unveils the dynamic relationships between physiological variables across different time scales.

ICML Conference 2025 Conference Paper

Refined generalization analysis of the Deep Ritz Method and Physics-Informed Neural Networks

  • Xianliang Xu
  • Ye Li
  • Zhongyi Huang

In this paper, we derive refined generalization bounds for the Deep Ritz Method (DRM) and Physics-Informed Neural Networks (PINNs). For the DRM, we focus on two prototype elliptic partial differential equations (PDEs): Poisson equation and static Schrödinger equation on the $d$-dimensional unit hypercube with the Neumann boundary condition. Furthermore, sharper generalization bounds are derived based on the localization techniques under the assumptions that the exact solutions of the PDEs lie in the Barron spaces or the general Sobolev spaces. For the PINNs, we investigate the general linear second order elliptic PDEs with Dirichlet boundary condition using the local Rademacher complexity in the multi-task learning setting. Finally, we discuss the generalization error in the setting of over-parameterization when solutions of PDEs belong to Barron space.

ICLR Conference 2025 Conference Paper

Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video

  • Xiaohao Xu
  • Tianyi Zhang 0014
  • Shibo Zhao
  • Xiang Li 0106
  • Sibo Wang
  • Yongqi Chen
  • Ye Li
  • Bhiksha Raj

We aim to redefine robust ego-motion estimation and photorealistic 3D reconstruction by addressing a critical limitation: the reliance on noise-free data in existing models. While such sanitized conditions simplify evaluation, they fail to capture the unpredictable, noisy complexities of real-world environments. Dynamic motion, sensor imperfections, and synchronization perturbations lead to sharp performance declines when these models are deployed in practice, revealing an urgent need for frameworks that embrace and excel under real-world noise. To bridge this gap, we tackle three core challenges: scalable data generation, comprehensive benchmarking, and model robustness enhancement. First, we introduce a scalable noisy data synthesis pipeline that generates diverse datasets simulating complex motion, sensor imperfections, and synchronization errors. Second, we leverage this pipeline to create Robust-Ego3D, a benchmark rigorously designed to expose noise-induced performance degradation, highlighting the limitations of current learning-based methods in ego-motion accuracy and 3D reconstruction quality. Third, we propose Correspondence-guided Gaussian Splatting (CorrGS), a novel method that progressively refines an internal clean 3D representation by aligning noisy observations with rendered RGB-D frames from clean 3D map, enhancing geometric alignment and appearance restoration through visual correspondence. Extensive experiments on synthetic and real-world data demonstrate that CorrGS consistently outperforms prior state-of-the-art methods, particularly in scenarios involving rapid motion and dynamic illumination. We will release our code and benchmark to advance robust 3D vision, setting a new standard for ego-motion estimation and high-fidelity reconstruction in noisy environments.

AIIM Journal 2025 Journal Article

TSFNet: A Temporal–Spectral Fusion Network for advanced speech emotion recognition in medical applications

  • Xinran Li
  • Peilin Huang
  • Xiaojiang Peng
  • Feng Sha
  • Xiaomao Fan
  • Ye Li

Speech emotion recognition (SER) is a critical component in enhancing communication systems and human–machine interaction, with significant potential for applications in the medical field. Although existing SER methods that combine temporal and spectral features have achieved notable advancements, they still encounter a big challenge in capturing emotional nuances, which are vital in medical diagnostics and patient care. In this study, we introduce a straightforward yet highly efficient network called TSFNet, which is the Temporal–Spectral Fusion Network via a Large-scale Pre-trained Model. This network is specifically designed to effectively process intricate emotional nuances by seamlessly integrating temporal and spectral information present in speech signals. By leveraging the capabilities of a large-scale pre-trained model, which serves as a powerful plug-and-play component for extracting and learning the temporal characteristics of speech, TSFNet enables a more accurate capture of complex emotional details crucial for medical applications. Extensive experiments are conducted on publicly available datasets, to evaluate the performance of TSFNet. Extensive experiments conducted on six public datasets demonstrate that TSFNet significantly outperforms existing baselines, achieving unweighted accuracies of 95. 57% for Savee, 92. 67% for Crema-D, 85. 71% for IEMOCAP, 100. 00% for Tess, 95. 86% for Emovo, and 80. 43% for Meld. It means that TSFNet has the potential in advancing medical diagnostic tools and patient monitoring systems.

ICLR Conference 2025 Conference Paper

UniDrive: Towards Universal Driving Perception Across Camera Configurations

  • Ye Li
  • Wenzhao Zheng
  • Xiaonan Huang
  • Kurt Keutzer

Vision-centric autonomous driving has demonstrated excellent performance with economical sensors. As the fundamental step, 3D perception aims to infer 3D information from 2D images based on 3D-2D projection. This makes driving perception models susceptible to sensor configuration (e.g., camera intrinsics and extrinsics) variations. However, generalizing across camera configurations is important for deploying autonomous driving models on different car models. In this paper, we present UniDrive, a novel framework for vision-centric autonomous driving to achieve universal perception across camera configurations. We deploy a set of unified virtual cameras and propose a ground-aware projection method to effectively transform the original images into these unified virtual views. We further propose a virtual configuration optimization method by minimizing the expected projection error between original and virtual cameras. The proposed virtual camera projection can be applied to existing 3D perception methods as a plug-and-play module to mitigate the challenges posed by camera parameter variability, resulting in more adaptable and reliable driving perception models. To evaluate the effectiveness of our framework, we collect a dataset on CARLA by driving the same routes while only modifying the camera configurations. Experimental results demonstrate that our method trained on one specific camera configuration can generalize to varying configurations with minor performance degradation.

IROS Conference 2025 Conference Paper

Using Upper Limb Carrying Exoskeleton with Dual-Model Torque Control Strategy to Reduce Load Impact

  • Daming Liu
  • Ye Li
  • Junchen Liu
  • Ziqi Wang
  • Jie Zhao 0003
  • Yanhe Zhu

Exoskeleton technology holds significant promise within the human-centric paradigm of Industry 5. 0 for mitigating work-related musculoskeletal disorders (WMSDs). However, existing systems often struggle with mismatched assistive torque and inefficient human-machine collaboration under dynamic loading conditions, largely due to insufficient motion intent recognition accuracy. This study proposes a dual-model-based multimodal fusion control strategy that integrates a bidirectional LSTM neural network (Bi-LSTM) with a transformer-based multi-task learning model (MTL) to enable real-time torque compensation and accurate prediction of dynamic load mass under varying conditions. The team developed a lightweight elbow joint exoskeleton prototype, leveraging multi-modal information to enhance assistive torque prediction accuracy. Experimental results show an 83. 7% reduction in agonist muscle activation under a 3. 5 kg load compared to conditions without the exoskeleton, underscoring its potential for industrial material handling scenarios.

EAAI Journal 2025 Journal Article

Vehicle real-time collision risk prediction: A multi-modal learning approach for diverse urban road scenarios based on a large-scale near-crash event dataset

  • Jipu Li
  • Yi He
  • Ye Li
  • Helai Huang
  • Dan Wu
  • Jieling Jin

The effectiveness of vehicle collision avoidance systems depends on the precision of collision risk prediction models. However, current models often neglect the driver's condition, resulting in their poor ability to predict near-crash events triggered by aggressive, fatigued, or distracted driving. Additionally, current models overlook the differences in modality, type, and variability of multi-source data, leading to insufficient feature extraction from input data, which in turn limits the model's prediction accuracy. To address these issues, we developed an end-to-end pre-trained deep framework (PM-Transformer) with a Transformer, consisting of multi-module recurrent convolutional neural networks. The framework includes four modules: (1) pre-trained time series module that extracts spatiotemporal information from traffic time series using one-dimensional convolutional neural network - long short-term memory; (2) pre-trained spectral module that learns visual temporal representations from traffic spectrograms using two-dimensional convolutional neural network - long short-term memory; (3) metadata module for vectorizing traffic metadata; (4) fusion module that semantically integrates features from the three modules using a Transformer. Results show that the proposed model can achieve the same prediction accuracy as other models using only 5 % of their training sample size. Compared to other traditional models, our model improves the accuracy of risk prediction by 6 %, 9 %, and 4 %, respectively, with small sample sizes (0. 5 s, 1 s, and 2 s in advance), while also maintaining the best performance with larger sample sizes. Findings of this study hold significant potential for improving the effectiveness of vehicle collision avoidance systems.

EAAI Journal 2024 Journal Article

A hybrid deep learning framework for conflict prediction of diverse merge scenarios at roundabouts

  • Ye Li
  • Chang Ge
  • Lu Xing
  • Chen Yuan
  • Fei Liu
  • Jieling Jin

The unique traffic situation at roundabouts causes complex interactions between merging vehicles, thereby increasing the likelihood of conflicts. Reliable prediction of conflict risk contributes to active safety improvement, but few studies have investigated the merge risk of roundabouts at a microscopic level. In light of this, this study develops a hybrid deep learning framework for predicting potential conflict risks in complex merging scenarios at roundabouts. Specifically, a roundabout coordinate system is devised to define vehicle characteristics based on trajectory data. Then, an improved 2D-TTC (time-to-collision) indicator is employed to identify two-dimensional merge conflicts. Since the surrounding vehicles may change as vehicles merge into a roundabout, this study analyzes several merging scenarios involving different vehicle groups and conflict durations in order to provide a comprehensive understanding of the conflict mechanism. For these scenarios, a hybrid model consisting of a convolutional neural network (CNN) and a long short-term memory network (LSTM) integrated with the convolutional block attention module (CBAM) is utilized to identify key features. The superiority of the proposed prediction method is demonstrated in comparisons with benchmark models. Results showed that segmental predictions were more accurate than overall predictions in terms of conflict duration. Furthermore, it is possible that a specific vehicle group has a decisive effect on the merging conflict risk, as indicated by the fact that information from multiple vehicle groups does not significantly improve the prediction performance. Another finding is that the driving state of vehicles merging at the roundabout varies considerably, but rarely with consecutive or multiple changes. The study provides novel insights into roundabout conflict prediction, which could serve as a tool for enhancing safety management involving complex traffic scenarios.

EAAI Journal 2024 Journal Article

An adaptive lightweight small object detection method for incremental few-shot scenarios of unmanned surface vehicles

  • Bo Wang
  • Peng Jiang
  • Zhuoyan Liu
  • Yueming Li
  • Jian Cao
  • Ye Li

Real-time and accurate detection of sea surface objects has become an important research topic for unmanned surface vehicles (USVs). During the execution of tasks, USVs sometimes need to upgrade their model to detect new categories, and the initial data is often limited. This requires quick adaptation to new categories under few-shot scenarios. We propose a lightweight neural network for detecting small sea surface objects named Shuffle-High-Resolution-Net (SHRDet), which integrates the enhanced Shuffle Block based on High-Resolution-Net (HRNet), lightweight feature fusion module, and Focal Efficient Intersection over Union loss. Based on SHRDet, a fast adaptation method named SHRDet-N for incremental few-shot categories is proposed. It generates category enhancement features through a cross-attention mechanism, and introduces elastic weight consolidation and feature distance to solve catastrophic forgetting when learning incremental few-shot categories. The algorithms have been applied to an intelligent USV platform for various surface missions, such as security patrol, ocean investigation, and marine engineering. The experimental results on public datasets indicate that SHRDet achieves 80. 7 % mean Average Precision (mAP) on the Water Surface Object Detection Dataset (WSODD) with only 0. 69 M parameters and less calculation quantity, and SHRDet is significantly superior to state-of-the-art methods in terms of lightweight and accuracy. Moreover, SHRDet-N effectively solves the learning problem of incremental few-shot categories. When the sample number of a new category is set to 20, the mAP of the base categories is 82. 5 % and that of the new category is 63. 8 %, which is 2. 8 % and 4. 2 % higher than that of state-of-the-art models like Sylph.

IROS Conference 2024 Conference Paper

AVM-SLAM: Semantic Visual SLAM with Multi-Sensor Fusion in a Bird's Eye View for Automated Valet Parking

  • Ye Li
  • Wenchao Yang
  • Dekun Lin
  • Qianlei Wang
  • Zhe Cui
  • Xiaolin Qin

Accurate localization in challenging garage environments—marked by poor lighting, sparse textures, repetitive structures, dynamic scenes, and the absence of GPS—is crucial for automated valet parking (AVP) tasks. Addressing these challenges, our research introduces AVM-SLAM, a cutting-edge semantic visual SLAM architecture with multi-sensor fusion in a bird’s eye view (BEV). This novel framework synergizes the capabilities of four fisheye cameras, wheel encoders, and an inertial measurement unit (IMU) to construct a robust SLAM system. Unique to our approach is the implementation of a flare removal technique within the BEV imagery, significantly enhancing road marking detection and semantic feature extraction by convolutional neural networks for superior mapping and localization. Our work also pioneers a semantic prequalification (SPQ) module, designed to adeptly handle the challenges posed by environments with repetitive textures, thereby enhancing loop detection and system robustness. To demonstrate the effectiveness and resilience of AVM-SLAM, we have released a specialized multi-sensor and high-resolution dataset of an underground garage, accessible at https://yale-cv.github.io/avm-slamdataset, encouraging further exploration and validation of our approach within similar settings.

IJCAI Conference 2024 Conference Paper

Causality-enhanced Discreted Physics-informed Neural Networks for Predicting Evolutionary Equations

  • Ye Li
  • Siqi Chen
  • Bin Shan
  • Sheng-Jun Huang

Physics-informed neural networks (PINNs) have shown promising potential for solving partial differential equations (PDEs) using deep learning. However, PINNs face training difficulties for evolutionary PDEs, particularly for dynamical systems whose solutions exhibit multi-scale or turbulent behavior over time. The reason is that PINNs may violate the temporal causality property since all the temporal features in the PINNs loss are trained simultaneously. This paper proposes to use implicit time differencing schemes to enforce temporal causality, and use transfer learning to sequentially update the PINNs in space as surrogates for PDE solutions in different time frames. The evolving PINNs are better able to capture the varying complexities of the evolutionary equations, while only requiring minor updates between adjacent time frames. Our method is theoretically proven to be convergent if the time step is small and each PINN in different time frames is well-trained. In addition, we provide state-of-the-art (SOTA) numerical results for a variety of benchmarks for which existing PINNs formulations may fail or be inefficient. We demonstrate that the proposed method improves the accuracy of PINNs approximation for evolutionary PDEs and improves efficiency by a factor of 4–40x. The code is available at https: //github. com/SiqiChen9/TL-DPINNs.

AAAI Conference 2024 Conference Paper

Component Fourier Neural Operator for Singularly Perturbed Differential Equations

  • Ye Li
  • Ting Du
  • Yiwen Pang
  • Zhongyi Huang

Solving Singularly Perturbed Differential Equations (SPDEs) poses computational challenges arising from the rapid transitions in their solutions within thin regions. The effectiveness of deep learning in addressing differential equations motivates us to employ these methods for solving SPDEs. In this paper, we introduce Component Fourier Neural Operator (ComFNO), an innovative operator learning method that builds upon Fourier Neural Operator (FNO), while simultaneously incorporating valuable prior knowledge obtained from asymptotic analysis. Our approach is not limited to FNO and can be applied to other neural network frameworks, such as Deep Operator Network (DeepONet), leading to potential similar SPDEs solvers. Experimental results across diverse classes of SPDEs demonstrate that ComFNO significantly improves accuracy compared to vanilla FNO. Furthermore, ComFNO exhibits natural adaptability to diverse data distributions and performs well in few-shot scenarios, showcasing its excellent generalization ability in practical situations.

JBHI Journal 2024 Journal Article

HGCTNet: Handcrafted Feature-Guided CNN and Transformer Network for Wearable Cuffless Blood Pressure Measurement

  • Zeng-Ding Liu
  • Ye Li
  • Yuan-Ting Zhang
  • Jia Zeng
  • Zu-Xian Chen
  • Ji-Kui Liu
  • Fen Miao

Biosignals collected by wearable devices, such as electrocardiogram and photoplethysmogram, exhibit redundancy and global temporal dependencies, posing a challenge in extracting discriminative features for blood pressure (BP) estimation. To address this challenge, we propose HGCTNet, a handcrafted feature-guided CNN and transformer network for cuffless BP measurement based on wearable devices. By leveraging convolutional operations and self-attention mechanisms, we design a CNN-Transformer hybrid architecture to learn features from biosignals that capture both local information and global temporal dependencies. Then, we introduce a handcrafted feature-guided attention module that utilizes handcrafted features extracted from biosignals as query vectors to eliminate redundant information within the learned features. Finally, we design a feature fusion module that integrates the learned features, handcrafted features, and demographics to enhance model performance. We validate our approach using two large wearable BP datasets: the CAS-BP dataset and the Aurora-BP dataset. Experimental results demonstrate that HGCTNet achieves an estimation error of 0. 9 $\pm$ 6. 5 mmHg for diastolic BP (DBP) and 0. 7 $\pm$ 8. 3 mmHg for systolic BP (SBP) on the CAS-BP dataset. On the Aurora-BP dataset, the corresponding errors are $-$ 0. 4 $\pm$ 7. 0 mmHg for DBP and $-$ 0. 4 $\pm$ 8. 6 mmHg for SBP. Compared to the current state-of-the-art approaches, HGCTNet reduces the mean absolute error of SBP estimation by 10. 68% on the CAS-BP dataset and 9. 84% on the Aurora-BP dataset. These results highlight the potential of HGCTNet in improving the performance of wearable cuffless BP measurements.

ICRA Conference 2024 Conference Paper

Influence of Camera-LiDAR Configuration on 3D Object Detection for Autonomous Driving

  • Ye Li
  • Hanjiang Hu
  • Zuxin Liu
  • Xiaohao Xu
  • Xiaonan Huang
  • Ding Zhao

Cameras and LiDARs are both important sensors for autonomous driving, playing critical roles in 3D object detection. Camera-LiDAR Fusion has been a prevalent solution for robust and accurate driving perception. In contrast to the vast majority of existing arts that focus on how to improve the performance of 3D target detection through cross-modal schemes, deep learning algorithms, and training tricks, we devote attention to the impact of sensor configurations on the performance of learning-based methods. To achieve this, we propose a unified information-theoretic surrogate metric for camera and LiDAR evaluation based on the proposed sensor perception model. We also design an accelerated high-quality framework for data acquisition, model training, and performance evaluation that functions with the CARLA simulator. To show the correlation between detection performance and our surrogate metrics, We conduct experiments using several camera-LiDAR placements and parameters inspired by selfdriving companies and research institutions. Extensive experimental results of representative algorithms on nuScenes dataset validate the effectiveness of our surrogate metric, demonstrating that sensor configurations significantly impact point-cloudimage fusion based detection models, which contribute up to 30% discrepancy in terms of the average precision.

NeurIPS Conference 2024 Conference Paper

Is Your LiDAR Placement Optimized for 3D Scene Understanding?

  • Ye Li
  • Lingdong Kong
  • Hanjiang Hu
  • Xiaohao Xu
  • Xiaonan Huang

The reliability of driving perception systems under unprecedented conditions is crucial for practical usage. Latest advancements have prompted increasing interest in multi-LiDAR perception. However, prevailing driving datasets predominantly utilize single-LiDAR systems and collect data devoid of adverse conditions, failing to capture the complexities of real-world environments accurately. Addressing these gaps, we proposed Place3D, a full-cycle pipeline that encompasses LiDAR placement optimization, data generation, and downstream evaluations. Our framework makes three appealing contributions. 1) To identify the most effective configurations for multi-LiDAR systems, we introduce the Surrogate Metric of the Semantic Occupancy Grids (M-SOG) to evaluate LiDAR placement quality. 2) Leveraging the M-SOG metric, we propose a novel optimization strategy to refine multi-LiDAR placements. 3) Centered around the theme of multi-condition multi-LiDAR perception, we collect a 280, 000-frame dataset from both clean and adverse conditions. Extensive experiments demonstrate that LiDAR placements optimized using our approach outperform various baselines. We showcase exceptional results in both LiDAR semantic segmentation and 3D object detection tasks, under diverse weather and sensor failure conditions.

EAAI Journal 2023 Journal Article

Contrastive knowledge integrated graph neural networks for Chinese medical text classification

  • Ge Lan
  • Mengting Hu
  • Ye Li
  • Yuzhi Zhang

This paper aims at medical text classification, where texts describe medicines, diseases, or other medical topics. This field is still challenging since medical texts contain intensive specialization and terminology, which require professional semantic and structured knowledge to classify. Based on our observations, medical knowledge graph (KG) can provide such knowledge although they may be ambiguous. To this end, we propose contrastive knowledge integrated graph neural networks (ConKGNN) to make full use of the above knowledge. Specifically, the proposed method builds two graphs for a medical text, i. e. text graph and text-specific subgraph, containing the text information and relevant KG information, respectively. Two graphs are merged into a united graph, which is jointly modeled by graph neural networks (GNN). In this way, our approach adequately learns interactions between neighbors. Meanwhile, it promotes the mutual influences between text and KG. We further propose graph-based supervised contrastive learning. By randomly cutting off nodes from the text graph, an augmented united graph is obtained. Learning it in a contrastive way could enhance the robustness of introducing KG information. Comprehensive experiments are conducted on five Chinese medical datasets and experimental results show our model outperforms strong baselines remarkably. Consequently, our model can serve as an efficient medical text classifier with excellent performance. We release the code at https: //github. com/nolongernome/ConKGNN.

JBHI Journal 2023 Journal Article

Cuffless Blood Pressure Measurement Using Smartwatches: A Large-Scale Validation Study

  • Zeng-Ding Liu
  • Ye Li
  • Yuan-Ting Zhang
  • Jia Zeng
  • Zu-Xian Chen
  • Zhi-Wei Cui
  • Ji-Kui Liu
  • Fen Miao

This study aimed to evaluate the performance of cuffless blood pressure (BP) measurement techniques in a large and diverse cohort of participants. We enrolled 3077 participants (aged 18–75, 65. 16% women, 35. 91% hypertensive participants) and conducted followed-up for approximately 1 month. Electrocardiogram, pulse pressure wave, and multiwavelength photoplethysmogram signals were simultaneously recorded using smartwatches; dual-observer auscultation systolic BP (SBP) and diastolic BP (DBP) reference measurements were also obtained. Pulse transit time, traditional machine learning (TML), and deep learning (DL) models were evaluated with calibration and calibration-free strategy. TML models were developed using ridge regression, support vector machine, adaptive boosting, and random forest; while DL models using convolutional and recurrent neural networks. The best-performing calibration-based model yielded estimation errors of 1. 33 $\pm$ 6. 43 mmHg for DBP and 2. 31 $\pm$ 9. 57 mmHg for SBP in the overall population, with reduced SBP estimation errors in normotensive (1. 97 $\pm$ 7. 85 mmHg) and young (0. 24 $\pm$ 6. 61 mmHg) subpopulations. The best-performing calibration-free model had estimation errors of $-$ 0. 29 $\pm$ 8. 78 mmHg for DBP and $-$ 0. 71 $\pm$ 13. 04 mmHg for SBP. We conclude that smartwatches are effective for measuring DBP for all participants and SBP for normotensive and younger participants with calibration; performance degrades significantly for heterogeneous populations including older and hypertensive participants. The availability of cuffless BP measurement without calibration is limited in routine settings. Our study provides a large-scale benchmark for emerging investigations on cuffless BP measurement, highlighting the need to explore additional signals or principles to enhance the accuracy in large-scale heterogeneous populations.

AAAI Conference 2023 Conference Paper

Implicit Stochastic Gradient Descent for Training Physics-Informed Neural Networks

  • Ye Li
  • Song-Can Chen
  • Sheng-Jun Huang

Physics-informed neural networks (PINNs) have effectively been demonstrated in solving forward and inverse differential equation problems, but they are still trapped in training failures when the target functions to be approximated exhibit high-frequency or multi-scale features. In this paper, we propose to employ implicit stochastic gradient descent (ISGD) method to train PINNs for improving the stability of training process. We heuristically analyze how ISGD overcome stiffness in the gradient flow dynamics of PINNs, especially for problems with multi-scale solutions. We theoretically prove that for two-layer fully connected neural networks with large hidden nodes, randomly initialized ISGD converges to a globally optimal solution for the quadratic loss function. Empirical results demonstrate that ISGD works well in practice and compares favorably to other gradient-based optimization methods such as SGD and Adam, while can also effectively address the numerical stiffness in training dynamics via gradient descent.

EAAI Journal 2022 Journal Article

Instance segmentation of biological images using graph convolutional network

  • Rongtao Xu
  • Ye Li
  • Changwei Wang
  • Shibiao Xu
  • Weiliang Meng
  • Xiaopeng Zhang

Instance segmentation in biological images is an important task in the field of biological images and biomedical analysis. Different from the instance segmentation of natural image scenes, this task is still challenging because there are a large number of overlapping objects with similar appearance as well as great variability in shape, size and texture in the foreground and background. In this paper, we propose a novel method for segmentation of graph-guided instances of biological images, which successfully addresses these peculiarities. Our method predicts the embedding at each pixel and uses clustering to recover instances during testing. Specifically, we design the Graph-guided Feature Fusion Module in response to overlapping instances. Our Graph-guided Feature Fusion Module combines fine deep features and coarse shallow features to learn the affinity matrix, and then uses graph convolutional network to guide the network to learn object-level local features. Next, we devise the Gated Spatial Attention Module to effectively learn key spatial information by introducing a gating mechanism. Furthermore, we give the Cluster Distance Loss that can effectively distinguish foreground objects from similar backgrounds. The effectiveness of our proposed method has been verified on various biological and biomedical datasets. The experimental results show that our method is superior to previous embedding-based instance segmentation methods. The SBD metric for our method reached 90. 8% on the plant phenotype dataset (CVPPP), 72. 5% on the cell nucleus dataset (DSB2018), and 81. 8% on the C. elegans dataset, all achieving state-of-the-art performance.

JBHI Journal 2022 Journal Article

Non-Contact Heartbeat Detection Based on Ballistocardiogram Using UNet and Bidirectional Long Short-Term Memory

  • Yaozong Mai
  • Zizhao Chen
  • Baoxian Yu
  • Ye Li
  • Zhiqiang Pang
  • Zhang Han

Benefiting from non-invasive sensing tech- nologies, heartbeat detection from ballistocardiogram (BCG) signals is of great significance for home-care applications, such as risk prediction of cardiovascular disease (CVD) and sleep staging, etc. In this paper, we propose an effective deep learning model for automatic heartbeat detection from BCG signals based on UNet and bidirectional long short-term memory (Bi-LSTM). The developed deep learning model provides an effective solution to the existing challenges in BCG-aided heartbeat detection, especially for BCG in low signal-to-noise ratio, in which the waveforms in BCG signals are irregular due to measured postures, rhythm and artifact motion. For validations, performance of the proposed detection is evaluated by BCG recordings from 43 subjects with different measured postures and heart rate ranges. The accuracy of the detected heartbeat intervals measured in different postures and signal qualities, in comparison with the R-R interval of ECG, is promising in terms of mean absolute error and mean relative error, respectively, which is superior to the state-of-the-art methods. Numerical results demonstrate that the proposed UNet-BiLSTM model performs robust to noise and perturbations (e. g. respiratory effort and artifact motion) in BCG signals, and provides a reliable solution to long term heart rate monitoring.

AIIM Journal 2022 Journal Article

Predicting disease progress with imprecise lab test results

  • Mei Wang
  • Zhihua Lin
  • Ruihua Li
  • Ye Li
  • Jianwen Su

Clinical lab tests play an important role in disease diagnose and medical treatment, the test results have been utilized widely in predictive modeling tasks in healthcare. However, in most existing works, the loss function implicitly assumes that the value of the sample used to be predicted is the only correct one. This assumption fails to hold for lab test data, which usually are within respective tolerable ranges or imprecision ranges. In addition, the historical lab test data is always organized based on their sequential position, the timestamps between the data are often neglected. In this paper, we study the issue of building robust models while simultaneously taking imprecision and timestamp of the data into account with better generalization. In particular, “IR loss” is proposed in which each data in imprecision range space has a certain probability to be the real value, participating in the loss calculation. The loss is then defined as the integral of the error of each point in the impression range space. The sampling and discretization methods are proposed for loss calculation. A heuristic learning algorithm is developed to learn the model parameters. We further apply IR loss for disease progress prediction while the input data is organized as sequence. We reformulate the prediction task with timestamp based on Long Short-Term Memory (LSTM) network. At the same time, the timestamp is readily combined with the proposed IR loss to avoid the change of predicted result caused by the change of the test values in small time range. We conducted the experiments based on two real world datasets. Experimental results show that the prediction method based on IR loss can provide more accurate prediction result for different kinds of task and diverse learning methods. Our method can also provide more stable and consistent results when test samples are generated from imprecision range and small time range.

NeurIPS Conference 2022 Conference Paper

Redundancy-Free Message Passing for Graph Neural Networks

  • Rongqin Chen
  • Shenghui Zhang
  • Leong Hou U
  • Ye Li

Graph Neural Networks (GNNs) resemble the Weisfeiler-Lehman (1-WL) test, which iteratively update the representation of each node by aggregating information from WL-tree. However, despite the computational superiority of the iterative aggregation scheme, it introduces redundant message flows to encode nodes. We found that the redundancy in message passing prevented conventional GNNs from propagating the information of long-length paths and learning graph similarities. In order to address this issue, we proposed Redundancy-Free Graph Neural Network (RFGNN), in which the information of each path (of limited length) in the original graph is propagated along a single message flow. Our rigorous theoretical analysis demonstrates the following advantages of RFGNN: (1) RFGNN is strictly more powerful than 1-WL; (2) RFGNN efficiently propagate structural information in original graphs, avoiding the over-squashing issue; and (3) RFGNN could capture subgraphs at multiple levels of granularity, and are more likely to encode graphs with closer graph edit distances into more similar representations. The experimental evaluation of graph-level prediction benchmarks confirmed our theoretical assertions, and the performance of the RFGNN can achieve the best results in most datasets.

ICLR Conference 2021 Conference Paper

Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval

  • Lee Xiong
  • Chenyan Xiong
  • Ye Li
  • Kwok-Fung Tang
  • Jialin Liu
  • Paul N. Bennett
  • Junaid Ahmed
  • Arnold Overwijk

Conducting text retrieval in a learned dense representation space has many intriguing advantages. Yet dense retrieval (DR) often underperforms word-based sparse retrieval. In this paper, we first theoretically show the bottleneck of dense retrieval is the domination of uninformative negatives sampled in mini-batch training, which yield diminishing gradient norms, large gradient variances, and slow convergence. We then propose Approximate nearest neighbor Negative Contrastive Learning (ANCE), which selects hard training negatives globally from the entire corpus. Our experiments demonstrate the effectiveness of ANCE on web search, question answering, and in a commercial search engine, showing ANCE dot-product retrieval nearly matches the accuracy of BERT-based cascade IR pipeline. We also empirically validate our theory that negative sampling with ANCE better approximates the oracle importance sampling procedure and improves learning convergence.

AIIM Journal 2020 Journal Article

Continuous blood pressure measurement from one-channel electrocardiogram signal using deep-learning techniques

  • Fen Miao
  • Bo Wen
  • Zhejing Hu
  • Giancarlo Fortino
  • Xi-Ping Wang
  • Zeng-Ding Liu
  • Min Tang
  • Ye Li

Continuous blood pressure (BP) measurement is crucial for reliable and timely hypertension detection. State-of-the-art continuous BP measurement methods based on pulse transit time or multiple parameters require simultaneous electrocardiogram (ECG) and photoplethysmogram (PPG) signals. Compared with PPG signals, ECG signals are easy to collect using wearable devices. This study examined a novel continuous BP estimation approach using one-channel ECG signals for unobtrusive BP monitoring. A BP model is developed based on the fusion of a residual network and long short-term memory to obtain the spatial-temporal information of ECG signals. The public multiparameter intelligent monitoring waveform database, which contains ECG, PPG, and invasive BP data of patients in intensive care units, is used to develop and verify the model. Experimental results demonstrated that the proposed approach exhibited an estimation error of 0. 07 ± 7. 77 mmHg for mean arterial pressure (MAP) and 0. 01 ± 6. 29 for diastolic BP (DBP), which comply with the Association for the Advancement of Medical Instrumentation standard. According to the British Hypertension Society standards, the results achieved grade A for MAP and DBP estimation and grade B for systolic BP (SBP) estimation. Furthermore, we verified the model with an independent dataset for arrhythmia patients. The experimental results exhibited an estimation error of −0. 22 ± 5. 82 mmHg, −0. 57 ± 4. 39 mmHg, and −0. 75 ± 5. 62 mmHg for SBP, MAP, and DBP measurements, respectively. These results indicate the feasibility of estimating BP by using a one-channel ECG signal, thus enabling continuous BP measurement for ubiquitous health care applications.

JBHI Journal 2020 Journal Article

Deep Multi-Scale Fusion Neural Network for Multi-Class Arrhythmia Detection

  • Ruxin Wang
  • Jianping Fan
  • Ye Li

Automated electrocardiogram (ECG) analysis for arrhythmia detection plays a critical role in early prevention and diagnosis of cardiovascular diseases. Extracting powerful features from raw ECG signals for fine-grained diseases classification is still a challenging problem today due to variable abnormal rhythms and noise distribution. For ECG analysis, the previous research works depend mostly on heartbeat or single scale signal segments, which ignores underlying complementary information of different scales. In this paper, we formulate a novel end-to-end Deep Multi-Scale Fusion convolutional neural network (DMSFNet) architecture for multi-class arrhythmia detection. Our proposed approach can effectively capture abnormal patterns of diseases and suppress noise interference by multi-scale feature extraction and cross-scale information complementarity of ECG signals. The proposed method implements feature extraction for signal segments with different sizes by integrating multiple convolution kernels with different receptive fields. Meanwhile, joint optimization strategy with multiple losses of different scales is designed, which not only learns scale-specific features, but also realizes cumulatively multi-scale complementary feature learning during the learning process. In our work, we demonstrate our DMSFNet on two open datasets (CPSC_2018 and PhysioNet/CinC_2017) and deliver the state-of-art performance on them. Among them, CPSC_2018 is a 12-lead ECG dataset and CinC_2017 is a single-lead dataset. For these two datasets, we achieve the F1 score $\text{82. 8}\%$ and $\text{84. 1}\%$ which are higher than previous state-of-art approaches respectively. The results demonstrate that our end-to-end DMSFNet has outstanding performance for feature extraction from a broad range of distinct arrhythmias and elegant generalization ability for effectively handling ECG signals with different leads.

JBHI Journal 2020 Journal Article

Multi-Sensor Fusion Approach for Cuff-Less Blood Pressure Measurement

  • Fen Miao
  • Zeng-Ding Liu
  • Ji-Kui Liu
  • Bo Wen
  • Qing-Yun He
  • Ye Li

Ambulatory blood pressure (BP) provides valuable information for cardiovascular risk assessment. The present cuff-based devices are intrusive for longterm BP monitoring, whereas cuff-less BP measurement methods based on pulse transit time or multi-parameter are inferior in robustness and reliability by using electrocardiogram (ECG) and photoplethysmogram signals. This study examined a multi-sensor fusion-based platform and algorithm for systolic BP (SBP), mean arterial pressure (MAP), and diastolic BP (DBP) estimation. The proposed multisensor platform was comprised of one ECG sensor and two pulse pressure wave sensors for simultaneous signal collection. After extracting 35 features from the collected signals, a weakly supervised feature selection method was proposed for dimension reduction because the reference oscillometric technique-based BP are intermittent and can be redeemed as coarse-grained labels. BP models were then established using a multi-instance regression algorithm. A total of 85 participants including 17 hypertensive and 12 hypotensive patients were enrolled. Experimental results showed that the proposed approach exhibited good accuracy for diverse population with an estimation error of 1. 62 ± 7. 76 mmHg for SBP, 1. 53 ± 6. 03 mmHg for MAP, and 1. 49 ± 5. 52 for DBP, which complied with the association for the advancement of medical instrumentation standards in BP estimation. Moreover, the estimation accuracy is with random daily fluctuations rather than long-term degradation through a maximum two-month follow-up period indicated good robustness performance. These results suggest that the proposed approach is with high reliability and robustness and thus provides a novel insight for cuff-less BP measurement.

AAAI Conference 2018 Conference Paper

Community Detection in Attributed Graphs: An Embedding Approach

  • Ye Li
  • Chaofeng Sha
  • Xin Huang
  • Yanchun Zhang

Community detection is a fundamental and widely-studied problem that finds all densely-connected groups of nodes and well separates them from others in graphs. With the proliferation of rich information available for entities in real-world networks, it is useful to discover communities in attributed graphs where nodes tend to have attributes. However, most existing attributed community detection methods directly utilize the original network topology leading to poor results due to ignoring inherent community structures. In this paper, we propose a novel embedding based model to discover communities in attributed graphs. Specifically, based on the observation of densely-connected structures in communities, we develop a novel community structure embedding method to encode inherent community structures via underlying community memberships. Based on node attributes and community structure embedding, we formulate the attributed community detection as a nonnegative matrix factorization optimization problem. Moreover, we carefully design iterative updating rules to make sure of finding a converging solution. Extensive experiments conducted on 19 attributed graph datasets with overlapping and non-overlapping ground-truth communities show that our proposed model CDE can accurately identify attributed communities and significantly outperform 7 stateof-the-art methods.

JBHI Journal 2018 Journal Article

Gait-Cycle-Driven Transmission Power Control Scheme for a Wireless Body Area Network

  • Weilin Zang
  • Ye Li

In a wireless body area network (WBAN), walking movements can result in rapid channel fluctuations, which severely degrade the performance of transmission power control (TPC) schemes. On the other hand, these channel fluctuations are often periodic and are time-synchronized with the user's gait cycle, since they are all driven from the walking movements. In this paper, we propose a novel gait-cycle-driven transmission power control (G-TPC) for a WBAN. The proposed G-TPC scheme reinforces the existing TPC scheme by exploiting the periodic channel fluctuation in the walking scenario. In the proposed scheme, the user's gait cycle information acquired by an accelerometer is used as beacons for arranging the transmissions at the time points with the ideal channel state. The specific transmission power is then determined by using received signal strength indication (RSSI). An experiment was conducted to evaluate the energy efficiency and reliability of the proposed G-TPC based on a CC2420 platform. The results reveal that compared to the original RSSI/link-quality-indication-based TPC, G-TPC reduces energy consumption by 25% on the sensor node and reduce the packet loss rate by 65%.

JBHI Journal 2018 Journal Article

Multiscaled Fusion of Deep Convolutional Neural Networks for Screening Atrial Fibrillation From Single Lead Short ECG Recordings

  • Xiaomao Fan
  • Qihang Yao
  • Yunpeng Cai
  • Fen Miao
  • Fangmin Sun
  • Ye Li

Atrial fibrillation (AF) is one of the most common sustained chronic cardiac arrhythmia in elderly population, associated with a high mortality and morbidity in stroke, heart failure, coronary artery disease, systemic thromboembolism, etc. The early detection of AF is necessary for averting the possibility of disability or mortality. However, AF detection remains problematic due to its episodic pattern. In this paper, a multiscaled fusion of deep convolutional neural network (MS-CNN) is proposed to screen out AF recordings from single lead short electrocardiogram (ECG) recordings. The MS-CNN employs the architecture of two-stream convolutional networks with different filter sizes to capture features of different scales. The experimental results show that the proposed MS-CNN achieves 96. 99% of classification accuracy on ECG recordings cropped/padded to 5 s. Especially, the best classification accuracy, 98. 13%, is obtained on ECG recordings of 20 s. Compared with artificial neural network, shallow single-stream CNN, and VisualGeometry group network, the MS-CNN can achieve the better classification performance. Meanwhile, visualization of the learned features from the MS-CNN demonstrates its superiority in extracting linear separable ECG features without hand-craft feature engineering. The excellent AF screening performance of the MS-CNN can satisfy the most elders for daily monitoring with wearable devices.

JBHI Journal 2017 Journal Article

A Novel Continuous Blood Pressure Estimation Approach Based on Data Mining Techniques

  • Fen Miao
  • Nan Fu
  • Yuan-Ting Zhang
  • Xiao-Rong Ding
  • Xi Hong
  • Qingyun He
  • Ye Li

Continuous blood pressure (BP) estimation using pulse transit time (PTT) is a promising method for unobtrusive BP measurement. However, the accuracy of this approach must be improved for it to be viable for a wide range of applications. This study proposes a novel continuous BP estimation approach that combines data mining techniques with a traditional mechanism-driven model. First, 14 features derived from simultaneous electrocardiogram and photoplethysmogram signals were extracted for beat-to-beat BP estimation. A genetic algorithm-based feature selection method was then used to select BP indicators for each subject. Multivariate linear regression and support vector regression were employed to develop the BP model. The accuracy and robustness of the proposed approach were validated for static, dynamic, and follow-up performance. Experimental results based on 73 subjects showed that the proposed approach exhibited excellent accuracy in static BP estimation, with a correlation coefficient and mean error of 0. 852 and −0. 001 ± 3. 102 mmHg for systolic BP, and 0. 790 and −0. 004 ± 2. 199 mmHg for diastolic BP. Similar performance was observed for dynamic BP estimation. The robustness results indicated that the estimation accuracy was lower by a certain degree one day after model construction but was relatively stable from one day to six months after construction. The proposed approach is superior to the state-of-the-art PTT-based model for an approximately 2-mmHg reduction in the standard derivation at different time intervals, thus providing potentially novel insights for cuffless BP estimation.

JBHI Journal 2016 Journal Article

Continuous Blood Pressure Measurement From Invasive to Unobtrusive: Celebration of 200th Birth Anniversary of Carl Ludwig

  • Xiao-Rong Ding
  • Ni Zhao
  • Guang-Zhong Yang
  • Roderic I. Pettigrew
  • Benny Lo
  • Fen Miao
  • Ye Li
  • Jing Liu

The year 2016 marks the 200th birth anniversary of Carl Friedrich Wilhelm Ludwig (1816-1895). As one of the most remarkable scientists, Ludwig invented the kymograph, which for the first time enabled the recording of continuous blood pressure (BP), opening the door to the modern study of physiology. Almost a century later, intraarterial BP monitoring through an arterial line has been used clinically. Subsequently, arterial tonometry and volume clamp method were developed and applied in continuous BP measurement in a noninvasive way. In the last two decades, additional efforts have been made to transform the method of unobtrusive continuous BP monitoring without the use of a cuff. This review summarizes the key milestones in continuous BP measurement; that is, kymograph, intraarterial BP monitoring, arterial tonometry, volume clamp method, and cuffless BP technologies. Our emphasis is on recent studies of unobtrusive BP measurements as well as on challenges and future directions.

IROS Conference 2006 Conference Paper

Motion Control of Underwater Vehicles Based on Robust Neural Network

  • Xiao Liang 0003
  • Ye Li
  • Lei Wan

Aiming at low response speed and sensitization to external disturbance in motion control of underwater vehicles by adopting neural network, a stable robust learning algorithm was presented based on variable structure control theory and error back propagation algorithm, and the global stability conditions were discussed in detail. Finally, simulation experiments were carried out on general detection remotely operated vehicle. The results show that it has good robustness to external noises and changing of learning-ratio, which reduces the abrasion of the mechanically-driven system greatly. It keeps learning of neural network fast and stable, which meets the requirement of real-time control and has theoretical and practical value

v2026.09.13