Arrow Research search

Author name cluster

Jiahui Pan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

JBHI Journal 2026 Journal Article

A Multi-Scale Attention-based Reconstruction Fusion Network for Motor Imagery Classification

  • Lina Qiu
  • You Hu
  • Minjin Wu
  • Baiqiang Long
  • Tianjian Chen
  • Jiahui Pan

Motor imagery (MI) is a widely used cognitive paradigm in brain-computer interface (BCI) systems, where accurate and efficient MI decoding is essential for real-time human-machine interaction. However, the non-stationary nature and pronounced inter-subject variability of electroencephalography (EEG) signals pose significant challenges to reliable decoding. To address these issues, we propose a multi-scale attention-based reconstruction fusion network (MSARFNet) for MI-EEG decoding. The proposed framework employs parallel multi-scale convolutional branches to extract discriminative spatio-temporal features at different temporal resolutions. An attention-based reconstruction fusion module is then introduced to selectively diminish non-dominant information while promoting effective interaction among multi-scale features. Furthermore, a local-global temporal encoding strategy is designed to enhance transient MI-related responses through local temporal context aggregation and subsequently capture long-range temporal dependencies via global temporal modeling. Subject-dependent experiments conducted on the BCI Competition IV 2a and 2b datasets demonstrate that MSARFNet achieves average classification accuracies of 84. 64% and 87. 96%, respectively, outperforming several state-of-the-art methods. These results indicate that MSARFNet provides an effective and robust solution for EEG-based MI decoding.

JBHI Journal 2025 Journal Article

A Multimodal Consistency-Based Self-Supervised Contrastive Learning Framework for Automated Sleep Staging in Patients With Disorders of Consciousness

  • Jiahui Pan
  • Yangzuyi Yu
  • Man Li
  • Wanxin Wei
  • Shuyu Chen
  • Heyi Zheng
  • Yanbin He
  • Yuanqing Li

Sleep is a fundamental human activity, and automated sleep staging holds considerable investigational potential. Despite numerous deep learning methods proposed for sleep staging that exhibit notable performance, several challenges remain unresolved, including inadequate representation and generalization capabilities, limitations in multimodal feature extraction, the scarcity of labeled data, and the restricted practical application for patients with disorder of consciousness (DOC). This paper proposes MultiConsSleepNet, a multimodal consistency-based sleep staging network. This network comprises a unimodal feature extractor and a multimodal consistency feature extractor, aiming to explore universal representations of electroencephalograms (EEGs) and electrooculograms (EOGs) and extract the consistency of intra- and intermodal features. Additionally, self-supervised contrastive learning strategies are designed for unimodal and multimodal consistency learning to address the current situation in clinical practice where it is difficult to obtain high-quality labeled data but has a huge amount of unlabeled data. It can effectively alleviate the model's dependence on labeled data, and improve the model's generalizability for effective migration to DOC patients. Experimental results on three publicly available datasets demonstrate that MultiConsSleepNet achieves state-of-the-art performance in sleep staging with limited labeled data and effectively utilizes unlabeled data, enhancing its practical applicability. Furthermore, the proposed model yields promising results on a self-collected DOC dataset, offering a novel perspective for sleep staging research in patients with DOC.

ICLR Conference 2025 Conference Paper

Reading Your Heart: Learning ECG Words and Sentences via Pre-training ECG Language Model

  • Jiarui Jin
  • Haoyu Wang
  • Hongyan Li 0002
  • Jun Li
  • Jiahui Pan
  • Shenda Hong

Electrocardiogram (ECG) is essential for the clinical diagnosis of arrhythmias and other heart diseases, but deep learning methods based on ECG often face limitations due to the need for high-quality annotations. Although previous ECG self-supervised learning (eSSL) methods have made significant progress in representation learning from unannotated ECG data, they typically treat ECG signals as ordinary time-series data, segmenting the signals using fixed-size and fixed-step time windows, which often ignore the form and rhythm characteristics and latent semantic relationships in ECG signals. In this work, we introduce a novel perspective on ECG signals, treating heartbeats as words and rhythms as sentences. Based on this perspective, we first designed the QRS-Tokenizer, which generates semantically meaningful ECG sentences from the raw ECG signals. Building on these, we then propose HeartLang, a novel self-supervised learning framework for ECG language processing, learning general representations at form and rhythm levels. Additionally, we construct the largest heartbeat-based ECG vocabulary to date, which will further advance the development of ECG language processing. We evaluated HeartLang across six public ECG datasets, where it demonstrated robust competitiveness against other eSSL methods. Our data and code are publicly available at https://github.com/PKUDigitalHealth/HeartLang.

IJCAI Conference 2024 Conference Paper

A Density-driven Iterative Prototype Optimization for Transductive Few-shot Learning

  • Jingcong Li
  • Chunjin Ye
  • Fei Wang
  • Jiahui Pan

Few-shot learning (FSL) poses a considerable challenge since it aims to improve the model generalization ability with limited labeled data. Previous works usually attempt to construct class-specific prototypes and then predict novel classes using these prototypes. However, the feature distribution represented by the limited labeled data is coarse-grained, leading to large information gap between the labeled and unlabeled data as well as biases in the prototypes. In this paper, we investigate the correlation between sample quality and density, and propose a Density-driven Iterative Prototype Optimization to acquire high-quality prototypes, and further improve few-shot learning performance. Specifically, the proposed method consists of two optimization strategies. The similarity-evaluating strategy is for capturing the information gap between the labeled and unlabeled data by reshaping the feature manifold for the novel feature distribution. The density-driven strategy is proposed to iteratively refine the prototypes in the direction of density growth. The proposed method could reach or even exceed the state-of-the-art performance on four benchmark datasets, including mini-ImageNet, tiered-ImageNet, CUB, and CIFAR-FS. The code will be available soon at https: //github. com/tailofcat/DIPO.

EAAI Journal 2024 Journal Article

An attention-based adaptive spatial–temporal graph convolutional network for long-video ergonomic risk assessment

  • Chengju Zhou
  • Jiayu Zeng
  • Lina Qiu
  • Shuxi Wang
  • Pingzhi Liu
  • Jiahui Pan

Ergonomic risk assessment (ERA) is commonly used to identify and analyze postures that are detrimental to the health of workers in industrial workplaces, which is vital to prevent work-related musculoskeletal disorders (WMSDs). Among the automatic approaches, algorithms based on graph convolutional networks (GCNs) have shown promising results in ERA using skeleton sequence as input. However, previous GCN-based methods still have certain limitations. First, the separated modeling of spatial and temporal information and the manually pre-defined topology of graph may restrict the representation diversity of the networks. Additionally, RNN-based temporal modeling often incurs high computational costs and fails to capture long-range temporal dependencies, thereby reducing flexibility in describing long videos. To overcome these challenges, in this study, we propose an attention-based adaptive spatial–temporal graph convolutional network (AAST-GCN), aiming to achieve effective and efficient action representation for ERA in long video. First, we employ an alternate modeling strategy to effectively capture the spatial–temporal information, and propose an improved adaptive adjacency matrix scheme to learn various coordination and relations of body-joints, thus enhancing the flexibility to model diverse postures. Furthermore, we introduce an efficient multi-scale temporal convolutional network as a replacement for RNN-based algorithms, enabling the network to extract various granularities of temporal features. Moreover, to make the network focuses on more valuable information, we employ a spatial–temporal interaction attention (STIA) module. Finally, the aforementioned modules are aggregated within a multi-task learning framework, with the action segmentation serving as the auxiliary task to further improve the accuracy of ERA. We conducted the ergonomic risk assessment on the UW-IOM and TUM Kitchen datasets using our network. Extensive experiments conducted on the most popular datasets UW-IOM and TUM Kitchen demonstrated that our proposed AAST-GCN outperforms other GCN-based methods. Ablation studies and visualization also prove the effectiveness of the individual sub-modules.

AAAI Conference 2024 Conference Paper

CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive Learning

  • Bingzhi Chen
  • Sisi Fu
  • Yishu Liu
  • Jiahui Pan
  • Guangming Lu
  • Zheng Zhang

Dental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intricacies. To bridge this gap, we release a hospital-scale panoramic dental X-ray benchmark, namely “CariesXrays”, to facilitate the advancements in high-precision computer-aided diagnosis for dental caries. It comprises 6,000 panoramic dental X-ray images, with a total of 13,783 instances of dental caries, all meticulously annotated by dental professionals. In this paper, we propose a novel Feature Pyramid Contrastive Learning (FPCL) framework, that jointly incorporates feature pyramid learning and contrastive learning within a unified diagnostic paradigm for automated dental caries detection. Specifically, a robust dual-directional feature pyramid network (D2D-FPN) is designed to adaptively capture rich and informative contextual information from multi-level feature maps, thus enhancing the generalization ability of caries detection across different scales. Furthermore, our model is augmented with an effective proposals-prototype contrastive regularization learning (P2P-CRL) mechanism, which can flexibly bridge the semantic gaps among diverse dental caries with varying appearances, resulting in high-quality dental caries proposals. Extensive experiments on our newly-established CariesXrays benchmark demonstrate the potential of FPCL to make a significant social impact on caries diagnosis.

IJCAI Conference 2024 Conference Paper

Innovative Directional Encoding in Speech Processing: Leveraging Spherical Harmonics Injection for Multi-Channel Speech Enhancement

  • Jiahui Pan
  • Pengjie Shen
  • Hui Zhang
  • Xueliang Zhang

Multi-channel speech enhancement leverages multiple microphones to extract target speech signals amid background noise. Effectively utilizing directional cues is key for robust enhancement. While deep learning shows promise for multi-channel speech processing, most methods operate on short-time Fourier transform (STFT) coefficients directly. We propose using spherical harmonics transform (SHT) coefficients as auxiliary inputs to models. which concisely represent spatial distributions. SHT allows signals from varying numbers of microphones to be converted into coefficients of a consistent dimension. The proposed technique enables a single model to generalize to microphone arrays with varying configurations, rather than requiring a specialized model for each array layout. We present two architectures with SHT-based auxiliary inputs: parallel and serial. Specifically, the parallel model contains two encoders - one for STFT and another for SHT. By fusing both encoders' outputs in the decoder to estimate the enhanced STFT, it effectively incorporates spatial context. For the serial approach, we first apply SHT to the signals and then take STFT of the transformed signals as network inputs. Evaluations of the TIMIT dataset under fluctuating noise and reverberation demonstrate our model outperforms established benchmarks. Remarkably, these results are attained with reduced computations and parameters. Furthermore, experiments on the MS-SNSD dataset show the proposed method can enhance the generalization ability of networks. The source code is publicly accessible at https: //github. com/Pandade1997/SH_injection.

EAAI Journal 2024 Journal Article

Multi-degree-of-freedom unmanned aerial vehicle control combining a hybrid brain-computer interface and visual obstacle avoidance

  • Shanghong Xie
  • Wei Gao
  • Zhen Zeng
  • Qingfu Wu
  • Qian Huang
  • Nianming Ban
  • Qian Wu
  • Jiahui Pan

Objective The difficulty of unmanned aerial vehicle (UAV) control recently lies in multidirectional movement in 3-dimensional space, improving control accuracy and manipulation safety. To address these challenges, a UAV control system that incorporates a hybrid brain-computer interface (hBCI), gyroscope and visual obstacle avoidance based on monocular depth estimation is proposed. Approach. We propose an efficient steady-state visual evoked potential (SSVEP) classification network (CL-NET) featuring a one-dimensional convolutional neural network, a long short-term memory module and an attention module to identify the user's intention for UAV movement in the front, back, left and right directions. The take-off, landing and rising control of the UAV is realized by an electrooculogram (EOG) signal detection algorithm, a blink state detector. In addition, the UAV can fly in an oblique state and rotate according to the current head posture detected by a gyroscope. Furthermore, an improved monocular depth estimation network is employed to design the autonomous obstacle avoidance module of the UAV, ensuring the safety of the brain-controlled system in practice. Main results. The proposed CL-NET delivers an accuracy of 98. 67% on the public dataset and an accuracy of 97. 92% on the self-collected dataset, both of which surpass the performance of state-of-the-art models. Additionally, we set up a brain control group and a remote control group to conduct practical experiments in a realistic environment. In the experiments involving sixteen subjects, the proposed UAV control system reached an average information transfer rate (ITR) of 44. 09 bits/min, and the brain control group had a lower collision rate than the remote control group. Significance. The hybrid control method ensures that the multi-degree-of-freedom (multi-DOF) UAV control system maintains outstanding performance while ensuring good safety.

YNIMG Journal 2024 Journal Article

Precise detection of awareness in disorders of consciousness using deep learning framework

  • Huan Yang
  • Hang Wu
  • Lingcong Kong
  • Wen Luo
  • Qiuyou Xie
  • Jiahui Pan
  • Wuxiu Quan
  • Lianting Hu

Diagnosis of disorders of consciousness (DOC) remains a formidable challenge. Deep learning methods have been widely applied in general neurological and psychiatry disorders, while limited in DOC domain. Considering the successful use of resting-state functional MRI (rs-fMRI) for evaluating patients with DOC, this study seeks to explore the conjunction of deep learning techniques and rs-fMRI in precisely detecting awareness in DOC. We initiated our research with a benchmark dataset comprising 140 participants, including 76 unresponsive wakefulness syndrome (UWS), 25 minimally conscious state (MCS), and 39 Controls, from three independent sites. We developed a cascade 3D EfficientNet-B3-based deep learning framework tailored for discriminating MCS from UWS patients, referred to as "DeepDOC", and compared its performance against five state-of-the-art machine learning models. We also included an independent dataset consists of 11 DOC patients to test whether our model could identify patients with cognitive motor dissociation (CMD), in which DOC patients were behaviorally diagnosed unconscious but could be detected conscious by brain computer interface (BCI) method. Our results demonstrate that DeepDOC outperforms the five machine learning models, achieving an area under curve (AUC) value of 0.927 and accuracy of 0.861 for distinguishing MCS from UWS patients. More importantly, DeepDOC excels in CMD identification, achieving an AUC of 1 and accuracy of 0.909. Using gradient-weighted class activation mapping algorithm, we found that the posterior cortex, encompassing the visual cortex, posterior middle temporal gyrus, posterior cingulate cortex, precuneus, and cerebellum, as making a more substantial contribution to classification compared to other brain regions. This research offers a convenient and accurate method for detecting covert awareness in patients with MCS and CMD using rs-fMRI data.

JBHI Journal 2024 Journal Article

ST-SCGNN: A Spatio-Temporal Self-Constructing Graph Neural Network for Cross-Subject EEG-Based Emotion Recognition and Consciousness Detection

  • Jiahui Pan
  • Rongming Liang
  • Zhipeng He
  • Jingcong Li
  • Yan Liang
  • Xinjie Zhou
  • Yanbin He
  • Yuanqing Li

In this paper, a novel spatio-temporal self-constructing graph neural network (ST-SCGNN) is proposed for cross-subject emotion recognition and consciousness detection. For spatio-temporal feature generation, activation and connection pattern features are first extracted and then combined to leverage their complementary emotion-related information. Next, a self-constructing graph neural network with a spatio-temporal model is presented. Specifically, the graph structure of the neural network is dynamically updated by the self-constructing module of the input signal. Experiments based on the SEED and SEED-IV datasets showed that the model achieved average accuracies of 85. 90% and 76. 37%, respectively. Both values exceed the state-of-the-art metrics with the same protocol. In clinical besides, patients with disorders of consciousness (DOC) suffer severe brain injuries, and sufficient training data for EEG-based emotion recognition cannot be collected. Our proposed ST-SCGNN method for cross-subject emotion recognition was first attempted in training in ten healthy subjects and testing in eight patients with DOC. We found that two patients obtained accuracies significantly higher than chance level and showed similar neural patterns with healthy subjects. Covert consciousness and emotion-related abilities were thus demonstrated in these two patients. Our proposed ST-SCGNN for cross-subject emotion recognition could be a promising tool for consciousness detection in DOC patients.

YNIMG Journal 2023 Journal Article

Identifying patients with cognitive motor dissociation using resting-state temporal stability

  • Hang Wu
  • Qiuyou Xie
  • Jiahui Pan
  • Qimei Liang
  • Yue Lan
  • Yequn Guo
  • Junrong Han
  • Musi Xie

Using task-dependent neuroimaging techniques, recent studies discovered a fraction of patients with disorders of consciousness (DOC) who had no command-following behaviors but showed a clear sign of awareness as healthy controls, which was defined as cognitive motor dissociation (CMD). However, existing task-dependent approaches might fail when CMD patients have cognitive function (e.g., attention, memory) impairments, in which patients with covert awareness cannot perform a specific task accurately and are thus wrongly considered unconscious, which leads to false-negative findings. Recent studies have suggested that sustaining a stable functional organization over time, i.e., high temporal stability, is crucial for supporting consciousness. Thus, temporal stability could be a powerful tool to detect the patient's cognitive functions (e.g., consciousness), while its alteration in the DOC and its capacity for identifying CMD were unclear. The resting-state fMRI (rs-fMRI) study included 119 participants from three independent research sites. A sliding-window approach was used to investigate global and regional temporal stability, which measured how stable the brain's functional architecture was across time. The temporal stability was compared in the first dataset (36/16 DOC/controls), and then a Support Vector Machine (SVM) classifier was built to discriminate DOC from controls. Furthermore, the generalizability of the SVM classifier was tested in the second independent dataset (35/21 DOC/controls). Finally, the SVM classifier was applied to the third independent dataset, where patients underwent rs-fMRI and brain-computer interface assessment (4/7 CMD/potential non-CMD), to test its performance in identifying CMD. Our results showed that global and regional temporal stability was impaired in DOC patients, especially in regions of the cingulo-opercular task control network, default-mode network, fronto-parietal task control network, and salience network. Using temporal stability as the feature, the SVM model not only showed good performance in the first dataset (accuracy = 90%), but also good generalizability in the second dataset (accuracy = 84%). Most importantly, the SVM model generalized well in identifying CMD in the third dataset (accuracy = 91%). Our preliminary findings suggested that temporal stability could be a potential tool to assist in diagnosing CMD. Furthermore, the temporal stability investigated in this study also contributed to a deeper understanding of the neural mechanism of consciousness.

v2026.09.13