Arrow Research search

Author name cluster

Yonghui Yu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
1 author row

Possible papers

3

EAAI Journal 2026 Journal Article

An emotion recognition approach using peripheral physiological signals based on hierarchical gated residuals and receptive field attention

  • Yonghui Yu
  • Hongji Xu
  • Zhikai Xu
  • Yupeng Duan
  • Renzhuo Wang
  • Hao Zheng
  • Yiran Li
  • Yipeng Xu

The field of emotion recognition (ER) based on peripheral physiological signals (PPSs) has gained significant attention due to the analytical capabilities of artificial intelligence (AI) and its extensive applications. However, a common challenge in many existing ER networks lies in effectively extracting and integrating features from the PPSs. Some networks rely solely on single-channel features, while others struggle with efficient multi-channel fusion, often resulting in suboptimal accuracy. Moreover, the lack of PPS-based public ER datasets, with most existing datasets focusing on the electroencephalogram signal, continues to pose an obstacle. The hierarchical gated residual and receptive field attention fusion (HGR-RFAF) network is proposed to address the above challenges. The HGR-RFAF network improves feature extraction and fusion through multiple multi-branch hierarchical residual fusion convolution (MB-HRFC) layers and a multi-scale dilated convolution (MS-DC) layer. Additionally, the I + Lab Emotion (ILEmo) dataset is constructed to address the scarcity of public ER datasets based on PPSs. To evaluate the performance of the HGR-RFAF network, experiments are conducted on the public K-EmoCon, DEAP, and self-constructed ILEmo datasets. By applying the valence-arousal (V-A) model for validation, the HGR-RFAF network achieves accuracies of 91. 15 % (A)/93. 03 % (V) for the K-EmoCon dataset and 81. 44 % (A)/82. 07 % (V) for the DEAP dataset, respectively. On the ILEmo dataset, HGR-RFAF achieves an accuracy of 91. 25 % with the discrete emotion model. Moreover, various PPSs and their combinations are evaluated, highlighting the benefits of fusing multi-channel PPSs along with their complementarities and contributions to ER.

AAAI Conference 2026 Conference Paper

End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer

  • Yonghui Yu
  • Jiahang Cai
  • Xun Wang
  • Wenwu Yang

Existing multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single-person pose estimation. This design relies on heuristic operations such as tracking, RoI cropping, and non-maximum suppression, limiting both accuracy and efficiency. In this paper, we present a fully end-to-end framework for multi-person 2D pose estimation in videos, effectively eliminating heuristic operations. A key challenge is to associate individuals across frames under complex and overlapping temporal trajectories. To address this, we introduce a novel Pose-Aware Video transformEr Network (PAVE-Net), which features a spatial encoder to model intra-frame relations and a spatiotemporal pose decoder to capture global dependencies across frames. To achieve accurate temporal association, we propose a pose-aware attention mechanism that enables each pose query to selectively aggregate features corresponding to the same individual across consecutive frames. Additionally, we explicitly model spatiotemporal dependencies among pose keypoints to improve accuracy. Notably, our approach is the first end-to-end method for multi-frame 2D human pose estimation. Extensive experiments show that PAVE-Net substantially outperforms prior image-based end-to-end methods, achieving a 6.0 mAP improvement on PoseTrack2017, and delivers accuracy competitive with state-of-the-art two-stage video-based approaches, while offering significant gains in efficiency.

EAAI Journal 2026 Journal Article

Temporal-spatial parallel multiscale network with sparse three-channel mixed attention for wearable sensor-based human activity recognition

  • Renzhuo Wang
  • Hongji Xu
  • Yiran Li
  • Yonghui Yu
  • Yupeng Duan
  • Zhikai Xu
  • Wentao Ai
  • Xinya Li

Human activity recognition (HAR) based on wearable sensors using deep learning (DL) models has garnered significant attention in recent years. However, existing models encounter several challenges in fully exploiting the information from multi-source sensor positions. Notably, they often fail to provide adequate interpretability in terms of both single-channel attention extraction and inter-channel attention mixing. This paper proposes a novel temporal-spatial parallel multiscale network with sparse three-channel mixed attention (TSPM-STCMA) designed to independently and in parallel learn temporal-spatial features from both single-channel and inter-channel relationships derived from multi-source sensors at various body positions. Specifically, the multiscale dynamic convolution with sparse three-channel mixed attention (MDC-STCMA) module is designed to enhance the interpretability of both single-channel attention and inter-channel attention from the same sensor position. Furthermore, the MDC-STCMA module integrates three components based on an attention mechanism. The three-channel dynamic convolution based on improved squeeze-and-excitation (TCDC-ISE) enables a single channel to generate distinct attention responses. The three-channel mixed attention (TCMA) extracts inter-channel correlations. The sparse attention (SA) mitigates overfitting by retraining features with low contributions. Compared with previous models, the TSPM-STCMA achieves superior recognition accuracies of 98. 72 %, 96. 83 %, and 98. 40 %, on three publicly available datasets, i. e. , Physical Activity Monitoring for Aging People (PAMAP2), OPPORTUNITY activity recognition (OPPORTUNITY), and University of California Irvine HAR (UCI-HAR), respectively, while requiring significantly fewer parameters.

v2026.09.13