Arrow Research search
Back to EAAI

EAAI 2025

Three-dimensional human pose estimation based on multi-scale spatial–temporal transformer

Journal Article journal-article Applied Artificial Intelligence · Artificial Intelligence

Abstract

Recently, transformer-based methods have become dominant in the domain of three-dimensional (3D) human pose estimation, yet the U-net model based on convolutional neural networks (CNN) struggles to model long temporal sequences, and sequence-to-frame (Seq2frame) and sequence-to-sequence (Seq2seq) approaches often fail to preserve dependencies at the start and end of sequences. To address these challenges, this paper proposes a Multi-Scale Spatial–Temporal Transformer network (MSST). This network utilizes Sequence Padding Module (SPM) to extract edge features of the first and last frames, and employs Spatial–Temporal Transformer (STT) to model the spatial–temporal correlations of keypoints. Additionally, we design a Multi-Scale Module (MSM) that analyzes the human skeletal topology to extract multi-scale features of keypoints, local information, and global information, and fuse semantic information at different scales. Finally, we utilize regression heads to project the processed keypoint feature information into 3D space. We conduct quantitative evaluations on two benchmark datasets using four evaluation metrics and design multiple sets of comparative experiments to validate the effectiveness of the proposed modules. Experimental results demonstrate that the proposed network achieves excellent performance.

Authors

Keywords

  • Multi-scale
  • Sequence padding
  • Spatial–temporal transformer
  • Human pose estimation

Context

Venue
Engineering Applications of Artificial Intelligence
Archive span
1988-2026
Indexed papers
13269
Paper id
881342587241651110
v2026.09.13