EAAI Journal 2025 Journal Article
Multimodal feature fusion for human activity recognition using human centric temporal transformer
- Samee Ullah Khan
- Maryam Sultana
- Sufyan Danish
- Deepak Gupta
- Norah Saleh Alghamdi
- Suchang Woo
- Dong-Gyu Lee
- Sangtae Ahn
In recent years, human activity recognition (HAR) has focused considerable interest due to its manifold monitoring applications. Mainstream HAR approaches often face challenges with the reliability of results when relying on a single data modality, especially when integrating heterogeneous data sources. A notable limitation of the implemented artificial intelligence (AI) models is their limited capability to handle dynamic scenarios, as they lack the necessary contextual information from multiple sources, which impedes the models’ adaptability and accuracy. This paper proposes a multi-modality framework for HAR that fuses human concern patterns using various spatiotemporal model flavors. In addition, to get the spatial features, a swin transformer with a dual attention concept is applied to process visual sensor data, while one-dimensional convolutional neural network leverages human skeleton information obtained from the detection model with numerous key points. Later, these multi-modality features are fused to improve the robust analysis and comprehension of activities. Next, these resulting features are passed to the human centric temporal transformer (HCTT), that has the capabilities to process multimodal sequence data for temporal learning. Moreover, the attention block of HCTT enables human-related attentive patterns followed by a dual fusion mechanism. The proposed model was evaluated on four open-access large-scale HAR datasets, where comprehensive ablation studies and comparative analyses demonstrated that our developed multimodal approach outperforms recent baseline HAR models. This underscores its potential for advancing AI applications and human activity analysis.