Arrow Research search
Back to IROS

IROS 2022

Robust Human Motion Forecasting using Transformer-based Model

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Comprehending human motion is a fundamental challenge for developing Human-Robot Collaborative applications. Computer vision researchers have addressed this field by only focusing on reducing error in predictions, but not taking into account the requirements to facilitate its implementation in robots. In this paper, we propose a new model based on Transformer that simultaneously deals with the real time 3D human motion forecasting in the short and long term. Our 2-Channel Transformer (2CH-TR) is able to efficiently exploit the spatio-temporal information of a shortly observed sequence (400ms) and generates a competitive accuracy against the current state-of-the-art. 2CH-TR stands out for the efficient performance of the Transformer, being lighter and faster than its competitors. In addition, our model is tested in conditions where the human motion is severely occluded, demonstrating its robustness in reconstructing and predicting 3D human motion in a highly noisy environment. Our experiment results show that the proposed 2CH-TR outperforms the ST-Transformer, which is another state-of-the-art model based on the Transformer, in terms of reconstruction and prediction under the same conditions of input prefix. Our model reduces in 8. 89% the mean squared error of ST-Transformer in short-term prediction, and 2. 57% in long-term prediction in Human3. 6M dataset with 400ms input prefix.

Authors

Keywords

  • Solid modeling
  • Three-dimensional displays
  • Computational modeling
  • Computer architecture
  • Predictive models
  • Transformers
  • Skeleton
  • Human Motion
  • Motion Forecasting
  • Mean Square Error
  • 3D Motion
  • Long-term Prediction
  • Short-term Prediction
  • Convolutional Network
  • Human Bone
  • Long Short-term Memory
  • Recurrent Neural Network
  • Attention Mechanism
  • Linear Interpolation
  • Generative Adversarial Networks
  • Language Model
  • Sequence Of Frames
  • Pose Estimation
  • Graph Convolutional Network
  • Discrete Cosine Transform
  • Axis Angle
  • Human Pose
  • Parts Of The Skeleton
  • Global Rotation
  • Bidirectional Information
  • Learnable Weight Matrix
  • Transformer Block
  • 3D Pose
  • Global Translation
  • Weight Matrix
  • Spatial Dependence
  • Fully Convolutional Network

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
290135134137114514
v2026.09.13