Arrow Research search
Back to ICRA

ICRA 2014

Action recognition using ensemble weighted multi-instance learning

Conference Paper RGB-D Perception: Recognition Artificial Intelligence ยท Robotics

Abstract

This paper deals with recognizing human actions in depth video data. Current state-of-the-art action recognition methods use hand-designed features, which are difficult to produce and time-consuming to extend to new modalities. In this paper, we propose a novel, 3. 5D representation of a depth video for action recognition. A 3. 5D graph of the depth video consists of a set of nodes that are the joints of the human body. Each joint is represented by a set of spatio-temporal features, which are computed by an unsupervised learning approach. However, if occlusions occur, the 3D positions of the joints are noisy which increases the intra-class variations in action classes. To address this problem, we propose the Ensemble Weighted Multi-Instance Learning approach (EnwMi) for the action recognition task. It considers the class imbalance and intra-class variations. We formulate the action recognition task with depth videos as a weighted multi-instance problem. We further integrate an ensemble learning method into the weighted multi-instance learning framework. Our approach is evaluated on Microsoft Research Action3D dataset, and the results show that it outperforms state-of-the-art methods.

Authors

Keywords

  • Joints
  • Three-dimensional displays
  • Training
  • Kernel
  • Histograms
  • Feature extraction
  • Action Recognition
  • Unsupervised Learning
  • Ensemble Method
  • Class Imbalance
  • Joint Position
  • Spatiotemporal Characteristics
  • 3D Position
  • Action Classes
  • Intra-class Variance
  • Unsupervised Learning Approach
  • Action Recognition Task
  • Graphical Representation
  • Hidden Markov Model
  • Feature Learning
  • Multi-label
  • Error Data
  • High-level Features
  • Depth Map
  • Depth Images
  • Depth Camera
  • Multiple Kernel Learning
  • Histogram Features
  • Ensemble Learning Approach
  • Units In Layer
  • Skeleton Data
  • Input Units
  • Kernel Images
  • Local Interests

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
13481466901506664
v2026.09.13