Arrow Research search
Back to IROS

IROS 2019

Learning Actions from Human Demonstration Video for Robotic Manipulation

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Learning actions from human demonstration is an emerging trend for designing intelligent robotic systems, which can be referred as video to command. The performance of such approach highly relies on the quality of video captioning. However, the general video captioning methods focus more on the understanding of the full frame, lacking of consideration on the specific object of interests in robotic manipulations. We propose a novel deep model to learn actions from human demonstration video for robotic manipulation. It consists of two deep networks, grasp detection network (GNet) and video captioning network (CNet). GNet performs two functions: providing grasp solutions and extracting the local features for the object of interests in robotic manipulation. CNet outputs the captioning results by fusing the features of both full frames and local objects. Experimental results on UR5 robotic arm show that our method could produce more accurate command from video demonstration than state-of-the-art work, thereby leading to more robust grasping performance.

Authors

Keywords

  • Accuracy
  • Fuses
  • Grasping
  • Artificial neural networks
  • Feature extraction
  • Market research
  • Manipulators
  • Robots
  • Intelligent robots
  • Videos
  • Active Learning
  • Robot Manipulator
  • Video Presentation
  • Deep Network
  • Local Features
  • Robotic System
  • Object Location
  • Frame Features
  • Video Captioning
  • Human Activities
  • Long Short-term Memory
  • Global Features
  • Stochastic Gradient Descent
  • Classification Network
  • Motion Capture
  • Word Embedding
  • Robotic Applications
  • Human Motion
  • Rectangular Box
  • Short Clips
  • Extract Visual Features
  • Inception V3
  • Kinect Camera
  • Human Pose
  • RGB-D Images
  • Long Short-term Memory Unit
  • Raw Video
  • Real Robot
  • End Of Frame
  • Robotic Tasks

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
298917097868309973
v2026.09.13