Arrow Research search
Back to IROS

IROS 2017

SMSnet: Semantic motion segmentation using deep convolutional neural networks

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Interpreting the semantics and motion of objects are prerequisites for autonomous robots that enable them to reason and operate in dynamic real-world environments. Existing approaches that tackle the problem of semantic motion segmentation consist of long multistage pipelines and typically require several seconds to process each frame. In this paper, we present a novel convolutional neural network architecture that learns to predict both the object label and motion status of each pixel in an image. Given a pair of consecutive images, the network learns to fuse features from self-generated optical flow maps and semantic segmentation kernels to yield pixel-wise semantic motion labels. We also introduce the Cityscapes-Motion dataset which contains over 2, 900 manually annotated semantic motion labels, which is the largest dataset of its kind so far. We demonstrate that our network outperforms existing approaches achieving state-of-the-art performance on the KITTI dataset, as well as in the more challenging Cityscapes-Motion dataset while being substantially faster than existing techniques.

Authors

Keywords

  • Semantics
  • Motion segmentation
  • Optical imaging
  • Computer vision
  • Computer architecture
  • Cameras
  • Neural Network
  • Convolutional Network
  • Convolutional Neural Network
  • Deep Network
  • Deep Neural Network
  • Deep Convolutional Neural Network
  • Semantic Segmentation
  • Image Pixels
  • Convolutional Neural Network Architecture
  • Optical Flow
  • Object Motion
  • Motion State
  • Consecutive Images
  • Semantic Labels
  • Convolutional Architecture
  • Label Of Pixel
  • KITTI Dataset
  • Semantic Segmentation Problem
  • Feature Maps
  • Feature Learning
  • Motion Features
  • Semantic Features
  • Semantic Annotation
  • Neural Network Training
  • Consecutive Frames
  • Xeon E5
  • Fully Convolutional Network
  • Conditional Random Field
  • Object Classification
  • Ground Truth Annotations

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
454779232137092229
v2026.09.13