Arrow Research search
Back to ICRA

ICRA 2021

GPU-Efficient Dense Convolutional Network for Real-time Semantic Segmentation

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

Real-time semantic segmentation is a challenging task as both accuracy and inference speed need to be considered simultaneously. In real-world applications, it is usually achieved by deploying a deep neural network in modern GPU device. However, most of the work focused on real-time semantic segmentation is designed by significantly reducing computation complexity and model size. There are other factors that have a significant impact on inference speed are overlooked, especially when the network is running in modern GPU device. In this paper, we focus on designing a GPU-efficient network as backbone for real-time semantic segmentation. Dense connectivity can preserve and accumulate feature maps of multiple receptive fields and is therefore ideal for semantic segmentation. Therefore, we design a GPU-efficient network (DenseENet) with dense connectivity. The proposed DenseENet shows an obvious advantage in balancing accuracy and inference speed in modern GPU device. Specifically, on Cityscapes test set, DenseENet with a simple FCN decoder achieves 75. 2% mIoU with 83. 6 FPS for an input of 1024 × 2048 resolution and 73. 6% mIoU with 132 FPS for an input of 768 × 1536 resolution on a single GTX 1080Ti card.

Authors

Keywords

  • Deep learning
  • Automation
  • Conferences
  • Computational modeling
  • Semantics
  • Graphics processing units
  • Real-time systems
  • Convolutional Network
  • Semantic Segmentation
  • Semantic Segmentation Network
  • Real-time Semantic Segmentation
  • Deep Neural Network
  • Feature Maps
  • Model Size
  • Receptive Field
  • Dense Connections
  • Fully Convolutional Network
  • Inference Speed
  • Modern Devices
  • Growth Rate
  • Computational Cost
  • Factorization
  • Input Features
  • Dense Layer
  • Convolution Operation
  • Bottleneck Layer
  • Pointwise Convolution
  • Convolutional Neural Network Model
  • Output Feature Map
  • Neural Architecture Search
  • Degree Of Parallelism
  • Number Of Input Features
  • Transition Layer
  • Batch Normalization
  • Kernel Parameters

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
390900663783911077
v2026.09.13