Arrow Research search
Back to IROS

IROS 2024

A Point-Based Approach to Efficient LiDAR Multi-Task Perception

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

Multi-task perception networks hold great potential as they can improve performance and computational efficiency compared to their single-task counterparts, facilitating online deployment. However, current multi-task architectures in point cloud perception combine multiple task-specific point cloud representations, each requiring a separate feature encoder, making the network significantly large and slow. In this work, we propose PAttFormer, an efficient multi-task learning architecture for joint semantic segmentation and object detection in point clouds, only relying on a point-based representation. The network builds on transformer-based feature encoders using neighborhood attention and grid-pooling, complemented with a query-based detection decoder using a novel 3D deformable-attention detection head topology. Unlike other LiDAR-based multi-task architectures, our proposed PAttFormer does not require separate feature encoders for multiple task-specific point cloud representations, resulting in a network that is 3× smaller and 1. 4× faster while achieving competitive performance on the nuScenes and KITTI benchmarks for autonomous driving perception. We perform extensive evaluations that show substantial improvement from multi-task learning, achieving +1. 7% in mIoU for LiDAR semantic segmentation and +1. 7% in mAP for 3D object detection on the nuScenes benchmark compared to the single-task models.

Authors

Keywords

  • Point cloud compression
  • Training
  • Three-dimensional displays
  • Laser radar
  • Semantic segmentation
  • Computer architecture
  • Object detection
  • Benchmark testing
  • Multitasking
  • Transformers
  • Benchmark
  • Point Cloud
  • Detection Function
  • Multi-task Learning
  • Feature Encoder
  • Multiple Representations
  • Segmentation Detection
  • Detection Head
  • 3D Object Detection
  • Point Cloud Representation
  • Multiple Cloud
  • Detection Performance
  • Frame Rate
  • Bounding Box
  • Detection Task
  • Similar Improvements
  • Feature Points
  • Segmentation Task
  • Semantic Labels
  • Voxel-based Methods
  • Point Cloud Features
  • Multi-task Training
  • KITTI Dataset
  • Multi-task Model
  • Multi-task Network
  • Semantic Segmentation Results
  • Rotation Error
  • Semantic Segmentation Models

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
293689654099260721
v2026.09.13