3D object instance recognition and pose estimation using triplet loss with dynamic margin

Sergey Zakharov; Wadim Kehl; Benjamin Planche; Andreas Hutter; Slobodan Ilic

Back to IROS

IROS 2017

3D object instance recognition and pose estimation using triplet loss with dynamic margin

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Details

Abstract

In this paper, we address the problem of 3D object instance recognition and pose estimation of localized objects in cluttered environments using convolutional neural networks. Inspired by the descriptor learning approach of Wohlhart et al. [1], we propose a method that introduces the dynamic margin in the manifold learning triplet loss function. Such a loss function is designed to map images of different objects under different poses to a lower-dimensional, similarity-preserving descriptor space on which efficient nearest neighbor search algorithms can be applied. Introducing the dynamic margin allows for faster training times and better accuracy of the resulting low-dimensional manifolds. Furthermore, we contribute the following: adding in-plane rotations (ignored by the baseline method) to the training, proposing new background noise types that help to better mimic realistic scenarios and improve accuracy with respect to clutter, adding surface normals as another powerful image modality representing an object surface leading to better performance than merely depth, and finally implementing an efficient online batch generation that allows for better variability during the training phase. We perform an exhaustive evaluation to demonstrate the effects of our contributions. Additionally, we assess the performance of the algorithm on the large BigBIRD dataset [2] to demonstrate good scalability properties of the pipeline with respect to the number of models.

Authors

Keywords

Three-dimensional displays
Training
Manifolds
Solid modeling
Pose estimation
Clutter
Two dimensional displays
Triplet Loss
Human Pose Estimation
Object Pose
3D Pose
3D Instance
3D Object Pose
Loss Function
Convolutional Neural Network
Background Noise
Performance Of Algorithm
Training Phase
Baseline Methods
Types Of Noise
Nearest Neighbor Search
In-plane Rotation
Surface Normals
Descriptor Space
Training Set
Input Image
RGB Channels
Objective View
Background Clutter
Real Background
Additional Degrees Of Freedom
Scaling Method
Similar Pose
Siamese Network
Real-world Scenarios
Field Of Image Processing

Context

Venue: IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span: 1988-2025
Indexed papers: 26578
Paper id: 805283393359292782