Arrow Research search
Back to ICRA

ICRA 2025

Transferring Visual Knowledge: Semi-Supervised Instance Segmentation for Object Navigation Across Varying Height Viewpoints

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

The object navigation task requires robots to understand the semantic regularities in their environments. However, existing modular object navigation frameworks rely on instance segmentation models trained at fixed camera height viewpoints, limiting generalization performance and increasing labeling costs for new height viewpoints. To tackle this issue, we propose a semi-supervised method that transfers knowledge from a source height to a target height, minimizing the need for additional labels. Our approach introduces three key innovations: i) a projection policy to enhance the teacher model's detection capabilities at the target height, ii) a dynamic weight mechanism that emphasizes high-confidence pseudo-labels to reduce overfitting, and iii) a prototype contrast transferring method to transfer knowl-edge effectively. Experiments on the Habitat- Matterport 3D (HM3D) dataset show our method outperforms state-of-the-art semi-supervised techniques, improving both segmentation accuracy and navigation performance. The code is available at: https://github.com/FreeformRobotics/TransferKnowledge.

Authors

Keywords

  • Instance segmentation
  • Visualization
  • Technological innovation
  • Accuracy
  • Three-dimensional displays
  • Navigation
  • Semantics
  • Robot vision systems
  • Cameras
  • Robots
  • Object Navigation
  • Prototype
  • Knowledge Transfer
  • Teacher Model
  • Dynamic Mechanism
  • Semi-supervised Methods
  • Modular Framework
  • Camera Height
  • Cross-entropy Loss
  • Stochastic Gradient Descent
  • Bounding Box
  • Recognition Accuracy
  • Average Precision
  • Unlabeled Data
  • 3D Point
  • Graph Neural Networks
  • Self-supervised Learning
  • Student Model
  • Modular Method
  • Weight Allocation
  • Viewpoint Changes
  • Soft Labels
  • Vision Transformer
  • Mask R-CNN
  • Memory Bank
  • Categorical Cross-entropy Loss
  • RGB-D Images
  • Confidence Score

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
407802455166728042
v2026.09.13