Arrow Research search
Back to IROS

IROS 2023

Optical Flow Boosts Unsupervised Localization and Segmentation

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Unsupervised localization and segmentation are long-standing robot vision challenges that describe the critical ability for an autonomous robot to learn to decompose images into individual objects without labeled data. These tasks are important because of the limited availability of dense image manual annotation and the promising vision of adapting to an evolving set of object categories in lifelong learning. Most recent methods focus on using visual appearance continuity as object cues by spatially clustering features obtained from self-supervised vision transformers (ViT). In this work, we leverage motion cues, inspired by the common fate principle that pixels that share similar movements tend to belong to the same object. We propose a new loss term formulation that uses optical flow in unlabeled videos to encourage self-supervised ViT features to become closer to each other if their corresponding spatial locations share similar movements, and vice versa. We use the proposed loss function to finetune vision transformers that were originally trained on static images. Our fine-tuning procedure outperforms state-of-the-art techniques for unsupervised semantic segmentation through linear probing, without the use of any labeled data. This procedure also demonstrates increased performance over original ViT networks across unsupervised object localization and semantic segmentation benchmarks. Our code is available at https://github.com/mlzxy/flowdino.

Authors

Keywords

  • Location awareness
  • Optical losses
  • Visualization
  • Semantic segmentation
  • Robot vision systems
  • Object segmentation
  • Transformers
  • Optical Flow
  • Unsupervised Localization
  • Fine-tuned
  • Object Location
  • Spatial Clustering
  • Lifelong Learning
  • Set Of Categories
  • Linear Probe
  • Density Imaging
  • Vision Transformer
  • Feature Maps
  • Visual Features
  • Bounding Box
  • Video Frames
  • Local Neighborhood
  • Evaluation Protocol
  • Correct Location
  • Still Images
  • Self-supervised Learning
  • Adjacent Frames
  • Similar Motion
  • Optical Loss
  • Qualitative Examples
  • Flow Map
  • Motion Information
  • PASCAL VOC
  • Dense Prediction
  • ImageNet Classification

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
83220513199505573
v2026.09.13