Arrow Research search
Back to ICRA

ICRA 2021

Robotic Indoor Scene Captioning from Streaming Video

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Robots are usually equipped with cameras to explore the indoor scene and it is expected that the robot can well describe the scene with natural language. Although some great success has been achieved in image and video captioning technology, especially on many public datasets, the caption generated from indoor scene video is still not informative and coherent enough. In this paper, we propose the problem of Indoor Scene Captioning from Streaming Video, which aims at generating a more accurate and informative caption from streaming video. To solve this problem, we firstly design an algorithm to organize the visual information of the indoor scene into a scene graph, and then implement a scene graph guided captioning method, which takes the scene graph and video frames as input to generate the caption from the video streaming. The proposed framework is evaluated both on the AI2THOR dataset and a real-world robotic platform, demonstrating the effectiveness of the framework.

Authors

Keywords

  • Visualization
  • Automation
  • Conferences
  • Robot vision systems
  • Natural languages
  • Streaming media
  • Cameras
  • Natural Language
  • Scene Graph
  • Video Captioning
  • Syntactic
  • Metadata
  • Object Detection
  • Shortest Path
  • Simulation Environment
  • Bounding Box
  • C=O Groups
  • Class Assignment
  • 3D Coordinates
  • Reference Object
  • Preset Threshold
  • Real-time Video
  • Graph Generation
  • List Of Objects
  • Room Type
  • Opening Sentence
  • Kinect Camera
  • Order Of Objects
  • Bounding Box Location
  • Semantic
  • Image Frames
  • Time Graph

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
640657447291255846
v2026.09.13