Arrow Research search
Back to IROS

IROS 2010

Using text-spotting to query the world

Conference Paper Computer Vision IV Artificial Intelligence ยท Robotics

Abstract

The world we live in is labeled extensively for the benefit of humans. Yet, to date, robots have made little use of human readable text as a resource. In this paper we aim to draw attention to text as a readily available source of semantic information in robotics by implementing a system which allows robots to read visible text in natural scene images and to use this knowledge to interpret the content of a given scene. The reliable detection and parsing of text in natural scene images is an active area of research and remains a non-trivial problem. We extend a commonly adopted approach based on boosting for the detection and optical character recognition (OCR) for the parsing of text by a probabilistic error correction scheme incorporating a sensor-model for our pipeline. In order to interpret the scene content we introduce a generative model which explains spotted text in terms of arbitrary search terms. This allows the robot to estimate the relevance of a given scene with respect to arbitrary queries such as, for example, whether it is looking at a bank or a restaurant. We present results from images recorded by a robot in a busy cityscape.

Authors

Keywords

  • Robots
  • Optical character recognition software
  • Training
  • Pixel
  • Mathematical model
  • Probabilistic logic
  • Pipelines
  • Source Of Information
  • Restaurants
  • Search Terms
  • Natural Images
  • Image Texture
  • Optical Character Recognition
  • Natural Scene Images
  • Rectangular
  • Training Set
  • Posterior Probability
  • Bounding Box
  • Light Signal
  • Training Round
  • Probabilistic Generative Model

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
317194554788471497
v2026.09.13