Arrow Research search
Back to ICRA

ICRA 2022

Audio-Visual Grounding Referring Expression for Robotic Manipulation

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

Referring expressions are commonly used when referring to a specific target in people's daily dialogue. In this paper, we develop a novel task of audio-visual grounding referring expression for robotic manipulation. The robot leverages both the audio and visual information to understand the referring expression in the given manipulation instruction and the corresponding manipulations are implemented. To solve the proposed task, an audio-visual framework is proposed for visual localization and sound recognition. We have also established a dataset which contains visual data, auditory data and manipulation instructions for evaluation. Finally, extensive experiments are conducted both offline and online to verify the effectiveness of the proposed audio-visual framework. And it is demonstrated that the robot performs better with the audio-visual data than with only the visual data.

Authors

Keywords

  • Location awareness
  • Visualization
  • Automation
  • Grounding
  • Task analysis
  • Robots
  • Robot Manipulator
  • Data Visualization
  • Visual Information
  • People’s Daily
  • Sound Detection
  • Target People
  • Audio Information
  • Natural Language
  • Visual Features
  • Recognition Accuracy
  • Target Object
  • Robotic Arm
  • Objects In The Scene
  • Audio Data
  • Auditory Modality
  • Sound Data
  • Robotic Hand
  • Instructional Text
  • Kinect Camera

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
77706183333549620
v2026.09.13