Arrow Research search
Back to IROS

IROS 2024

Robot Synesthesia: A Sound and Emotion Guided Robot Painter

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

If a picture paints a thousand words, sound may voice a million. While recent robotic painting and image synthesis methods have achieved progress in generating visuals from text inputs, the translation of sound into images is vastly unexplored. Generally, sound-based interfaces and sonic interactions have the potential to expand accessibility and control for the user and provide a means to convey complex emotions and the dynamic aspects of the real world. In this paper, we propose an approach for using sound and speech to guide a robotic painting process, known here as robot synesthesia. For general sound, we encode the simulated paintings and input sounds into the same latent space. For speech, we decouple speech audio into its transcribed text and the tone of the speech. Whereas we use the text to control the content, we estimate the emotions from the tone to guide the mood of the painting. Our approach has been fully integrated with FRIDA, a robotic painting framework, adding sound and speech to FRIDA’s existing input modalities such as text and style. In two surveys, participants were able to correctly guess the emotion or natural sound used to generate a given painting more than twice as likely as random chance. On our sound-guided image manipulation and music-guided paintings, we discuss the results qualitatively.

Authors

Keywords

  • Surveys
  • Visualization
  • Translation
  • Image synthesis
  • Mood
  • Generators
  • Speech processing
  • Robots
  • Intelligent robots
  • Painting
  • Synaesthesia
  • Use Of Imaging
  • Latent Space
  • Input Modalities
  • User Control
  • Input Text
  • Natural Sounds
  • Loss Function
  • Neural Network
  • Semantic
  • Spoken Language
  • Image Generation
  • Robotic System
  • Visual Arts
  • Emotion Type
  • Thunderstorm
  • Emotional Context
  • Emotional Prosody
  • Brush Strokes
  • Audio Input
  • Mel-frequency Cepstral Coefficients
  • Speech Input

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
200840794607112835
v2026.09.13