S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information
2026-07-28 • Robotics
Robotics
AI summaryⓘ
The authors introduce new robot tasks that require using sounds to find and choose objects to manipulate. They created a learning method called S2A2 that combines what the robot sees with where sounds come from and their characteristics. Their experiments in simulations showed S2A2 works best for tasks needing both sound location and quality recognition. They also tested it on real robots, proving it can work outside of simulations.
acoustic manipulationsound source localizationimitation learningmultimodal learningrobotic manipulationspatial audioaudio signal processingpolicy learningtimbre recognitionreal-robot experiments
Authors
Kaneyoshi Hiratsuka, Benjamin Yen, Ryosuke Kojima
Abstract
Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and $π_0$, into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.