Journals & Magazines >IEEE Transactions on Image Pr... >Volume: 31

Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene Understanding

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

Acoustic images are an emergent data modality for multimodal scene understanding. Such images have the peculiarity of distinguishing the spectral signature of the sound c...Show More

Metadata

Abstract:

Acoustic images are an emergent data modality for multimodal scene understanding. Such images have the peculiarity of distinguishing the spectral signature of the sound coming from different directions in space, thus providing a richer information as compared to that derived from single or binaural microphones. However, acoustic images are typically generated by cumbersome and costly microphone arrays which are not as widespread as ordinary microphones. This paper shows that it is still possible to generate acoustic images from off-the-shelf cameras equipped with only a single microphone and how they can be exploited for audio-visual scene understanding. We propose three architectures inspired by Variational Autoencoder, U-Net and adversarial models, and we assess their advantages and drawbacks. Such models are trained to generate spatialized audio by conditioning them to the associated video sequence and its corresponding monaural audio track. Our models are trained using the data collected by a microphone array as ground truth. Thus they learn to mimic the output of an array of microphones in the very same conditions. We assess the quality of the generated acoustic images considering standard generation metrics and different downstream tasks (classification, cross-modal retrieval and sound localization). We also evaluate our proposed models by considering multimodal datasets containing acoustic images, as well as datasets containing just monaural audio signals and RGB video frames. In all of the addressed downstream tasks we obtain notable performances using the generated acoustic data, when compared to the state of the art and to the results obtained using real acoustic images as input.

Published in: IEEE Transactions on Image Processing ( Volume: 31)

Page(s): 7102 - 7115

Date of Publication: 08 November 2022

ISSN Information:

PubMed ID: 36346862

DOI: 10.1109/TIP.2022.3219228

Contents

References is not available for this document.

Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene Understanding

Abstract:

Metadata

Abstract:

ISSN Information:

References

IEEE Account

Purchase Details

Profile Information

Need Help?

Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene Understanding

Alerts

Abstract:

Metadata

Abstract:

ISSN Information:

References

IEEE Account

Purchase Details

Profile Information

Need Help?