Conferences >2012 IEEE Spoken Language Tec...

What makes this voice sound so bad? A multidimensional analysis of state-of-the-art text-to-speech systems

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

This paper presents research on perceptual quality dimensions of synthetic speech. We generated 57 stimuli from 16/19 female/male German text-to-speech systems (TTS) and ...Show More

Metadata

Abstract:

This paper presents research on perceptual quality dimensions of synthetic speech. We generated 57 stimuli from 16/19 female/male German text-to-speech systems (TTS) and asked listeners to judge the perceptual distances between them in a sorting task. Through a subsequent multidimensional scaling algorithm, we extracted three dimensions. Via expert listening and a comparison to ratings gathered on 16 attribute scales, the three dimensions can be assigned to naturalness of voice, temporal distortions and calmness. These dimensions are discussed in detail and compared to the perceptual quality dimensions from previous multidimensional analyses. Moreover, the results are analyzed depending on the type of TTS system. The identified dimensions will be used in the future to build a dimension-based quality predictor for synthetic speech.

Published in: 2012 IEEE Spoken Language Technology Workshop (SLT)

Date of Conference: 02-05 December 2012

Date Added to IEEE Xplore: 31 January 2013

ISBN Information:

DOI: 10.1109/SLT.2012.6424229

Conference Location: Miami, FL, USA

Contents

References is not available for this document.

What makes this voice sound so bad? A multidimensional analysis of state-of-the-art text-to-speech systems

Abstract:

Metadata

Abstract:

References

IEEE Account

Purchase Details

Profile Information

Need Help?

What makes this voice sound so bad? A multidimensional analysis of state-of-the-art text-to-speech systems

Alerts

Abstract:

Metadata

Abstract:

References

IEEE Account

Purchase Details

Profile Information

Need Help?