2021/an-objective-evaluation-of-the-effects-of-recording-conditions-and-speaker-characteristics-in-multi-speaker-deep-neural-speech-synthesis

An objective evaluation of the effects of recording conditions and speaker characteristics in multi-speaker deep neural speech synthesis

Multi-speaker spoken datasets enable the creation of text-to-speech synthesis(TTS) systems which can output several voice identities. The multi-speaker(MSPK) scenario also enables the use of fewer training samples per speaker.However, in the resulting acoustic model, not all speakers exhibit the samesynthetic quality, and some of the voice identities cannot be used at all. In this paper we evaluate the influence of the recording conditions, speakergender, and speaker particularities over the quality of the synthesised outputof a deep neural TTS architecture, namely Tacotron2. The evaluation is possibledue to the use of a large Romanian parallel spoken corpus containing over 81hours of data. Within this setup, we also evaluate the influence of differenttypes of text representations: orthographic, phonetic, and phonetic extendedwith syllable boundaries and lexical stress markings. We evaluate the results of the MSPK system using the objective measures ofequal error rate (EER) and word error rate (WER), and also look into thedistances between natural and synthesised t-SNE projections of the embeddingscomputed by an accurate speaker verification network. The results show thatthere is indeed a large correlation between the recording conditions and thespeaker's synthetic voice quality. The speaker gender does not influence theoutput, and that extending the input text representation with syllableboundaries and lexical stress information does not equally enhance thegenerated audio across all speaker identities. The visualisation of the t-SNEprojections of the natural and synthesised speaker embeddings show that theacoustic model shifts some of the speakers' neural representation, but not allof them. As a result, these speakers have lower performances of the outputspeech.

Related projects

No projects linked.

Attachments

No attachments yet.