2020/recoapy-data-recording-pre-processing-and-phonetic-transcription-for-end-to-end-speech-based-applications

RECOApy: Data recording, pre-processing and phonetic transcription for end-to-end speech-based applications

Deep learning enables the development of efficient end-to-end speechprocessing applications while bypassing the need for expert linguistic andsignal processing features. Yet, recent studies show that good quality speechresources and phonetic transcription of the training data can enhance theresults of these applications. In this paper, the RECOApy tool is introduced.RECOApy streamlines the steps of data recording and pre-processing required inend-to-end speech-based applications. The tool implements an easy-to-useinterface for prompted speech recording, spectrogram and waveform analysis,utterance-level normalisation and silence trimming, as well grapheme-to-phonemeconversion of the prompts in eight languages: Czech, English, French, German,Italian, Polish, Romanian and Spanish. The grapheme-to-phoneme (G2P) converters are deep neural network (DNN) basedarchitectures trained on lexicons extracted from the Wiktionary onlinecollaborative resource. With the different degree of orthographic transparency,as well as the varying amount of phonetic entries across the languages, theDNN's hyperparameters are optimised with an evolution strategy. The phoneme andword error rates of the resulting G2P converters are presented and discussed.The tool, the processed phonetic lexicons and trained G2P models are madefreely available.

Related projects

No projects linked.

Attachments

No attachments yet.