Methodologies for designing and recording speech databases for corpus based synthesis. de Oliveira, L. C.; Paulo, S.; Figueira, L.; Mendes, C.; Nunes, A.; and Godinho, J. In LREC 2008. Proceedings of the 6th International Conference on Language Resources and Evaluation. Marrakech, Morocco, May 26 - June 1, 2008.
Methodologies for designing and recording speech databases for corpus based synthesis [pdf]Paper  abstract   bibtex   
In this paper we share our experience and describe the methodologies that we have used in designing and recording large speech databases for applications requiring speech synthesis. Given the growing demand for customized and domain specific voices for use in corpus based synthesis systems, we believe that good practices should be established for the creation of these databases which are a key factor in the quality of the resulting speech synthesizer. We will focus on the designing of the recording prompts, on the speaker selection procedure, on the recording setup and on the quality control of the resulting database. One of the major challenges was to assure the uniformity of the recordings during the 20 two hour recording sessions that each speaker had to perform, to produce a total of 13 hours of recorded speech for each of the four speakers. This work was conducted in the scope of the Tecnovoz project that brought together 4 speech research centers and 9 companies with the goal of integrating speech technologies in a wide range of applications.
@inproceedings{de_oliveira_methodologies_2008,
	Author = {de Oliveira, Luís Caldas and Paulo, Sérgio and Figueira, Luís and Mendes, Carlos and Nunes, Ana and Godinho, Joaquim},
	Booktitle = {LREC 2008. Proceedings of the 6th International Conference on Language Resources and Evaluation},
	Date = {2008},
	Date-Modified = {2016-09-23 19:22:38 +0000},
	File = {Attachment:files/8652/Oliveira et al. - 2008 - Methodologies for designing and recording speech databases for corpus based synthesis.pdf:application/pdf},
	Keywords = {Portuguese, speech synthesis, speech technology},
	Publisher = {Marrakech, Morocco, May 26 - June 1, 2008},
	Title = {Methodologies for designing and recording speech databases for corpus based synthesis},
	Url = {http://www.inesc-id.pt/pt/indicadores/Ficheiros/5011.pdf},
	Abstract = {In this paper we share our experience and describe the methodologies that we have used in designing and recording large speech databases for applications requiring speech synthesis. Given the growing demand for customized and domain specific voices for use in corpus based synthesis systems, we believe that good practices should be established for the creation of these databases which are a key factor in the quality of the resulting speech synthesizer. We will focus on the designing of the recording prompts, on the speaker selection procedure, on the recording setup and on the quality control of the resulting database. One of the major challenges was to assure the uniformity of the recordings during the 20 two hour recording sessions that each speaker had to perform, to produce a total of 13 hours of recorded speech for each of the four speakers. This work was conducted in the scope of the Tecnovoz project that brought together 4 speech research centers and 9 companies with the goal of integrating speech technologies in a wide range of applications.},
	Bdsk-File-1 = {YnBsaXN0MDDUAQIDBAUGJCVYJHZlcnNpb25YJG9iamVjdHNZJGFyY2hpdmVyVCR0b3ASAAGGoKgHCBMUFRYaIVUkbnVsbNMJCgsMDxJXTlMua2V5c1pOUy5vYmplY3RzViRjbGFzc6INDoACgAOiEBGABIAFgAdccmVsYXRpdmVQYXRoWWFsaWFzRGF0YV8QZC4uLy4uLy4uL0JpYmxpb2dyYWZpYS9QYXBlcnMvT2xpdmVpcmEvTWV0aG9kb2xvZ2llcyBmb3IgZGVzaWduaW5nIGFuZCByZWNvcmRpbmcgc3BlZWNoIGRhdGFiYXNlcy5wZGbSFwsYGVdOUy5kYXRhTxECZAAAAAACZAACAAAMTWFjaW50b3NoIEhEAAAAAAAAAAAAAAAAAAAAy/YfzkgrAAAQhnL9H01ldGhvZG9sb2dpZXMgZm9yICMxMDg2NzJGRi5wZGYAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAABCGcv/UCdO6AAAAAAAAAAAAAwAEAAAJIAAAAAAAAAAAAAAAAAAAAAhPbGl2ZWlyYQAQAAgAAMv2A64AAAARAAgAANQJt5oAAAABABQQhnL9EIZljgAF/EcABfuYAADARgACAGZNYWNpbnRvc2ggSEQ6VXNlcnM6AGpvYXF1aW1fbGxpc3RlcnJpOgBCaWJsaW9ncmFmaWE6AFBhcGVyczoAT2xpdmVpcmE6AE1ldGhvZG9sb2dpZXMgZm9yICMxMDg2NzJGRi5wZGYADgB+AD4ATQBlAHQAaABvAGQAbwBsAG8AZwBpAGUAcwAgAGYAbwByACAAZABlAHMAaQBnAG4AaQBuAGcAIABhAG4AZAAgAHIAZQBjAG8AcgBkAGkAbgBnACAAcwBwAGUAZQBjAGgAIABkAGEAdABhAGIAYQBzAGUAcwAuAHAAZABmAA8AGgAMAE0AYQBjAGkAbgB0AG8AcwBoACAASABEABIAc1VzZXJzL2pvYXF1aW1fbGxpc3RlcnJpL0JpYmxpb2dyYWZpYS9QYXBlcnMvT2xpdmVpcmEvTWV0aG9kb2xvZ2llcyBmb3IgZGVzaWduaW5nIGFuZCByZWNvcmRpbmcgc3BlZWNoIGRhdGFiYXNlcy5wZGYAABMAAS8AABUAAgAY//8AAIAG0hscHR5aJGNsYXNzbmFtZVgkY2xhc3Nlc11OU011dGFibGVEYXRhox0fIFZOU0RhdGFYTlNPYmplY3TSGxwiI1xOU0RpY3Rpb25hcnmiIiBfEA9OU0tleWVkQXJjaGl2ZXLRJidUcm9vdIABAAgAEQAaACMALQAyADcAQABGAE0AVQBgAGcAagBsAG4AcQBzAHUAdwCEAI4A9QD6AQIDagNsA3EDfAOFA5MDlwOeA6cDrAO5A7wDzgPRA9YAAAAAAAACAQAAAAAAAAAoAAAAAAAAAAAAAAAAAAAD2A==},
	Bdsk-Url-1 = {http://www.inesc-id.pt/pt/indicadores/Ficheiros/5011.pdf}}
Downloads: 0