The database recorded 12,814 sentences (182,020 words) from a male voice talent. The total audio duration is about 27.71 hours, including the original silence at the beginning and ending (about 300 ms each).
The recorded content is organized into 26 texts, including multiple fields, such as news, baike, dialog, weather, etc. We used ar-msa_sampa & en-us_cmu phone set for labeling.
The voice talent was born and raised in Lebanon in 1968, with standard Arabic and good English, and was 56 years old when recording the database. The recording has even speech rate.
proofreading -- based on individual word, the accuracy is 99%
phonetic labeling -- based on individual phone, the accuracy is 99.5%
prosody labeling -- based on individual symbol, the accuracy is 98%
POS labeling -- based on individual symbol, the accuracy is 98%
phone labeling -- the boundary error is less than 10 ms, the accuracy is 99%
vowel labeling -- based on individual word, the accuracy is 99%
Samples
Audio
King-TTS-321-03000388
King-TTS-321-09000095
King-TTS-321-30000279
King-TTS-321-32000010
People also searched for
English Multi-speaker Speech Synthesis Corpus (Natural Conversation)
The database recorded 1496 sentences (40,553 words) from 3 voice talents(2 male talents and 1 female talent ). The total audio duration is about 4.01 hours, including the original silence at the beginning and ending (about 350 ms each).
American English Male and Female Speech Synthesis Corpus (Customer and Audiobook)
The database recorded 581 sentences (15,617 words) from a male and a female voice talents. The total audio duration is about 2.05 hours, including the original silence at the beginning and ending (about 300 ms each).
The recorded content is organized into 2 texts. The female speakers' texts are related to customer service, while the male speakers' texts are related to audio books.
Brazilian Portuguese Male and Female Speech Synthesis Corpus
The database recorded 2,924 sentences (49,025 words) from 3 voice talents(2 females and 1 male). The total audio duration is about 6 hours, including the original silence at the beginning and ending (about 300 ms each).
The recorded content is organized into 11 texts, F048-03 including multiple fields, such as news, letters, digit, etc. We used pt-BR_xsampa phone set for labeling.
The voice talents were born and raised in Brazil, in 1969/1973/1997, with standard Brazilian Portuguese and were 56/51/48 years old when recording the database, with a good line foundation. The recordings have even speech rate.
Spain Spanish Male and Female Speech Synthesis Corpus
The database recorded 3,068 sentences (52,494 words) from 3 voice talents(2 females and 1 male). The total audio duration is about 6 hours, including the original silence at the beginning and ending (about 300 ms each).
The recorded content is organized into 11 texts, F021-04 including multiple fields , such as news, letter, digit, etc. We used es-es_sampa phone set for labeling.
The voice talents were born and raised in Spain in1962/1984/1987, with standard Spanish, and were 63/41/37 years old when recording the database, with a good line foundation. The recording have even speech rate.