South African English Male and Female Multi-emotion Speech Synthesis Corpus (Free Talk)

The database recorded 2,362 sentences (62,401 words) from two male voice talents and a female talent. The total audio duration is about 6.32 hours, including the original silence at the beginning and ending (about 300ms each). The recorded content is organized into 5 texts, including 5 emotions. They are excited, sad, angry, fear and empathy. The female and male speakers were born and raised in South Africa in 1977, 1981, and 1999 respectively. They all speak Standard South African English, and were 25, 42, and 47 years old respectively when recording the database.
Specifications:
ID:
King-TTS-350
Size:
6.32 hours
Language:
English
Sample rate & bit depth
48 kHz,24bit
Recording environment
Studio
Speaker
1 male and 1 female
Devices:
Studio
Accuracy Rate
proofreading -- based on individual word, the accuracy is 99%
Samples
Audio
King-TTS-350-F1_0100025_00005
King-TTS-350-F1_0200015_00001
King-TTS-350-M2_0100001_00003
King-TTS-350-M2_0300008_00006

People also searched for

English Multi-speaker Speech Synthesis Corpus (Natural Conversation)
The database recorded 1496 sentences (40,553 words) from 3 voice talents(2 male talents and 1 female talent ). The total audio duration is about 4.01 hours, including the original silence at the beginning and ending (about 350 ms each).
American English Male and Female Speech Synthesis Corpus (Customer and Audiobook)
The database recorded 581 sentences (15,617 words) from a male and a female voice talents. The total audio duration is about 2.05 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 2 texts. The female speakers' texts are related to customer service, while the male speakers' texts are related to audio books.
Brazilian Portuguese Male and Female Speech Synthesis Corpus
The database recorded 2,924 sentences (49,025 words) from 3 voice talents(2 females and 1 male). The total audio duration is about 6 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 11 texts, F048-03 including multiple fields, such as news, letters, digit, etc. We used pt-BR_xsampa phone set for labeling. The voice talents were born and raised in Brazil, in 1969/1973/1997, with standard Brazilian Portuguese and were 56/51/48 years old when recording the database, with a good line foundation. The recordings have even speech rate.
Spain Spanish Male and Female Speech Synthesis Corpus
The database recorded 3,068 sentences (52,494 words) from 3 voice talents(2 females and 1 male). The total audio duration is about 6 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 11 texts, F021-04 including multiple fields , such as news, letter, digit, etc. We used es-es_sampa phone set for labeling. The voice talents were born and raised in Spain in1962/1984/1987, with standard Spanish, and were 63/41/37 years old when recording the database, with a good line foundation. The recording have even speech rate.

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.