Burmese Multi-speaker Speech Synthesis Corpus (Multi-emotion)

The database recorded 8,550 sentences from a total of four speakers: two male and two female. The total audio duration is approximately 14.47 hours, including the initial and final segments of clear silence (each about 300 milliseconds). Each speaker completed no less than 3 hours of recording, covering five emotions: neutral, happy, sad, angry, and surprised. Each speaker's recording duration for the neutral emotion is no less than 2 hours, while the happy, sad, angry, and surprised emotions are each no less than 0.15 hours. The recorded content is organized into 10 texts, covering a range of fields and emotions, such as news, encyclopedia, technology, conversation, novels, entertainment, happiness, sadness, anger, and surprise. The recordings were annotated using the my_mm_xsampa and en-us_cmu phoneme sets.
Specifications:
ID:
King-TTS-056
Size:
14.47 hours
Language:
Burmese
Sample rate & bit depth
48kHz, 24bit
Recording environment
Professional recording studio
Speaker
2 males and 2 females
Devices:
Studio
Accuracy Rate
proofreading -- based on individual word, the accuracy is 99% phonetic labeling -- based on individual phone, the accuracy is 99.5% prosody labeling -- based on individual symbol, the accuracy is 98%
Samples
Audio
King-TTS-056-01000095
King-TTS-056-01000128
King-TTS-056-03000022
King-TTS-056-07000112

People also searched for

American English Male and Female Speech Synthesis Corpus (Customer and Audiobook)
The database recorded 581 sentences (15,617 words) from a male and a female voice talents. The total audio duration is about 2.05 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 2 texts. The female speakers' texts are related to customer service, while the male speakers' texts are related to audio books.
Brazilian Portuguese Male and Female Speech Synthesis Corpus
The database recorded 2,924 sentences (49,025 words) from 3 voice talents(2 females and 1 male). The total audio duration is about 6 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 11 texts, F048-03 including multiple fields, such as news, letters, digit, etc. We used pt-BR_xsampa phone set for labeling. The voice talents were born and raised in Brazil, in 1969/1973/1997, with standard Brazilian Portuguese and were 56/51/48 years old when recording the database, with a good line foundation. The recordings have even speech rate.
Spain Spanish Male and Female Speech Synthesis Corpus
The database recorded 3,068 sentences (52,494 words) from 3 voice talents(2 females and 1 male). The total audio duration is about 6 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 11 texts, F021-04 including multiple fields , such as news, letter, digit, etc. We used es-es_sampa phone set for labeling. The voice talents were born and raised in Spain in1962/1984/1987, with standard Spanish, and were 63/41/37 years old when recording the database, with a good line foundation. The recording have even speech rate.
New Zealand English Female Speech Synthesis Corpus
The database recorded 1,600 sentences (17,808 words) from a female voice talent. The total audio duration is about 2.04 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is organized into 1 texts, news. The voice talent was born and raised in New Zealand in 1989, with standard New Zealand English. She is a professional voice talent who has many years of experience in dubbing and broadcasting , with a good line foundation.

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.