TTS

Search our off-the-shelf datasets.

Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios
Malay Multi-speaker Speech Synthesis Corpus (Multi-emotion)
The database recorded 5994 sentences ( 72583 words) from 3 voice talents(2 males and 1 female). The total audio duration is about 10.84 hours, including the cleared silent segments at the beginning and ending (about 300 ms each). The recorded content is organized into 8 texts, including multiple fields and emotions, such as news, encyclopedia, dialog, travel and happy, angry, sad, surprise, etc. We used ms-my_xsampa phone set for labeling. The voice talent were born and raised in Malaysia, majoring in broadcasting and performing arts, and have a solid foundation in reciting lines. The recorded voice timbre is natural, and the speaking speed is even.
Malaysian Female Speech Synthesis Corpus (Customer Service Style)
The database recorded 7,697 sentences (77,709 words) from a female voice talent. The total audio duration is about 10.38 hours, including the cleared silent segments at the beginning and ending (about 350 ms each). The recorded content is organized into 16 texts, including multiple fields, such as news, novel, dialog, etc. We used ms-my_xsampa phone set for labeling.
Maltese Female Speech Synthesis Corpus
The database recorded 2,000 sentences (19,821 words) from a female voice talent. The total audio duration is about 2.96 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is news text and dialog text. The voice talent was born and raised in Malta. She has a standard Maltese pronunciation and is a professional broadcaster. The recording has a warm timbre and even speech rate.
Maltese Male Speech Synthesis Corpus
The database recorded 2,000 sentences (19,776 words) from a male voice talent. The total audio duration is about 2.86 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is news text and dialog text. The voice talent was born and raised in Malta. He has a standard Maltese pronunciation and is a professional broadcaster. The recording has a steady timbre and even speech rate.
Mandarin 10-Speaker Song Synthesis Corpus (30 Songs)
The database recorded 887 sentences ( 6768 words) from 10 voice talents(5 males and 5 females). The total audio duration is about 1.92 hours (including the corresponding silence section of the accompaniment), about 1.34 hours for the vocal duration ( excluding the silent at the beginning and end). The database has 30 songs, each voice talent records one pop song and two bel canto songs.
Mandarin 100-Speaker Speech Synthesis Corpus
The database recorded 102,174 sentences (1,641,924 words) from 104 voice talents(54 females and 50 males). The total audio duration is about 124.90 hours, including the cleared silent segments at the beginning and ending (about 400 ms each). The recorded content is organized into 104 texts, including multiple fields, such as news, dialog,children's books, etc. We used zh-cn_pinyin phone set for labeling.
Mandarin 11-Speaker Multi-style Speech Synthesis Corpus (Virtual Lovers)
The database recorded 5,790 sentences ( 74,353 words) from 11 voice talents(5 males and 6 females). The total audio duration is about 5.12 hours, including the cleared silent segments at the beginning and ending (about 200 ms each). The recorded content is organized into 12 texts, with the scenarios set as virtual lovers' interactions and daily conversations.
Mandarin 6-Speaker Song Synthesis Corpus (300 Pop Songs)
The database recorded 3106 sentences ( 46388 words) from six non-professional voice talents(3 males and 3 females). The total audio duration is about 16.99 hours (consistent with the accompaniment), about 6.33 hours for the vocal duration (excluding the silent at the beginning and end). The recorded content is organized into 50 songs. All the voice talents are were born and grew up in China.
Mandarin Boy Speech Synthesis Corpus (Film Dubbing)
The database recorded 3,318 sentences (38,235 words) from a female voice talent. The total audio duration is about 3.01 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 1 text. We used zh-cn_pinyin phone set for labeling. The voice talent was born and raised in China, with standard Mandarin. She works in broadcasting and performance industry, with a good line foundation, imitating the voice of boys in this database .

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.

Filter by
Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios