TTS

Search our off-the-shelf datasets.

Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios
Mandarin Capable Woman Speech Synthesis Corpus
The database recorded 5913 sentences (89497 words) from a female voice talent. The total audio duration is about 6.48 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is organized into 17 texts, including multiple emotions, such as happy, angry, sad, surprise, hate, fear, etc. We used zh-cn_pinyin phone set for labeling. The voice talent was born and raised in China, and was 40 years old when recording the database. She has a standard Mandarin pronunciation and is a professional broadcaster. The role she imitated in the recording was a capable white-collar lady.
Mandarin Ethereal Female Song Synthesis Corpus (205 Pop Songs)
The database recorded 1799 sentences (36462 words) from a female voice talent, including 205 songs (1505 sentences) and 294 sentences. The total audio duration is about 15.09 hours, including 14.09 hours for the song section (consistent with the accompaniment), and about 3.31 hours for the vocal duration (excluding the silence at the beginning and end); the voice section is about 1 hour, including the silence at the beginning and ending (about 200 ms each). The recorded content is organized into 2 texts, the first part are pop songs which are recorded along with the accompaniment, but sing one verse and one chorus unaccompanied. And the second part is free talking about various topics.
Mandarin Female Speech Synthesis Corpus
The database recorded 19,509 sentences (185,766 words) from a female voice talent. The total audio duration is about 15.11 hours, around 14.64 hours in Chinese, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 26 texts, mainly of general type, including commonly used sentences, numbers and letters. We used zh-cn_pinyin & en-us_CMU phone set for labeling.
Mandarin Female Speech Synthesis Corpus
The database recorded 22546 sentences (329141 words) from a female voice talent, including 20156 Chinese sentences, 2390 English and mixed Chinese-English sentences. The total audio duration is about 29.08 hours, including 25.64 hours in Chinese, 3.44 hours in English and 0.34 hours in mixed Chinese-English, including the original silence at the beginning and ending (about 350 ms each). The recorded content is organized into 28 texts, including multiple fields, such as news, time, dialogue etc. We used zh-cn_pinyin & en-us_CMU phone set for labeling.
Mandarin Female Speech Synthesis Corpus (Audiobook)
The database recorded 7219 sentences (133557 words) from a female voice talent. Among them, 6404 sentences were recorded into 47 pieces of audio in paragraph format. The total audio duration is about 11.27 hours, including the original silence at the beginning and ending (about1.5 seconds each). The recorded content is organized into a single script, which includes excerpts from novels along with additional general-purpose sentences. We used zh-cn_pinyin phone set for labeling. The voice talent was born and grew up in Shaanxi, China in 1981, and speaks standard Mandarin. As a professional broadcasting practitioner, the speaker has a solid foundation in delivering lines. The recorded voice has a warm and clear timbre, with a steady and even speaking pace.
Mandarin Female Speech Synthesis Corpus (Elderly Empress)
The database recorded 6019 sentences (102731 words) from a female voice talent. The total audio duration is about 11.08 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is organized into 17 texts, including multiple emotions, such as happy, angry, sad, surprise, hate, fear, etc. We used zh-cn_pinyin phone set for labeling. The voice talent was born and raised in China, and was 40 years old when recording the database. She has a standard Mandarin pronunciation and is a professional broadcaster. The role she imitated in the recording was the old queen mother.
Mandarin Female Speech Synthesis Corpus (Funny Game Streamer)
The database recorded 722 sentences (16111 words) from a female voice talent. The total audio duration is about 1.06 hours, including the cleared silence at the beginning and ending (about 300 ms each). The recorded content is organized into 1 texts. We used zh-cn_pinyin and en-us_cmu phone set for labeling. The voice talent is a professional broadcaster and speaks standard Mandarin. She was born in China, with professional broadcasting background and sweet and lively voice.
Mandarin Female Speech Synthesis Corpus (Interview of Customer Service Style)
The database recorded 1010 sentences (41529 words) from a female voice talent. The total audio duration is about 3.37 hours, including the cleared silent segments at the beginning and ending (about 200 ms each). The database designed 157 questions to conduct interviews with the voice talent. The female voice talent is a Chinese professional dubbing actress, with a friendly and natural voice.
Mandarin Female Speech Synthesis Corpus (Livestreaming Style)
The database recorded 3,587 sentences (90,045 words) from a female voice talent. The total audio duration is about 5.06 hours, including the original silence at the beginning and ending (about 150 ms each). The recorded content is organized into 1 text, focusing on live streaming for product promotion. We used zh-cn_pinyin & en-us_cmu phone set for labeling. The voice talent was born and grew up in Inner Mongolia, China, and was 21 years old when recording the database. She has standard Mandarin pronunciation and a good command of English. She is a practitioner in the broadcasting industry. The recording has a positive and amiable timbre, with a relatively fast speech rate.

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.

Filter by
Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios