TTS

Search our off-the-shelf datasets.

Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios
Bulgarian Male Speech Synthesis Corpus
The database recorded 1936 sentences (19311 words) from a male voice talent. The total audio duration is about 2.79 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is organized into 2 texts, including news, dialog. We used bg-bg_xsampa phone set for labeling. The voice talent is a professional broadcaster. He was native speaker in Bulgarian, with professional broadcasting background and steady voice.
Burmese Multi-speaker Speech Synthesis Corpus (Multi-emotion)
The database recorded 8,550 sentences from a total of four speakers: two male and two female. The total audio duration is approximately 14.47 hours, including the initial and final segments of clear silence (each about 300 milliseconds). Each speaker completed no less than 3 hours of recording, covering five emotions: neutral, happy, sad, angry, and surprised. Each speaker's recording duration for the neutral emotion is no less than 2 hours, while the happy, sad, angry, and surprised emotions are each no less than 0.15 hours. The recorded content is organized into 10 texts, covering a range of fields and emotions, such as news, encyclopedia, technology, conversation, novels, entertainment, happiness, sadness, anger, and surprise. The recordings were annotated using the my_mm_xsampa and en-us_cmu phoneme sets.
Canadian English Male and Female Multi-emotion Speech Synthesis Corpus (Free Talk)
The database recorded 3,034 sentences (83,905 words) from two male voice talents and three female talents. The total audio duration is about 8.56 hours, including the original silence at the beginning and ending (about 300ms each). The recorded content is organized into 5 texts, including 5 emotions. They are excited, sad, angry, fear and empathy. The five speakers were born and raised in Canada in 1958, 1962, 1982, 1993 and 2004, respectively. They all speak Standard Canadian English, and were between the ages of 21 and 68 when recording the database.
Canadian English Male and Female Multi-style Speech Synthesis Corpus
The database recorded 1421 sentences (31049 words) from a male voice talent and a female voice talent. The total audio duration is about 3.51 hours, including the original silence at the beginning and ending (about 300ms each). The recorded content is organized into 4 texts, including 4 fields. They are voice assistant, podcast, audiobook and Elearning. The female and male speakers were born in 1993 and 1958, respectively, with standard Canadian English. They were 31 and 68 years old when recording the database.
Canadian French Female Speech Synthesis Corpus
The database recorded 1,952 sentences (19,787 words) from a female voice talent. The total audio duration is about 2.34 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is general Canadian French text. The voice talent was born and raised in Quebec, Canada and has a standard Canadian French pronunciation. She studied in broadcasting and performance, with a good line foundation. The recording has a deep timbre and even speech rate.
Canadian French Male Speech Synthesis Corpus
The database recorded 2,011 sentences (20,967 words) from a male voice talent. The total audio duration is about 2.09 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is general Canadian French text. The voice talent was born and raised in Quebec, Canada and has a standard Canadian French pronunciation. He studied in broadcasting and performance, with a good line foundation. The recording has a deep timbre and even speech rate.
Czech Female Speech Synthesis Corpus
The database recorded 6,538 sentences (80,142 words) from a female voice talent. The total audio duration is about 10.86 hours, including the cleared silence at the beginning and ending (about 350 ms each). The recorded content is organized into 22 texts, including multiple fields, such as news, novel, dialog, etc. We used cs-cz_xsampa phone set for labeling.
Czech Male Speech Synthesis Corpus
The database recorded 6,526 sentences (79,915 words) from a male voice talent. The total audio duration is about 8.65 hours, including the cleared silence at the beginning and ending (about 350 ms each). The recorded content is organized into 22 texts, including multiple fields, such as news, novel, dialog, etc. We used cs-cz_xsampa phone set for labeling.
Danish Female Speech Synthesis Corpus
The database recorded 6,177 sentences (79,766 words) from a female voice talent. The total audio duration is about 8.41 hours, including the silence at the beginning and ending (about 350 ms each). The recorded content is organized into 20 texts, including multiple fields, such as news, digit, dialog, etc.

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.

Filter by
Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios