This dataset was recorded by a 30-year-old female speaker with authentic pronunciation and a friendly, soft vocal quality in a professional recording studio. The recorded texts span the full range of phonemes, and the annotators have a professional linguistic background, ensuring the data meets the research and development needs for voice synthesis.
Mandarin Chinese with Natural Style Voice Synthesis Corpus
The dataset includes 237 speakers, covering a variety of voice qualities such as mature female, middle-aged male, bass, falsetto, etc., and spans across young, middle-aged, and elderly ages. The audio is clear and natural, which can greatly enhance the naturalness and expressiveness of the model.
Mandarin Chinese with Multi-Emotion & Multi-Timbre Synthesis Corpus
The dataset includes 142 distinctive speakers with various emotions such as happiness, sadness, anger, surprise, calmness, dislike, fear, etc.; it can greatly enhance the naturalness and expressiveness of the model.
Chinese Multi-speaker – Amateur Multi-Emotion, Multi-Style
This dataset consists of 11 hours of recordings with a balanced gender ratio. It has been meticulously labeled, including pronunciation, prosody, and voice quality labeling. The voice samples, recorded by non-professional speakers, offer a higher degree of naturalness and are categorized and labeled according to voice gender, perceived age, voice description, vocal cord condition, and pronunciation location. The topic includes multi-emotional data, covering emotions such as calm, happy, angry, sad, and more.