This dataset consists of 11 hours of recordings with a balanced gender ratio. It has been meticulously labeled, including pronunciation, prosody, and voice quality labeling. The voice samples, recorded by non-professional speakers, offer a higher degree of naturalness and are categorized and labeled according to voice gender, perceived age, voice description, vocal cord condition, and pronunciation location. The topic includes multi-emotional data, covering emotions such as calm, happy, angry, sad, and more.
Chinese Female Speech Synthesis Corpus – Live Streaming for Sales with Multi Styles
【Features】Two styles: Deep and uplifting; covers a variety of product categories including food, clothing, beauty, personal care, electronics, and home goods.