American pop Songs Speech Synthesis Corpus (150 Songs)

The database recorded 4,940 sentences (72,059 words) from 5 speakers (4 females and 1 male). The total audio duration is about 13.30 hours. The recording content is divided into two parts: singing and reading. The singing part includes 150 American pop songs (36,193 words). The songs are recorded in their entirety, with a total duration of approximately 8.62 hours, including the cleared silent segments at the beginning and end (the silent segment at the beginning corresponds to the length of the song's instrumental prelude), and the dry voice duration is 6.32 hours. The reading part consists of 4,790 sentences (35,866 words), with a total duration of approximately 4.68 hours, including the silent segments at the beginning and end (each approximately 350 ms). We used en-us_cmu phone set for labeling.
Specifications:
ID:
King-TTS-126
Size:
11 hours
Language:
English
Sample rate & bit depth
48 kHz, 24bit
Recording environment
Professional recording studio
Speaker
1 male and 1 female
Devices:
Studio
Accuracy Rate
proofreading -- based on individual word, the accuracy is 99% phonetic labeling -- based on individual phone, the accuracy is 99.5% phone labeling -- the boundary error is less than 10 ms, the accuracy is 99%
Samples
Audio
King-TTS-126-F01001001
King-TTS-126-F02001
King-TTS-126-F08001
King-TTS-126-M04001001

People also searched for

American English Male and Female Speech Synthesis Corpus (Customer and Audiobook)
This database contains 2000 sentences from one female speaker and one male speaker, with a total audio duration of approximately 2 hours. The texts include customer and audiobook field.
Brazilian Portuguese Male and Female Speech Synthesis Corpus
The database recorded 2,924 sentences (49,025 words) from 3 voice talents(2 females and 1 male). The total audio duration is about 6 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 11 texts, F048-03 including multiple fields, such as news, letters, digit, etc. We used pt-BR_xsampa phone set for labeling. The voice talents were born and raised in Brazil, in 1969/1973/1997, with standard Brazilian Portuguese and were 56/51/48 years old when recording the database, with a good line foundation. The recordings have even speech rate.
Spain Spanish Male and Female Speech Synthesis Corpus
The database recorded 3,068 sentences (52,494 words) from 3 voice talents(2 females and 1 male). The total audio duration is about 6 hours, including the original silence at the beginning and ending (about 300 ms each). The recorded content is organized into 11 texts, F021-04 including multiple fields , such as news, letter, digit, etc. We used es-es_sampa phone set for labeling. The voice talents were born and raised in Spain in1962/1984/1987, with standard Spanish, and were 63/41/37 years old when recording the database, with a good line foundation. The recording have even speech rate.
New Zealand English Female Speech Synthesis Corpus
The database recorded 1,600 sentences (17,808 words) from a male voice talent. The total audio duration is about 2.04 hours, including the original silence at the beginning and ending (about 350 ms each). The recorded content is organized into 1 texts, news. The voice talent was born and raised in New Zealand in 1989, with standard New Zealand English. She is a professional voice talent who has many years of experience in dubbing and broadcasting , with a good line foundation.

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.