All Datasets

Search our off-the-shelf datasets.

Filter by
Category
Category
Afghanistan Dari Pronunciation Lexicon
This Afghanistan Dari Pronunciation Lexicon, curated by DataoceanAI Inc., offers a wealth of linguistic resources tailored specifically for the Dari language as spoken in Afghanistan. With 30,075 meticulously crafted entries and an impressive 95.00% entry accuracy rate, this lexicon provides accurate pronunciation transcription in the popular XSAMPA phonemic system. It serves as indispensable training data for speech recognition, speech synthesis, and other language processing applications.
Afghanistan Pashto Pronunciation Lexicon
This Afghanistan Pashto Pronunciation Lexicon, curated by DataoceanAI Inc., offers a wealth of linguistic resources tailored specifically for the Pashto language as spoken in Afghanistan. With 50,170 meticulously crafted entries and an impressive 95.00% entry accuracy rate, this lexicon provides accurate pronunciation transcription in the popular XSAMPA phonemic system. It serves as indispensable training data for speech recognition, speech synthesis, and other language processing applications.
African American English Speech Recognition Corpus (Mobile)
This database is collected over Mobile phones in quiet (office) environment, which were from 200 speakers, including 100 male and 100 female. The total pure recording time is about 120.5 hours, including the reasonable leading and trailing silence.
Afrikaans Speech Recognition Corpus (Mobile)
This database is collected over Mobile phones in quiet (office/home) environment, which were from 402 speakers, including 88 male and 314 female. The total pure recording time is about 235 hours, including the reasonable leading and trailing silence.
Afrikaans Speech Recognition Corpus (Mobile)
This database is collected over Mobile phones in quiet (home) environment, which were from 125 speakers, including 124 male and 1 female. The total pure recording time is about 30 hours, including the reasonable leading and trailing silence.
AIGC-Portrait image dataset
The images of human figures in this dataset have been transformed into four different styles based on the original human images, including 3D cartoon, comic, watercolor painting, and sketch. There are four skin colors included: dark, fair, brown, and light. Each style-generated image contains all four skin colors.
Albania Albanian Pronunciation Lexicon
This Albania Albanian Pronunciation Lexicon, curated by DataoceanAI Inc., offers a wealth of linguistic resources tailored specifically for the Albanian language as spoken in Albania. With 53,542 meticulously crafted entries and an impressive 95.00% entry accuracy rate, this lexicon provides accurate pronunciation transcription in the popular XSAMPA phonemic system. It serves as indispensable training data for speech recognition, speech synthesis, and other language processing applications.
Albanian Conversational Speech Recognition Corpus (Mobile)
This database is collected over Mobile phones in quiet (office/home) environment, which were from 20 speakers, including 9 male and 11 female. The total pure recording time is about 22 hours, including the reasonable leading and trailing silence.
Albanian Speech Recognition Corpus (Mobile)
This database is collected over Mobile phones in quiet (office/home) environment, which were from 400 speakers, including 206 male and 194 female. The total pure recording time is about 223.59 hours, including the reasonable leading and trailing silence.

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.

Filter by
Category
Category