Text corpora in Simplified Cantonese

This dataset is collecting Cantonese dialect of Guangdong
Specifications:
ID:
King-NLP-175
Size:
101931 sentences
Language:
Cantonese
Size
100000 sentences
Country
China
Language Code
ZH-CN
Accuracy Rate
Sentence Accuracy 95%

People also searched for

Chinese-English Parallel Corpus
This dataset is Chinese-English parallel corpus
English-Chinese Parallel Corpus
This dataset is English-Chinese parallel corpus
Chinese-Malay Parallel Corpus
This dataset is Chinese-Malay parallel corpus
English-Korean Parallel Corpus
This dataset is English-Korean parallel corpus

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.