HK Cantonese Text Corpus with POS Tagging

This dataset is given Cantonese text, perform word segmentation and part-of-speech tagging
Specifications:
ID:
King-NLP-056
Size:
300000 sentences
Language:
Cantonese
Quantity
The corpus contains 300,000 entries.
Country
China
Language Code
CT-HK
Accuracy Rate
Word Accuracy 97%

People also searched for

Chinese-English Parallel Corpus
This dataset is Chinese-English parallel corpus
English-Chinese Parallel Corpus
This dataset is English-Chinese parallel corpus
Chinese-Malay Parallel Corpus
This dataset is Chinese-Malay parallel corpus
English-Korean Parallel Corpus
This dataset is English-Korean parallel corpus

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.