Traditional Chinese SMS Corpus with Word Segmentation

This dataset is Taiwanese traditional Chinese text message data, segmented and labeled with words
Specifications:
ID:
King-NLP-039
Size:
881524 sentences
Language:
Chinese
Quantity
The corpus contains 881524 entries.
Country
Taiwan, China
Language Code
ZH-TW
Accuracy Rate
Word Accuracy 97%

People also searched for

Chinese-English Parallel Corpus
This dataset is Chinese-English parallel corpus
English-Chinese Parallel Corpus
This dataset is English-Chinese parallel corpus
Chinese-Malay Parallel Corpus
This dataset is Chinese-Malay parallel corpus
English-Korean Parallel Corpus
This dataset is English-Korean parallel corpus

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.