Chinese Personal SMS Corpus with Word Segmentation

This dataset is Chinese text data, with word segmentation annotation
Specifications:
ID:
King-NLP-004
Size:
1266486 items
Language:
Chinese
Quantity
This dataset contains 1266486 sets.
Data format
TXT
Country
China
Language Code
ZH-CN
Accuracy Rate
Word Accuracy 97%

People also searched for

Chinese-English Parallel Corpus
This dataset is Chinese-English parallel corpus
English-Chinese Parallel Corpus
This dataset is English-Chinese parallel corpus
Chinese-Malay Parallel Corpus
This dataset is Chinese-Malay parallel corpus
English-Korean Parallel Corpus
This dataset is English-Korean parallel corpus

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.