All Datasets

Search our off-the-shelf datasets.

Filter by
Category
Category
2D Flat Anti‑Spoofing Attack Video Dataset
This dataset collects live attack video data of Chinese people. The collection scenarios cover indoor and outdoor environments and various light conditions. The shooting distance will be dynamically adjusted according to the actual shooting situation, mainly changing between full-body shots and head-and-shoulder shots. The types of attack methods collected include static attacks and dynamic attacks, totaling 7 categories. Static attack forms include photos, A4 color prints, and black-and-white prints. Dynamic attack forms include playing videos on iPads, Android phones, iPhones, and laptops. Each person will shoot 6 live videos and 22 attack videos in a normal and natural state, with each video lasting approximately 15 seconds. The collection tools are collected using a handheld method.
3D Face & Body Keypoint Annotated Dataset
This dataset was collected using a 3D scanning device, which captured the facial and limb movements of the model wearing multiple pieces of clothing in an indoor setting. The 3D data was annotated with 57 key points for the face and 62 key points for the limbs. The collection was conducted using a 3D scanner, with each person completing the designated movements indoors. Each person was photographed and scanned in 16 postures, including 2 standard T-pose postures (tight clothing and personal attire). There were 62 key points for the human body and 57 key points for the face. The age range covered 7 to 60 years old, with a relatively balanced gender ratio, and the fluctuation was within 10%.
3D Hand Gesture Keypoint Annotated Dataset
This dataset was collected in the XR perspective (first-person perspective) and contains 3D hand joint point data (21 points). It consists of two parts: static gestures and dynamic gestures. Both types of gestures include the left and right hands, totaling 22 different gestures. Moreover, the collection covers hands of different sizes.
3D modeling data or market sence
The data collection equipment for this dataset is a 3D space scanning device. It was collected in Chinese supermarkets. Using a 3D space scanning camera device, 10 complete 3D scene data of medium and small-sized supermarkets were collected, and a 3D point cloud data model was generated for each supermarket.
50K Multi-Skin-Tone High-Definition Face Dataset
This dataset contains facial image data. It includes three different skin colors: dark, light, and others (excluding fair). All the subjects in the collection are adults, with an equal male-to-female ratio. The image resolution is all above 1024×1024. This dataset covers common facial expressions, including natural expressions, happiness, excitement, etc. The facial angles mainly depict frontal views with the smallest deviation (yaw/pitch/roll all less than 30 degrees). The images are real, clear, without gray scale, and without watermarks.
840 person image collection by front-facing camera and face 21points labeling
This dataset is a collection and annotation library of multi-posture self-portrait pictures taken by Chinese people in natural scenes. Each collector has taken pictures in at least 4 different scenarios, including 3 indoor and 1 outdoor ones. Under different background environments, lighting conditions and distances, each collector completed 11 types of self-portraits while wearing glasses and without wearing glasses respectively. Each person took 100 pictures and 88 designated actions, as well as 12 actions that were freely performed by themselves. The collected portrait pictures were labeled with 21 key points.
9 Languages OCR Dataset
This dataset includes various categories of scenes, including Attraction、Cards、Document、Manual、Menu、Packing、Screenshot、Street sign
Accented English Pronunciation Evaluation Corpus (Word Level)
This database is collected over Mobile phones in quiet (office/home) environment, which were from 22 speakers, including 11 male and 11 female. The total pure recording time is about 11.38 hours, including the reasonable leading and trailing silence.
Afghani Dari Conversational Speech Recognition Corpus (Telephone)
This database is collected over Telephone in quiet (office/home) environment, which were from 40 speakers, including 20 male and 20 female. The total pure recording time is about 40 hours, including the reasonable leading and trailing silence.

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.

Filter by
Category
Category