CV

Search our off-the-shelf datasets.

Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios
More
Image Segmentation Corpus
Indonesia Image Corpus
Indonesia Life Image Corpus
Indoor Scenes Image Collection
Indoor Tracking Video Corpus
Lip-movement Video Corpus
The corpus uses high-definition cameras to capture lip speech video data from approximately 208 individuals. The capture scenario is an indoor quiet environment, simulating various types of lighting, including normal light, strong light, backlight, and weak light. The shooting distance includes 0.5m and 1m, with a primary focus on 0.5m, accounting for about 90% of the recordings. The shooting angle is frontal, with the imaging size focusing mainly on the upper body. In addition to solo collections, the collection also simulates queue scenarios, with about 30% of each person's video data being collected in multi-person scenarios, where the number of people appearing in the multi-person scenes is mostly two. The collectors primarily speak Mandarin (prioritizing northern pronunciation individuals, with some collectors having better southern Mandarin pronunciation), some collectors may have a slight local accent, speaking at a normal pace, recording 10 sentences per person, with an average of 10 to 15 characters per sentence. The collectors' ages range from 7 to over 60 years old, mainly children and middle-aged and young people, with a balanced gender ratio. While the video is being recorded, there is also a front-facing interface microphone recording synchronized with the collector, and the other audio file comes from the collected video.
Lip-reading Speech Video Corpus
The corpus uses six cameras and two microphone arrays to simultaneously capture the lip speech video data of speakers. The capture and filming scenario simulates the interior of a cockpit, with diverse shooting angles and lighting. Data is collected from 250 individuals, all of whom are adults, primarily middle-aged and young people. Each person's target effective recording time is approximately 0.5 hours, with an average of about 600 short sentences per person. The product library also extracts audio from any one of the six video routes captured for each ID, saving it as a separate audio file. The results from the six cameras will be aligned with an error of less than 30 milliseconds, and the two microphone results will also be synchronized with the camera results.
Masked Facial Image Corpus
Multi Angle Facial Corpus

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.

Filter by
Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios
More