Lip-reading Speech Video Corpus

The corpus uses six cameras and two microphone arrays to simultaneously capture the lip speech video data of speakers. The capture and filming scenario simulates the interior of a cockpit, with diverse shooting angles and lighting. Data is collected from 250 individuals, all of whom are adults, primarily middle-aged and young people. Each person's target effective recording time is approximately 0.5 hours, with an average of about 600 short sentences per person. The product library also extracts audio from any one of the six video routes captured for each ID, saving it as a separate audio file. The results from the six cameras will be aligned with an error of less than 30 milliseconds, and the two microphone results will also be synchronized with the camera results.
Specifications:
ID:
King-VD-018
Size:
1.05 TB

People also searched for

High-Definition Dance Video Corpus
Product Features: This dataset has collected 100,000 dance videos, each averaging 30 seconds in length, at 4K resolution, including adults and teenagers with a foundation in dance, with a balanced gender ratio. It includes both solo and group dances, with high richness in videos from various angles such as front, side, back, and turning. Dance types include folk dance, jazz, street dance, and more. Application Fields: This dataset can be applied to virtual humans, VR, dance education, video production, and other fields, promoting the application and development of multimodal technology in the corresponding areas.
Telephoto Landscape Corpus
【Product Features】 High-quality images of architecture and plants, with no blurring within the full size of the image, ensuring that both the foreground and background show clear textures even when enlarged; no more than 5 images of the same subject from different angles to ensure diversity in the content captured. 【Image Specifications】 Resolution above 4k (shoot in the highest quality mode with the camera); focal length within the range of 185mm to 235mm.
DMS with Multi-skin color Drivers Corpus
【Collector Information】 Ethnicity: Divided into two categories, Black and White. Among them, Black includes (black, brown, olive), and White includes (fair, medium, very fair). Nationality: Involves more than 39 countries (Switzerland, Colombia, Peru, Paris, Ghana, Brazil, Latin America, Latvia, Samoa, South Africa, etc.). Age: Collectors cover the age range of 18-60+, with a majority being middle-aged and young adults. 【Video Information】 Each video segment is at least 20 seconds long, with a resolution of no less than 720P.【Data Collection Information】 Daytime: Includes (front lighting, backlighting, side lighting, dappled sunlight, overcast, rainy, snowy weather) Nighttime: Includes (interior vehicle lighting, street lamp lighting, oncoming vehicle high and low beam) Facial Expressions and Actions: Eyes open, mouth open and closed, exaggerated mouth open and closed, exaggerated expressions, mouth twisted, making faces, etc. Other Actions: (smoking, drinking water, using a mobile phone, hand occlusions, etc.) Accessories: All subjects wear accessories, including (glasses, hats, etc.)
Object Segmentation

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.