Multimodal

Search our off-the-shelf datasets.

Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios
AIGC-Portrait image dataset
The images of human figures in this dataset have been transformed into four different styles based on the original human images, including 3D cartoon, comic, watercolor painting, and sketch. There are four skin colors included: dark, fair, brown, and light. Each style-generated image contains all four skin colors.
Conference video action collection dataset
This dataset is a simulated video dataset of meeting actions. It captures various skin colors including dark, brown, fair, and light, recorded in a bright meeting room environment, with a single person's frontal video, using 4 different collection devices. Each person has 4 video clips and 2 pictures collected.
Lip speech video was collected for 250 people
This dataset uses six cameras and two microphones to simultaneously collect and record the lip-voiced video data of the speaker. The shooting scene is simulated in a cockpit environment, and the shooting angles and lighting conditions are diverse.
Lip-movement dataset
This dataset uses cameras to collect audio and video data of lip movements. The shooting scene is an indoor quiet environment, and various light conditions are simulated, including normal light, strong light, backlight, and weak light, with the shooting distance ranging from 0.5m to 1m, with 0.5m accounting for approximately 90%. The shooting angle is frontal, and the image size mainly covers the upper body. In addition to single-person collection, the collection also simulates queuing scenarios. Approximately 30% of the data in each person's video is collected by multiple people, and in the multi-person scenarios, the number of people on screen is mostly 2. The collectors mainly speak Mandarin at a normal speed, and some collectors may have slight local accents. Each person records 10 sentences of text, with an average of 10 to 15 words per sentence. The ages of the collectors cover multiple age groups, with a majority being children and middle-aged people, and the gender ratio is balanced. While the video is being recorded, there is also a front-facing interface microphone recording synchronously, and the audio file comes from the collected video.
Multi-pose facial video dataset
This dataset collects video data of the head posture and expressions of human figures. The collection was conducted in various indoor living and working scenarios such as offices, meeting rooms, homes, dormitories, and corridors. Each person was filmed for a video, with the human figure in the video approximately the size of a headshot. The video content included raising and lowering the head, moving left and right, shaking the head, opening and closing the mouth, and various combinations of posture and movements. Each video was approximately 1 minute long. The lighting conditions included normal, low light, and backlighting, and the human figures were clearly visible.
Multimodal 3D Sign Language Dataset
The hand gesture vocabulary domain covered by this dataset includes national common hand gesture vocabulary, sports, place names, computers, fine arts, physics, etc. The hand gesture method adopted was collected in accordance with the guidance of common vocabulary of national common hand gesture, and the collection equipment was an inertial motion capture device.

Join our newsletter to stay updated

Thank you for signing up!

Stay informed and ahead with the latest updates, insights, and exclusive content delivered straight to your inbox.

By subscribing you agree to with our Privacy Policy and provide consent to receive updates from our company.

Filter by
Filter by
Language
Filter by Languages
Language
Devices
Devices
Applicable Fields
Applicable Fields
More
Applicable Scenarios
Applicable Scenarios