This dataset contains a collection of adult and teenage facial expression images. The indoor scenes feature plain colors or other patterns as the background fabric, and three cameras are fixed to capture the images, with different angles of elevation. Each collector performs 14 facial expressions. The collector's race is of Asian descent, with the age range from 13 to 80 years old, and the gender ratio is roughly balanced. The shooting distance is head and shoulder shots or half-body shots.
This dataset covers more than 10 types of scenarios, including residential settings, offices, retail stores, restaurants, manufacturing facilities, warehouses, agricultural environments, and entertainment venues. It includes over 200 tasks and provides rich first-person action and human–environment interaction data. The videos have a resolution of 1920 × 1080 and a frame rate of 60 fps. The dataset can be used for embodied AI model training, egocentric behavior understanding, environmental perception and interaction, and other related applications.
Hospital Multi‑Scene Patient Action Simulation Video Dataset II
This dataset consists of video data collected in a simulated hospital environment, covering approximately 10 different scenarios. It includes shots of at least three types of different hospital beds. The subjects are shown in three postures: lying on the bed, sitting on the bed, and standing outside the bed. Each model is filmed three times. The first time, they are filmed in their own clothes. The second time, they change into hospital gowns (usually white or blue) for the shot. The third time, they can freely express themselves in the hospital gowns.
Hospital Multi‑Scene Patient Action Simulation Video Dataset I
This dataset consists of video data collected in a simulated hospital environment. The collectors captured the three postures of the subjects: lying on the bed, sitting on the bed, and standing outside the bed. A total of 118 individuals were included, and each person had 3 video segments collected. The resolution of the videos included 720P and 1080P.
This dataset collects various images and textual descriptions of human behaviors in different indoor and outdoor scenarios for people of multiple skin tones, covering common facial expressions and a wide range of body movements. It also includes multiple images and descriptions of human actions from different shooting angles and across different age groups (all adults).