π€ AI Summary
This work addresses the scarcity of event-based datasets capturing unscripted human activities in naturalistic settings by introducing EventKitchen, a novel multimodal dataset recorded in real-world kitchen environments. Using head-mounted sensors, the dataset captures stereoscopic event streams from ten participants engaged in spontaneous cooking tasks across thirteen kitchens, synchronized with RGB, depth, and IMU data. EventKitchen provides the first large-scale, human-centric, and unscripted event-camera recordings of everyday kitchen activities, comprising 5.5 hours of stereo event data, 10,762 action annotations, and 13,482 object bounding boxes. It supports benchmarking for action recognition, object detection, and stereo depth estimation, and includes baseline model evaluations to establish a new standard for neuromorphic vision in daily-life perception.
π Abstract
Event cameras, also known as neuromorphic cameras, have gained significant attention in recent years due to their high temporal resolution, high dynamic range, and low power consumption. While many studies and datasets in neuromorphic vision have focused on automotive and drone applications, human-centric daily-life scenarios remain largely underrepresented, despite their importance for developing and benchmarking event-based perception systems. Moreover, the few existing event-based human activity datasets are typically recorded with scripted human actions, limiting their ability to capture natural human behaviors. In this paper, we introduce EventKitchen, a large-scale stereo event camera benchmark dataset of human cooking activities in the kitchen. EventKitchen is egocentrically collected from 10 participants in 13 diverse kitchens, where the participants wear a helmet with multiple sensors and naturally perform cooking activities, without any scripted actions. EventKitchen comprises 5.5 hours of stereo event recordings with synchronized RGB, depth, and IMU data. We provide human annotations for 10,762 action segments and 13,482 bounding boxes. We train baseline models on EventKitchen to perform multiple event-based tasks, including action recognition, object detection, and stereo depth estimation. By capturing natural, real-world human activities, EventKitchen establishes a challenging benchmark for neuromorphic vision beyond autonomous driving.