Score
Designs and builds evaluation benchmarks and curated datasets for egocentric (first‑person or body‑worn) video, including large‑scale, diagnostic, hardware‑authentic, and streaming evaluation collections. Tasks include specifying data‑collection protocols and hardware‑authentic capture settings, choosing label taxonomies and high‑stakes action classes, creating annotation schemas with precise timestamps or per‑second action labels and QA pairs, and producing benchmark tasks, splits, and evaluation metrics for localization, semantic reasoning, and diagnostic analysis.
Existing video models struggle to accurately recognize critical actions—such as drawing a weapon—in high-risk, highly dynamic body-worn police camera scenarios, lacking fine-grained understanding of real-world law enforcement interactions. To address this gap, this work introduces the first egocentric video benchmark tailored to high-stakes policing contexts, constructed from rigorously curated real officer-civilian encounters. The dataset features second-level annotations of rare yet pivotal enforcement behaviors and includes classification and multiple-choice question-answering tasks to evaluate model performance under conditions of intense motion and complex contextual dynamics. Experiments reveal that even state-of-the-art video foundation models, such as Gemini 2.5 Pro, exhibit significant shortcomings on these tasks, underscoring the benchmark’s difficulty and establishing a technical foundation for efficient human-in-the-loop review of执法 footage.
This work addresses the fragmentation in first-person video capture caused by heterogeneous devices and the lack of a unified, low-cost cross-platform solution. The authors propose an open-source toolkit supporting six device categories—including smartphones, tablets, smart glasses, and XR headsets—that enables consistent data collection through a cross-platform SDK, a standardized video interface, 26-joint hand tracking compliant with OpenXR, USB-C peripheral expansion, and a uniform logging format. For the first time, the system synchronously captures high-quality egocentric videos from both eye- and wrist-mounted perspectives across diverse hardware, while XR devices additionally provide aligned head pose and hand-tracking data. This approach substantially enhances compatibility and reproducibility in first-person data acquisition.
This work addresses the scarcity of large-scale, temporally synchronized multi-view data for first-person dynamic scene reconstruction and novel-view synthesis. To this end, we introduce Ego-1K, a dataset comprising nearly 1,000 high-synchronization first-person multi-view video sequences captured using a custom VR headset equipped with 12 surrounding cameras and 4 head-mounted cameras, focusing on close-range dynamic interactions such as hand motions and hand-object manipulations. Ego-1K provides the first large-scale, high-quality benchmark tailored to complex egocentric interactions, ensured by precise spatiotemporal calibration and an automated processing pipeline for data alignment and usability. Experiments demonstrate that Ego-1K poses significant challenges to existing 3D/4D novel-view synthesis methods, effectively exposing their limitations under large viewpoint disparities and rapid motion, thereby establishing a critical foundation for future research.
Current vision-language models lack a unified benchmark for evaluating multimodal temporal reasoning—spanning retrospective, online, and prospective understanding—in first-person streaming video. To address this gap, this work proposes EgoSAT, the first comprehensive evaluation framework tailored to egocentric video streams, integrating 165 hours of video (1,997 clips) and approximately 4,800 high-quality question-answer pairs. EgoSAT further introduces an assessment mechanism that explicitly accounts for question answerability. Experimental results reveal that existing models exhibit significant weaknesses in both prospective and retrospective reasoning and consistently suffer from poor confidence calibration—often being highly confident yet incorrect. These findings provide critical diagnostic insights and establish a foundational benchmark for advancing research in streaming egocentric vision-language understanding.
Existing first-person video benchmarks struggle to evaluate models’ fine-grained, action-centric reasoning capabilities and lack mechanisms to verify whether such reasoning is grounded in explicit spatiotemporal evidence. To address this gap, this work proposes the first multimodal large language model benchmark that supports verifiable, fine-grained chain-of-action reasoning. Leveraging a spatiotemporal scene graph (STSG)-guided data generation framework augmented with expert annotations, the authors construct a dataset of 3,172 question-answer pairs spanning perception, retrospection, prediction, and high-level reasoning. Experiments reveal that while current models often produce correct answers, their explanations frequently misalign with actual evidence, underscoring the benchmark’s unique value in assessing reasoning groundedness and providing a reliable platform for future research.
This work addresses the high cost and limited scalability of existing first-person data collection systems by proposing and open-sourcing a low-cost (under $200 per unit), extensible head-mounted stereo visual-inertial sensing platform. The system integrates hardware-synchronized global-shutter stereo cameras, a 6-axis IMU, an embedded Linux board, and a real-time microcontroller, enabling high-quality crowdsourced data acquisition in unstructured environments. It features a complete software-hardware stack, including hardware-accelerated recording, an IMU sampling daemon, and precise time synchronization. The project contributes approximately 550 hours of synchronized stereo video with IMU data, annotated with full-temporal free-form action captions and per-frame 3D hand reconstructions—marking the first open-source integration of affordable hardware, real-time feedback, accurate time synchronization, and large-scale annotations.
Existing egocentric video datasets struggle to effectively capture users’ internal states—such as intent, emotion, and memory—thereby limiting the natural interaction capabilities of AI assistants. To address this gap, this work introduces EgoIntrospect, the first user-driven multimodal dataset, which synchronously collects video, audio, eye-tracking, motion, and physiological signals across devices and incorporates user-provided self-annotations to reveal subjective states during human–AI interactions. Leveraging this dataset, we establish a benchmark for internal state inference tailored to multimodal large language models. Experimental results demonstrate that current models still face significant challenges in accurately inferring internal states through effective fusion of multimodal signals. This study provides a comprehensive resource comprising 180 hours of data from 60 participants, along with an evaluation framework, thereby filling a critical void in the field.
Existing benchmarks struggle to jointly evaluate the multimodal perception, multi-hop tool use, and dynamic human–agent interaction capabilities required by AI agents in open-ended environments. To address this gap, this work introduces an interactive multimodal benchmark grounded in first-person videos, encompassing 1,045 everyday tasks within user–agent–tool interaction scenarios. The benchmark enforces the integration of visual perception and tool-augmented multi-hop reasoning through a three-stage collaborative reasoning pipeline and employs multi-agent simulations to generate high-fidelity user feedback. It presents the first framework capable of jointly assessing these three core competencies, featuring a deterministic joint validation protocol and a dual-dimensional evaluation mechanism that considers both process and outcome equivalence. Evaluations of eight state-of-the-art video-based multimodal large language models across four everyday scenarios reveal a best accuracy of only 30.62% (average: 19.43%), highlighting significant capability bottlenecks in current agents.
Existing video benchmarks struggle to evaluate the capability of vision-language models in recognizing worker behaviors and reasoning about safety rules under real-world industrial surveillance conditions—such as low illumination, occlusion, and long-range viewing. This work proposes SteelBench, the first multidimensional diagnostic benchmark tailored to authentic steel plant environments. Constructed from 149 hours of surveillance footage, it comprises 1,345 densely annotated video clips curated via temporal deduplication, category balancing, and visibility-aware sampling. The dataset encompasses actions, personal protective equipment (PPE) attributes, spatial context, and explicit safety rules, and introduces a novel annotation provenance auditing mechanism. Experiments reveal that even the best-performing model achieves only 42.6% accuracy on action recognition (versus 84.6% for humans), exhibits safety judgment error rates of 37–58%, and fails to pass more than two diagnostic tests. Moreover, unaudited model-generated labels can inflate reported accuracy by up to 17 percentage points for related models.
Current benchmarks struggle to independently evaluate agents’ ability to select actions based solely on egocentric perspectives in multi-agent scenarios. This work introduces a new capability—Ego-centric Action Selection (EAS)—and presents EgoGapBench, a diagnostic benchmark designed to isolate first-person visual inputs from confounding influences of others’ behaviors, thereby specifically assessing an agent’s capacity to make reasonable action decisions using only its own viewpoint. Experimental results show that humans perform robustly on this benchmark, whereas both open- and closed-source multimodal large language models exhibit substantially lower accuracy and frequently misattribute actions to themselves that are actually performed by others. Fine-tuning exclusively on EgoGapBench yields modest performance gains but remains far below human-level competence, revealing a systematic deficiency in current models’ EAS capabilities.