Event ActivityNet: A Large-Scale Simulated-Event Benchmark for Untrimmed Action Understanding

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing long-form action understanding methods, which are hindered by the scarcity of large-scale, natively captured event-stream datasets with dense temporal annotations and often rely on short, trimmed video clips. To bridge this gap, we introduce Event ActivityNet—the first large-scale simulated event benchmark for untrimmed long videos—comprising 3,263 videos, 200 action classes, and 106.94 hours of content, accompanied by 5/9-bin event voxels, temporal annotations, and timestamped captions. We propose a progressive nested training strategy and an adaptive event frame-formation method that preserves original frame-rate metadata to enable precise temporal mapping, facilitating event-language alignment and multimodal localization. Experiments demonstrate that our approach achieves a Top-1 action recognition accuracy of 66.42% and a mean average precision of 29.0 for online temporal localization, with pretraining on simulated events significantly outperforming training from scratch or within a single domain, thereby effectively alleviating the data scarcity bottleneck in event-based vision.
📝 Abstract
Long-horizon event-based action understanding remains underexplored because existing datasets largely comprise short, trimmed clips, while collecting native event streams with dense temporal annotations is costly. We introduce Event ActivityNet, a large-scale simulated-event benchmark derived from human-annotated, untrimmed ActivityNet videos. It comprises 3,263 videos, 200 action classes, and 106.94 hours, with matched 5-bin and 9-bin event-voxel representations, temporal action annotations, and timestamped captions. The benchmark supports annotated-segment action recognition, auxiliary event-language alignment, and causal online temporal action localization. We generate event voxels directly from non-interpolated source videos in decoded frame order, retain per-video rational nominal or average frame-rate metadata for approximate time mapping, and use action-center reconstruction LPIPS as a soft diagnostic of retained reconstructable content. We establish baselines for adaptive event framing, prompt-caption alignment, and event-only, RGB-only, and RGB-event localization. Under a progressive nested-scale training protocol, recognition Top-1 accuracy increases from 52.25 to 66.42, while online temporal localization average mAP improves from 21.7 to 29.0. Moreover, staged Event ActivityNet pretraining followed by native-event fine-tuning consistently outperforms target-only and joint-from-scratch training across multiple supervision budgets. Event ActivityNet provides a scalable benchmark for long-horizon event modeling, although native-camera evaluation remains essential for deployment-oriented conclusions.
Problem

Research questions and friction points this paper is trying to address.

event-based action understanding
long-horizon
untrimmed videos
temporal action localization
simulated-event benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

event-based vision
untrimmed action understanding
event voxel
temporal action localization
simulated-event benchmark
🔎 Similar Papers
No similar papers found.