🤖 AI Summary
This study addresses the lack of counterfactual video generation evaluation capabilities in existing roadside datasets, where modifying specific traffic participants often compromises road topology and irrelevant traffic consistency. We construct the first roadside counterfactual video generation benchmark and propose an actor-level procedural counterfactual representation method alongside a four-dimensional validity verification criterion, supporting behavior reasoning, intervention editing, and conditional generation tasks. Experimental results demonstrate that the optimal reasoner achieves a macro F1-score of 80.4%. Furthermore, a comprehensive conditional interface improves the end-to-end success rate of the best generator from 23.3% to 55.0%, revealing that conditional video execution constitutes the primary bottleneck constraining generation performance.
📝 Abstract
Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topology and unrelated traffic. We introduce TrafficImag, the first benchmark for counterfactual roadside traffic video generation. TrafficImag combines a large-scale roadside dataset (9,022 annotated images, 7,043 deduplicated video clips, and 31,145 actor-centered history-future samples) with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is represented as an actor-level program describing the target actor, intended behavior, legal route, interaction order, and temporal constraints, enabling a unified evaluation interface across heterogeneous foundation models. TrafficImag evaluates four complementary validity dimensions: initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation, and considers an end-to-end counterfactual successful only when all four are satisfied. Across state-of-the-art foundation models, the strongest reasoner reaches 80.4% macro F1, the complete condition interface raises end-to-end success from 23.3% to 55.0% for the best generator. Oracle studies further show that conditional video execution is the primary remaining bottleneck. TrafficImag provides a reproducible benchmark for evaluating and diagnosing counterfactual traffic video generation beyond perceptual video quality.