Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges

📅 2026-09-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of interpreting language model gains in video action prediction, where confounding variables obscure causal attributions. To this end, we propose a systematic evaluation framework. Methodologically, we construct an evidence-aware design graph to delineate task paradigms and intervention points, and introduce the Backbone-aware Comparison and Ablation Protocol (BCAP), which reframes large language model advantages as testable hypotheses rather than causal conclusions. The framework further integrates vision-language models, multi-axis taxonomies, and counterfactual diagnostic techniques. As key contributions, we conduct comprehensive audits on datasets including Ego4D, release versioned catalogs, and clarify open challenges within the field.
📝 Abstract
Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce large language models (LLMs), vision-language models (VLMs), or language-derived semantics at different stages, but reported gains are difficult to interpret when task formulation, visual pretraining, supervision, decoder design, and evaluation code change simultaneously. The central contribution of this review is an evidence-aware design map that crosses task regime with the point at which language-derived information intervenes. We characterise task regimes along six axes. These axes organise the literature into five broad task families: single-action, sequence, object-interaction, cross-view, and planning-oriented settings. C1-C3 locate interventions in context construction, goal/intention modelling, and future decoding, while C4 is treated as an adjacent, emerging grounding/executability extension. Unlike a generic processing pipeline, the map links each intervention to an appropriate counterfactual, failure diagnosis, and permissible evidence claim. Supporting contributions include a protocol-level audit of Ego4D-LTA and EPIC-KITCHENS-100, a multidimensional evidence profile, and the Backbone-Aware Comparison and Ablation Protocol (BCAP). The unresolved EK-100 record is treated as a reporting-comparability case study and is not used as a leaderboard. Evidence for LLM benefits, goal ambiguity, and horizon effects is therefore formulated as testable hypotheses requiring matched validation, not as causal conclusions. The accompanying package contains the coded evidence, source locators, protocol metadata, and versioned catalogue used in the review.
Problem

Research questions and friction points this paper is trying to address.

Action Anticipation
Language Augmentation
Benchmark Audit
Evidence Comparability
Task Formulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action Anticipation
Evidence-Aware Design Map
Vision-Language Models
BCAP
Intervention Taxonomy
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13
M
Mahsa Mohammadi
Department of Computer Science, University of Exeter, Harrison Building, Streatham Campus, North Park Road, Exeter, EX4 4QF, Devon, United Kingdom
Zeyu Fu
Zeyu Fu
Lecturer, Department of Computer Science, University of Exeter
Multimedia ComputingMedical Image AnalysisAI4Science
S
Sareh Rowlands
Department of Computer Science, University of Exeter, Harrison Building, Streatham Campus, North Park Road, Exeter, EX4 4QF, Devon, United Kingdom