Score
Design, build, and evaluate models that detect, classify, and temporally localize human actions and facial/body expressions in visual temporal data (video), producing per-frame, segment-level, or video-level labels and predictions; may also include models that predict future actions or integrate visual and language inputs to reason about actions.
This work addresses the limitation of single-modality approaches—using either video or motion data alone—in comprehensively capturing human behavioral semantics and fine-grained actions. To this end, we propose ViMoNet, a novel multimodal joint-training framework that, for the first time, simultaneously models high-fidelity 3D motion sequences and general-purpose video spatiotemporal features, while leveraging large language models to achieve cross-modal semantic alignment. Complementing this, we introduce VIMOS, a new multimodal dataset featuring dual-track annotations: motion–text and video–text pairs. Extensive experiments demonstrate that ViMoNet significantly outperforms state-of-the-art methods on behavior captioning, action understanding, and semantic reasoning tasks. Furthermore, we establish ViMoNet-Bench—a dedicated benchmark for fine-grained behavior understanding—which validates ViMoNet’s strong generalization capability and robustness across diverse scenarios.
This survey systematically categorizes video action understanding into three temporal regimes: full-action recognition, partial-observation prediction, and unobserved forecasting—first unifying their distinct modeling essences. Addressing limitations in long-horizon causal reasoning and cross-modal synergy, it comprehensively reviews deep temporal modeling, contrastive learning, multimodal alignment, generative video modeling, and zero-shot transfer, integrating Transformer, CNN-LSTM, and diffusion-based paradigms. Synthesizing over 100 seminal works, it identifies critical performance bottlenecks and evaluation biases, constructing a structured knowledge graph. The core contribution is the proposal of the “dynamic reasoning” paradigm—a conceptual and methodological shift that advances action understanding from static classification toward systematic, causal, anticipatory, and generalizable modeling.
Current Video-LLMs exhibit fundamental limitations in modeling abstract temporal concepts—such as long-range dependencies, causal relationships, and event evolution—due to the absence of explicit temporal annotations in video datasets, domain-specific biases, and a temporal misalignment between visual encoders and LLMs. Method: We systematically diagnose performance gaps in cross-segment event association and causal reasoning; propose a novel spatio-temporal semantic joint modeling paradigm; and construct a multi-source video data framework with explicit temporal annotations. We further conduct temporal attribution analysis, multimodal fusion modeling, and bias assessment to validate our approach. Contribution/Results: Our method significantly enhances temporal awareness in Video-LLMs, demonstrating measurable improvements in temporal reasoning tasks. The framework is fully reproducible and empirically verifiable, offering a principled technical pathway toward next-generation Video-LLMs with robust, interpretable temporal understanding capabilities.
This work addresses the challenge of building video-language understanding systems with human-like perceptual capabilities, enabling synergistic modeling of linguistic and dynamic visual temporal sequences. We systematically survey model architectures, training paradigms, and data construction methodologies in this domain, and introduce— for the first time—a unified, cross-perspective taxonomy that exposes core challenges including multimodal temporal alignment and dataset bias. Leveraging Transformer-based fusion, contrastive/generative pretraining, synthetic data augmentation, and benchmarks such as How2QA and Ego4D, we conduct a comprehensive, reproducible horizontal evaluation of state-of-the-art models under a standardized assessment protocol. Our key contributions are: (1) the first structured analytical framework for joint video-language modeling; (2) clear identification of critical research directions; and (3) a practical, deployable technology roadmap for embodied intelligence.
Existing video understanding benchmarks overestimate multimodal foundation models’ (MFMs) temporal reasoning capabilities, as their questions can often be answered using single frames, sparse frame sampling, or out-of-order frames—failing to assess continuous dynamic modeling. Method: We introduce TOMATO, the first rigorous video temporal reasoning benchmark, comprising 1,417 human-annotated videos with high temporal dependency across six dynamic understanding tasks. We propose three novel evaluation principles—multi-frame gain, frame-order sensitivity, and frame-information disparity—to design structured temporal question-answering tasks and establish cross-task human baselines. Contribution/Results: Experiments reveal that current MFMs severely lack frame-order awareness and dynamic integration capability: the best-performing model lags human accuracy by 57.3%. TOMATO establishes a diagnostic, quantitative paradigm for evaluating temporal reasoning in video understanding models.
研究通过对比人类和视觉-语言模型在分析以人为中心的视频时的表现,探讨了这些模型独立完成任务的能力及人机协作流程的有效性。
This work addresses the challenge of fine-grained, structured annotation of human body language—including pose and emotion—in video. We propose an end-to-end pipeline leveraging dual vision-language models (VLMs): Qwen2.5-VL-7B and Llama-4-Scout-17B. Our method integrates visual tokenization, multimodal Transformer attention, and instruction tuning to achieve frame-level person detection (pixel-accurate bounding boxes), prompt-conditioned emotion recognition, and cross-frame ID-consistent modeling, augmented by a schema-driven output validation module ensuring structural compliance. Methodologically, we are the first to systematically disentangle critical boundaries—syntactic validity versus semantic correctness, structural validation versus geometric precision, and local frame-level ID assignment versus cross-frame tracking—explicitly guided by VLM architectural properties. Experiments demonstrate reproducibility, interface robustness, and evaluation reliability, establishing a novel paradigm for controllable VLM deployment in embodied perception tasks.
Action recognition models often exhibit background bias—over-relying on contextual cues at the expense of motion semantics—thereby compromising generalization and robustness. This work presents the first systematic evaluation of background bias across three model families: standard classification models, contrastive vision-language pre-trained models (e.g., CLIP), and video large language models (VLLMs). To mitigate this bias, we propose a dual-path disentanglement framework: (1) input purification via human instance segmentation to explicitly remove background interference, and (2) prompt tuning that synergistically combines handcrafted priors with automated search to steer attention toward human pose and motion dynamics. We introduce a quantitative bias evaluation framework to rigorously assess mitigation efficacy. Experiments show a 3.78% reduction in background bias for classification models and a 9.85% improvement in human-centric focus for VLLMs on action discrimination tasks, yielding substantial gains in cross-background robustness.
This work addresses key challenges in long-form video understanding with multimodal large language models—namely sparse evidence, cross-modal misalignment, and computational constraints—by proposing a human-centric unified framework that integrates “watching, memory, and reasoning.” The framework systematically models structured relationships among perceptual representations, memory states, and reasoning trajectories. It combines fine-grained audiovisual perception, hybrid offline and streaming memory mechanisms, and joint text-video reasoning, enabling end-to-end training and efficient processing of long videos. To advance the field, the authors introduce a comprehensive evaluation benchmark and dataset spanning five video domains, clearly delineating core challenges and charting future research directions, thereby establishing both theoretical and practical foundations for multimodal large language models in video understanding.
This work addresses the limitation of existing short-term human pose forecasting methods, which predominantly rely on geometric motion cues while neglecting the influence of emotion on movement dynamics. The authors propose a lightweight autoregressive world model that integrates emotion embeddings—extracted from facial expressions—with pose keypoints through a learnable normalized gating mechanism, enabling 15-step roll-out predictions. Built upon a two-layer LSTM architecture, the model is evaluated on two emotion-annotated pose video datasets. Experimental results demonstrate that the normalized gating mechanism substantially improves prediction accuracy for emotion-driven motion sequences. Furthermore, counterfactual perturbation analyses reveal that predicted trajectories are sensitive to emotional inputs, confirming that emotion embeddings serve as effective conditioning signals rather than redundant features.