activity recognition modeling

Design, build, and evaluate models that detect, classify, and temporally localize human actions and facial/body expressions in visual temporal data (video), producing per-frame, segment-level, or video-level labels and predictions; may also include models that predict future actions or integrate visual and language inputs to reason about actions.

activityrecognitionmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video

Aug 13, 2025
RD
Rajan Das Gupta
🏛️ AIUB | Multimedia University | Washington University of Science & Technology

This work addresses the limitation of single-modality approaches—using either video or motion data alone—in comprehensively capturing human behavioral semantics and fine-grained actions. To this end, we propose ViMoNet, a novel multimodal joint-training framework that, for the first time, simultaneously models high-fidelity 3D motion sequences and general-purpose video spatiotemporal features, while leveraging large language models to achieve cross-modal semantic alignment. Complementing this, we introduce VIMOS, a new multimodal dataset featuring dual-track annotations: motion–text and video–text pairs. Extensive experiments demonstrate that ViMoNet significantly outperforms state-of-the-art methods on behavior captioning, action understanding, and semantic reasoning tasks. Furthermore, we establish ViMoNet-Bench—a dedicated benchmark for fine-grained behavior understanding—which validates ViMoNet’s strong generalization capability and robustness across diverse scenarios.

Combining motion and video data for human behavior understandingCreating a new dataset and benchmark for behavior analysisDeveloping a multimodal framework for action comprehension and inference

About Time: Advances, Challenges, and Outlooks of Action Understanding

Nov 22, 2024
AS
Alexandros Stergiou
🏛️ University of Twente | Utrecht University

This survey systematically categorizes video action understanding into three temporal regimes: full-action recognition, partial-observation prediction, and unobserved forecasting—first unifying their distinct modeling essences. Addressing limitations in long-horizon causal reasoning and cross-modal synergy, it comprehensively reviews deep temporal modeling, contrastive learning, multimodal alignment, generative video modeling, and zero-shot transfer, integrating Transformer, CNN-LSTM, and diffusion-based paradigms. Synthesizing over 100 seminal works, it identifies critical performance bottlenecks and evaluation biases, constructing a structured knowledge graph. The core contribution is the proposal of the “dynamic reasoning” paradigm—a conceptual and methodological shift that advances action understanding from static classification toward systematic, causal, anticipatory, and generalizable modeling.

Address challenges in action modeling and video representationOutline future directions for video action understandingReview advances in uni- and multi-modal action understanding

Do Language Models Understand Time?

Dec 18, 2024
XD
Xi Ding
🏛️ Australian National University

Current Video-LLMs exhibit fundamental limitations in modeling abstract temporal concepts—such as long-range dependencies, causal relationships, and event evolution—due to the absence of explicit temporal annotations in video datasets, domain-specific biases, and a temporal misalignment between visual encoders and LLMs. Method: We systematically diagnose performance gaps in cross-segment event association and causal reasoning; propose a novel spatio-temporal semantic joint modeling paradigm; and construct a multi-source video data framework with explicit temporal annotations. We further conduct temporal attribution analysis, multimodal fusion modeling, and bias assessment to validate our approach. Contribution/Results: Our method significantly enhances temporal awareness in Video-LLMs, demonstrating measurable improvements in temporal reasoning tasks. The framework is fully reproducible and empirically verifiable, offering a principled technical pathway toward next-generation Video-LLMs with robust, interpretable temporal understanding capabilities.

Assess LLMs' understanding of temporal relationships in videosIdentify limitations in LLMs' modeling of long-term dependenciesPropose solutions for enhancing temporal comprehension in LLMs

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Jun 09, 2024
TN
Thong Nguyen
🏛️ National University of Singapore | Nanyang Technological University

This work addresses the challenge of building video-language understanding systems with human-like perceptual capabilities, enabling synergistic modeling of linguistic and dynamic visual temporal sequences. We systematically survey model architectures, training paradigms, and data construction methodologies in this domain, and introduce— for the first time—a unified, cross-perspective taxonomy that exposes core challenges including multimodal temporal alignment and dataset bias. Leveraging Transformer-based fusion, contrastive/generative pretraining, synthetic data augmentation, and benchmarks such as How2QA and Ego4D, we conduct a comprehensive, reproducible horizontal evaluation of state-of-the-art models under a standardized assessment protocol. Our key contributions are: (1) the first structured analytical framework for joint video-language modeling; (2) clear identification of critical research directions; and (3) a practical, deployable technology roadmap for embodied intelligence.

Analyze challenges in model training for video-language tasksCompare performance and future research directionsSurvey video-language understanding systems' model architectures

TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Oct 30, 2024
ZS
Ziyao Shangguan
🏛️ Yale University | Allen Institute for AI

Existing video understanding benchmarks overestimate multimodal foundation models’ (MFMs) temporal reasoning capabilities, as their questions can often be answered using single frames, sparse frame sampling, or out-of-order frames—failing to assess continuous dynamic modeling. Method: We introduce TOMATO, the first rigorous video temporal reasoning benchmark, comprising 1,417 human-annotated videos with high temporal dependency across six dynamic understanding tasks. We propose three novel evaluation principles—multi-frame gain, frame-order sensitivity, and frame-information disparity—to design structured temporal question-answering tasks and establish cross-task human baselines. Contribution/Results: Experiments reveal that current MFMs severely lack frame-order awareness and dynamic integration capability: the best-performing model lags human accuracy by 57.3%. TOMATO establishes a diagnostic, quantitative paradigm for evaluating temporal reasoning in video understanding models.

Assessing visual temporal reasoning capabilities in multimodal foundation modelsEvaluating if models truly understand temporal context in video sequencesMeasuring models' ability to interpret frames as continuous sequences

Latest Papers

What's happening recently
View more

This work addresses the challenge of fine-grained, structured annotation of human body language—including pose and emotion—in video. We propose an end-to-end pipeline leveraging dual vision-language models (VLMs): Qwen2.5-VL-7B and Llama-4-Scout-17B. Our method integrates visual tokenization, multimodal Transformer attention, and instruction tuning to achieve frame-level person detection (pixel-accurate bounding boxes), prompt-conditioned emotion recognition, and cross-frame ID-consistent modeling, augmented by a schema-driven output validation module ensuring structural compliance. Methodologically, we are the first to systematically disentangle critical boundaries—syntactic validity versus semantic correctness, structural validation versus geometric precision, and local frame-level ID assignment versus cross-frame tracking—explicitly guided by VLM architectural properties. Experiments demonstrate reproducibility, interface robustness, and evaluation reliability, establishing a novel paradigm for controllable VLM deployment in embodied perception tasks.

Address semantic correctness and system constraints in video-to-artifact pipelinesDetect visible people and emotions from video frames using vision-language modelsGenerate structured bounding box outputs with prompt-conditioned attributes

Seeing Beyond the Scene: Analyzing and Mitigating Background Bias in Action Recognition

Dec 17, 2025
EZ
Ellie Zhou
🏛️ Westmont High School | Princeton University

Action recognition models often exhibit background bias—over-relying on contextual cues at the expense of motion semantics—thereby compromising generalization and robustness. This work presents the first systematic evaluation of background bias across three model families: standard classification models, contrastive vision-language pre-trained models (e.g., CLIP), and video large language models (VLLMs). To mitigate this bias, we propose a dual-path disentanglement framework: (1) input purification via human instance segmentation to explicitly remove background interference, and (2) prompt tuning that synergistically combines handcrafted priors with automated search to steer attention toward human pose and motion dynamics. We introduce a quantitative bias evaluation framework to rigorously assess mitigation efficacy. Experiments show a 3.78% reduction in background bias for classification models and a 9.85% improvement in human-centric focus for VLLMs on action discrimination tasks, yielding substantial gains in cross-background robustness.

Analyzes background bias in action recognition modelsExplores prompt tuning for human-focused reasoningProposes mitigation strategies to reduce bias

This work addresses key challenges in long-form video understanding with multimodal large language models—namely sparse evidence, cross-modal misalignment, and computational constraints—by proposing a human-centric unified framework that integrates “watching, memory, and reasoning.” The framework systematically models structured relationships among perceptual representations, memory states, and reasoning trajectories. It combines fine-grained audiovisual perception, hybrid offline and streaming memory mechanisms, and joint text-video reasoning, enabling end-to-end training and efficient processing of long videos. To advance the field, the authors introduce a comprehensive evaluation benchmark and dataset spanning five video domains, clearly delineating core challenges and charting future research directions, thereby establishing both theoretical and practical foundations for multimodal large language models in video understanding.

long-range dependenciesmultimodal alignmentmultimodal large language models

This work addresses the limitation of existing short-term human pose forecasting methods, which predominantly rely on geometric motion cues while neglecting the influence of emotion on movement dynamics. The authors propose a lightweight autoregressive world model that integrates emotion embeddings—extracted from facial expressions—with pose keypoints through a learnable normalized gating mechanism, enabling 15-step roll-out predictions. Built upon a two-layer LSTM architecture, the model is evaluated on two emotion-annotated pose video datasets. Experimental results demonstrate that the normalized gating mechanism substantially improves prediction accuracy for emotion-driven motion sequences. Furthermore, counterfactual perturbation analyses reveal that predicted trajectories are sensitive to emotional inputs, confirming that emotion embeddings serve as effective conditioning signals rather than redundant features.

emotion conditioningfacial expressionhuman pose forecasting

Hot Scholars

AF

Antonino Furnari

Assistant Professor at the University of Catania
Computer Vision
KP

Kunyu Peng

Karlsruhe Institute of Technology
video understandingopen set recognitiongeneralizable deep learning
HG

Harsh Goel

University of Texas at Austin
Reinforcement LearningRoboticsGenerative AINeurosymbolic AI
AY

Angela Yao

National University of Singapore
computer visiondeep learningmachine learning