semantic-aware activity recognition

Design and build models that recognize activities or events from temporal sensory data by integrating semantic information—aligning global activity predictions with local semantic cues and using multi-loss semantic supervision. Implement semantic-guided temporal reasoning (e.g., transformer-based modules) to improve label alignment and robustness to low-visibility or ambiguous inputs.

semantic-awareactivityrecognition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of semantic-temporal misalignment and insufficient multi-scale temporal modeling in few-shot action recognition by proposing the STAR framework, which enforces fine-grained semantic-temporal consistency through two core modules: semantic alignment and temporal awareness. The framework innovatively integrates Temporal Semantic Attention (TSA) and a semantics-guided Mamba block, jointly optimized with temporal dependency descriptors generated by a large language model to align semantic and dynamic temporal representations. Additionally, a Semantic-Temporal Prototype Refiner (STPR) and a multi-frequency temporal sampling strategy are introduced to enhance cross-video temporal modeling. Extensive experiments demonstrate that STAR significantly outperforms existing methods across five benchmarks, achieving absolute gains of 8.1%, 6.7%, and 7.3% under the 1-shot setting on SSv2-Full, SSv2-Small, and HMDB51, respectively.

few-shot action recognitionmulti-scale temporal dynamicssemantic-temporal misalignment

Do Language Models Understand Time?

Dec 18, 2024
XD
Xi Ding
🏛️ Australian National University

Current Video-LLMs exhibit fundamental limitations in modeling abstract temporal concepts—such as long-range dependencies, causal relationships, and event evolution—due to the absence of explicit temporal annotations in video datasets, domain-specific biases, and a temporal misalignment between visual encoders and LLMs. Method: We systematically diagnose performance gaps in cross-segment event association and causal reasoning; propose a novel spatio-temporal semantic joint modeling paradigm; and construct a multi-source video data framework with explicit temporal annotations. We further conduct temporal attribution analysis, multimodal fusion modeling, and bias assessment to validate our approach. Contribution/Results: Our method significantly enhances temporal awareness in Video-LLMs, demonstrating measurable improvements in temporal reasoning tasks. The framework is fully reproducible and empirically verifiable, offering a principled technical pathway toward next-generation Video-LLMs with robust, interpretable temporal understanding capabilities.

Assess LLMs' understanding of temporal relationships in videosIdentify limitations in LLMs' modeling of long-term dependenciesPropose solutions for enhancing temporal comprehension in LLMs

SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition

Oct 14, 2024
ZL
Zechen Li
🏛️ University of New South Wales | University of Tokyo

Large language models (LLMs) struggle to process raw motion sensor time-series data due to semantic sparsity, numerical input incompatibility, and computational constraints. To address this, we propose SensorLLM—a two-stage sensor-to-language alignment framework. Its core contributions are: (1) channel-specific special tokens coupled with auto-generated trend-oriented textual descriptions, enabling semantic encoding of multichannel, variable-length numeric sequences; and (2) an integrated pipeline combining textualized sequence representation, special token embedding, instruction tuning, and task-aware LoRA adaptation—enabling zero-shot human activity recognition (HAR). Evaluated across multiple benchmarks, SensorLLM achieves or surpasses state-of-the-art performance, demonstrating high accuracy, cross-device transferability, and strong generalization capability.

Achieves state-of-the-art performance in HAR classification.Addresses challenges in processing numerical sensor inputs.Enables LLMs to recognize human activities from sensor data.

LASER: A Neuro-Symbolic Framework for Learning Spatial-Temporal Scene Graphs with Weak Supervision

Apr 15, 2023
JH
Jiani Huang
🏛️ University of Pennsylvania | University of Central Florida

This work addresses the high annotation cost of video spatio-temporal scene graphs (STSGs) by proposing a weakly supervised learning framework that relies solely on video-caption pairs. Methodologically, it introduces the first differentiable symbolic reasoning module jointly optimized with contrastive, temporal, and semantic losses to generate logic-guided STSGs; additionally, it leverages large language models (LLMs) to automatically distill spatio-temporal logical rules, forming a neuro-symbolic architecture. Contributions include: (1) the first end-to-end weakly supervised paradigm for STSG generation without manual STSG annotations; (2) an LLM-driven mechanism for automatic spatio-temporal logical rule induction; and (3) state-of-the-art performance on Something-Something V2, MUGEN, and OpenPVSG, demonstrating substantial improvements in fine-grained video semantic representation.

Aligning predicted graphs with logical specifications from captionsLearning spatio-temporal scene graphs without annotated videosUsing video captions as weak supervision for training

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Jun 09, 2024
TN
Thong Nguyen
🏛️ National University of Singapore | Nanyang Technological University

This work addresses the challenge of building video-language understanding systems with human-like perceptual capabilities, enabling synergistic modeling of linguistic and dynamic visual temporal sequences. We systematically survey model architectures, training paradigms, and data construction methodologies in this domain, and introduce— for the first time—a unified, cross-perspective taxonomy that exposes core challenges including multimodal temporal alignment and dataset bias. Leveraging Transformer-based fusion, contrastive/generative pretraining, synthetic data augmentation, and benchmarks such as How2QA and Ego4D, we conduct a comprehensive, reproducible horizontal evaluation of state-of-the-art models under a standardized assessment protocol. Our key contributions are: (1) the first structured analytical framework for joint video-language modeling; (2) clear identification of critical research directions; and (3) a practical, deployable technology roadmap for embodied intelligence.

Analyze challenges in model training for video-language tasksCompare performance and future research directionsSurvey video-language understanding systems' model architectures

Latest Papers

What's happening recently
View more

This work addresses the challenge of achieving zero-shot semantic understanding of high-speed, fine-grained human actions in the absence of labeled data. The authors propose a training-free inference framework that integrates pretrained video-language models with large language models and systematically demonstrate, for the first time, the critical role of temporal resolution in zero-shot action understanding. By analyzing videos captured at multiple frame rates (120/60/30 Hz) and fusing them with pose information derived from human joint tracking, the method substantially enhances the stability and interpretability of semantic representations. In high-velocity action scenarios such as kendo, high frame-rate inputs significantly improve the separability of action semantics, with nearest-class prototype evaluation confirming the approach’s superior performance.

high-speed visionhuman action recognitionsemantic interpretation

This work addresses the challenges of human activity recognition in smart homes, where sparse sensor signals and similar local patterns hinder accurate modeling of semantically complex daily behaviors. To overcome these limitations, the authors propose TRACE, a framework that reframes activity recognition as a context-aware temporal reasoning task rather than isolated local classification. By integrating multi-source sensor evidence with user-specific contextual priors, TRACE enables coherent and robust semantic inference. The approach effectively mitigates prediction fragmentation and significantly improves recognition accuracy for complex activities on both public benchmarks and real-world deployments. Furthermore, it demonstrates consistent robustness under cross-domain scenarios and in the presence of missing modalities.

Contextual ReasoningHuman Activity RecognitionSensor Ambiguity

This work addresses the challenge of learning video representations that are both temporally coherent and semantically meaningful in the absence of labeled data. To this end, the authors propose a momentum-guided semantic prediction framework that performs self-supervised learning by predicting future latent embeddings across randomly sampled temporal intervals, thereby eliminating reliance on pixel-level reconstruction or task-specific alignment. A contrastive regularization term is further introduced to enhance temporal consistency and prevent representation collapse. Experimental results demonstrate that the learned embedding space on UCF101 not only exhibits strong temporal stability but also explicitly captures action semantics and motion patterns, reflecting a superior capacity for structured representation learning.

latent forecastingself-supervised learningsemantic forecasting

This work addresses the lack of interpretability in existing skeleton-based action recognition models, which typically operate as black boxes. The authors propose a concept-driven interpretable framework that reformulates action recognition as first-order logical reasoning grounded in motion primitives. Their approach employs a spatiotemporal skeleton encoder and a concept decoder to learn differentiable spatiotemporal motion concepts, which are instantiated as logical predicates. By integrating a large language model to align atomic action semantics, the method constructs a shared conceptual space bridging perception and reasoning. This is the first effort to incorporate differentiable first-order logic into skeleton-based action recognition, achieving competitive accuracy on the NTU RGB+D 60/120 and NW-UCLA benchmarks while generating human-readable logical rules that enable explicit, interpretable action understanding.

interpretabilitylogical reasoningmotion primitives

ExOAR: Expert-Guided Object and Activity Recognition from Textual Data

Dec 03, 2025
IB
Iris Beerepoot
🏛️ Utrecht University

To address the challenge that unstructured text inadequately supports object-centric process mining, this paper proposes an interactive human-in-the-loop framework. First, a large language model (LLM) generates candidate object types, instances, and activities from textual input. Subsequently, domain-specific contextual information—such as users’ professional backgrounds—is integrated via staged prompting, expert feedback, and iterative refinement to ensure semantic accuracy. This approach overcomes the semantic ambiguity inherent in fully automated extraction, substantially improving both the accuracy and interpretability of structured event logs. Experiments conducted on activity-window data from five real users demonstrate that the resulting logs exhibit well-defined object–activity semantic associations, effectively bridging unstructured text with object-centric process analysis requirements. The framework establishes a novel paradigm for process mining in low-resource settings where high semantic fidelity is critical.

Bridges textual data to structured logs for process miningCombines LLMs with human verification for accuracyExtracts objects and activities from unstructured text

Hot Scholars

JK

Juhi Kulshrestha

Department of Computer Science, Aalto University
Computational Social ScienceSocial computingOnline Social MediaInformation Consumption on the Web
MK

Mohammad Kazemi

Imperial College London
Machine LearningInformation TheorySignal ProcessingWireless Communications
MD

Marc Delcroix

NTT Communication Science Laboratories
Speech processingRobust ASRSpeech enhancementTarget speech extraction
SY

Stephen Yang

Stanford University
Distributed SystemsLow-Latency Systems
SP

Shwetak Patel

University of Washington, Washington Research Foundation Endowed Professor, Computer Science
Ubiquitous ComputingHuman-Computer InteractionSensorsEmbedded Systems