Score
Detecting and classifying human activities or actions from sensor and video streams in real time, including identifying risky behaviors, falls, or gait patterns while accounting for privacy and noisy field conditions.
This paper addresses the long-standing performance stagnation in wearable sensor-based human activity recognition (HAR). Methodologically, it introduces a novel paradigm that integrates foundational model world knowledge into HAR, unifying time-series signal processing, self-supervised pretraining, multimodal fusion, and prompt-based fine-tuning within a single end-to-end framework—designed to support both novice users and domain experts. The contributions are threefold: (1) it identifies and analyzes the root causes of performance saturation on mainstream HAR benchmarks; (2) it empirically validates the proposed paradigm across multiple public datasets, demonstrating significant improvements in classification accuracy and cross-dataset generalization; and (3) it releases comprehensive open-source tutorials and a methodological survey, substantially lowering the barrier to practical HAR deployment.
This paper addresses the challenge of disentangling sensor observations from individual identities in human activity recognition (HAR) within multi-resident environments—a problem exacerbated by identity ambiguity, activity overlap, and insufficient collaborative modeling. We present the first systematic survey dedicated to multi-occupant HAR, proposing a comprehensive evaluation framework tailored to realistic living settings. The framework integrates heterogeneous sensing modalities—including PIR, WiFi CSI, and wearable data—and unifies temporal modeling (LSTM/Transformer), multi-instance learning, and unsupervised identity separation. Through analysis of over 120 studies, we find that state-of-the-art methods achieve only ~78% average accuracy in co-located multi-person scenarios. We identify federated learning and self-supervised representation learning as pivotal avenues for improving robustness and scalability. This work provides both theoretical foundations and practical guidelines for advancing reliable, deployable multi-occupant HAR systems.
This study addresses the challenges of cross-user and cross-scenario generalization and privacy preservation in human activity recognition within home environments. Existing approaches rely heavily on camera-based, pre-defined labels that often misalign with practical sensing capabilities. To overcome this, the work proposes a privacy-preserving activity discovery framework centered on non-visual sensors—such as radar, thermal imaging, and LiDAR—that stably capture natural signal patterns. By adaptively invoking vision-language models only on key frames to interpret scenes, the method autonomously discovers discrete activity categories, thereby shifting away from the conventional camera-centric annotation paradigm and substantially reducing reliance on visual models. Experiments with 12 participants demonstrate that environmental sensors alone achieve 79% accuracy in recognizing 4–5 coarse-grained activities; integrating wearable and depth sensors yields 73% accuracy for 8–9 fine-grained activities (averaging 77%), while reducing visual queries by 90%, thus balancing privacy and generalization effectively.
Existing approaches for implicit class recognition in streaming signals employ multiple heterogeneous-accuracy classifiers under fixed scheduling, yet fail to effectively fuse their outputs or account for temporal dynamics. Method: We propose a real-time state-space filtering model that treats multi-classifier probabilistic outputs as observations and models the true class label—including its temporal evolution—as a latent state, enabling online estimation via Bayesian recursion. The model jointly addresses classifier heterogeneity, temporal dependencies, and strict real-time computational constraints. Results: Evaluated on activity recognition using wearable IMU data, our method achieves significant accuracy improvements over baselines (+3.2%–5.8%), while maintaining inference latency consistently below 20 ms—demonstrating both high precision and strong real-time performance.
This work addresses the lack of a unified, reproducible evaluation platform for human activity recognition (AR) using binary sensor data. Methodologically, we propose the first modular end-to-end AR pipeline, comprising a data-driven three-stage framework: robust data cleaning, sliding-window-based adaptive temporal segmentation, and lightweight personalized classification (using XGBoost or LSTM variants), enabling plug-and-play substitution of methods, datasets, and evaluation protocols. Our key contribution is the first modular AR architecture specifically designed for binary sensing—ensuring full pipeline reproducibility, customization, and rigorous evaluation. Extensive validation across multiple public datasets demonstrates significant improvements in cross-user accuracy and generalization robustness. The platform establishes a standardized experimental baseline for AR research and supports rapid prototyping and deployment.
This study addresses the challenge of jointly optimizing AI-powered video analytics and privacy preservation in community security. Methodologically, it proposes a lightweight, interpretable, privacy-by-design Smart Video Solution (SVS) that replaces raw video transmission with anonymized skeletal pose features extracted via edge-deployed lightweight pose estimation, coupled with statistical time-series anomaly detection. Novel visualization paradigms—including an Occupancy Indicator and bird’s-eye-view heatmaps—are introduced to support multi-stakeholder decision-making (e.g., law enforcement response and urban planning). An edge–cloud collaborative architecture integrates time-series databases and real-time notification services. Deployed across 16 CCTV channels, the system operated stably for 21 hours, achieving 16.5 FPS throughput and an end-to-end latency of 26.76 seconds—demonstrating robustness and practical feasibility. The core contributions are (1) the first privacy-enhanced behavioral understanding framework for community surveillance, and (2) a governance-oriented visual insight translation mechanism.
Existing fall detection methods suffer from poor portability, high privacy risks, and weak generalization—particularly under low-resolution video inputs and complex daily activities. To address these limitations, we propose a privacy-preserving, edge-deployable fall detection framework based on lightweight skeletal sequence extraction. Our approach introduces a structure-aware discriminative feature aggregation mechanism that jointly models joint-level geometric topology and motion dynamics. We further propose a novel Separable-Convolution-Enhanced Graph Convolutional Network (SE-GCN), which achieves superior discriminability while significantly reducing computational overhead. Evaluated on five benchmark datasets, our method achieves state-of-the-art accuracy, with inference speed substantially accelerated and FLOPs reduced by over 60%, enabling real-time execution on resource-constrained edge devices.
To address low accuracy and poor robustness in real-time multimodal anomaly detection caused by audio-video asynchrony and modality mismatch in industrial settings, this paper proposes a unified multimodal fusion framework. It introduces a bidirectional cross-modal attention mechanism for fine-grained alignment between video (YOLOv8/DETR + ByteTrack) and audio (AST/Wav2Vec2/HuBERT) streams, and integrates hybrid object detection with multi-strategy anomaly discrimination to enable synchronous streaming inference. Evaluated on general surveillance and industrial safety benchmarks, the system achieves >25 FPS on standard GPUs, with improvements of +4.2% mAP, +7.8% anomaly detection rate, and −12.3% false positive rate. Key contributions include: (i) the first end-to-end audio-visual synchronized anomaly detection system designed specifically for industrial deployment; and (ii) an extensible cross-modal attention architecture coupled with a lightweight audio integration scheme.
This work addresses the challenges of high latency, privacy risks, and computational bottlenecks that hinder efficient violence detection in existing edge-based video analytics systems for public safety. To overcome these limitations, the authors propose a hybrid edge action detection architecture that innovatively integrates skeleton-based pose analysis with the semantic reasoning capabilities of vision-language foundation models. Deployed on GPU-enabled edge devices, the system enables low-latency, low-overhead real-time inference while supporting context-aware and zero-shot detection. It dynamically orchestrates motion- and semantics-driven paradigms according to scene requirements. Experimental results demonstrate that the approach achieves a favorable trade-off among latency, resource consumption, and accuracy, highlighting the complementary strengths of the two modalities and offering a practical solution for real-world deployment.
This study addresses the challenges of insufficient activity monitoring and poor treatment adherence among elderly individuals undergoing home-based rehabilitation in resource-limited healthcare settings. To this end, the authors propose a low-cost, privacy-preserving human activity recognition method leveraging wearable inertial sensors—specifically accelerometers and gyroscopes. The key innovation lies in the adoption of a Support Tensor Machine (STM) model, which effectively preserves the spatiotemporal dynamics of human actions through tensor representation, enabling robust classification under low-resource conditions. Experimental results demonstrate that the STM achieves accuracies of 96.67% on the test set and 98.50% in cross-validation, significantly outperforming classical approaches such as logistic regression, random forest, support vector machines, and k-nearest neighbors. These findings underscore the STM’s promising potential for telehealth and in-home health monitoring applications.
This work addresses the deployment challenges of real-time fall detection for older adults—particularly privacy concerns, computational overhead, and bandwidth constraints—and tackles the significant performance degradation of supervised keypoint-based methods under occlusion or low-visibility conditions. To this end, we propose a privacy-preserving fall detection framework leveraging unsupervised keypoints extracted locally, combined with a variational recurrent neural network for motion sequence prediction. Fall events are classified at the sequence level by measuring discrepancies between observed and predicted motion. We present the first systematic comparison of supervised versus unsupervised keypoints in real-world fall detection, demonstrating that the latter exhibits markedly superior robustness under occlusion and subject invisibility. Additionally, we introduce a prediction-driven bandwidth compression mechanism. Experiments on the UR Fall Detection and Human Fall datasets show that our method substantially outperforms supervised approaches in occluded scenarios—where the latter miss nearly 50% of falls—with even greater advantages under bandwidth limitations.
Real-world continuous fall monitoring faces challenges including unknown fall event patterns, computational constraints on wearable devices, and inaccurate performance evaluation under streaming conditions. Method: We propose a lightweight, prior-free real-time fall detection framework that integrates IMU-based streaming data processing with cost-sensitive learning to dynamically optimize decision thresholds—thereby balancing false negatives and false positives—and employs an efficient streaming classifier trained and validated end-to-end on the FARSEEING real-world dataset. Contribution/Results: Our approach achieves perfect recall (1.00), precision of 0.84, and an F1-score of 0.91, with average inference latency under 5 ms per sample. To the best of our knowledge, this is the first work to simultaneously achieve high robustness and ultra-low latency in realistic continuous streaming scenarios, offering a practical, deployable solution for resource-constrained wearable systems.