HSEmotion Team at ABAW-10 Competition: Facial Expression Recognition, Valence-Arousal Estimation, Action Unit Detection and Fine-Grained Violence Classification

📅 2026-03-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses four challenging tasks in real-world scenarios: facial expression recognition, valence-arousal estimation, action unit detection, and fine-grained violent behavior classification. The authors propose an efficient two-stage prediction framework that leverages EfficientNet-based pretrained models to extract facial embeddings, followed by confidence-thresholded frame-level predictions using multilayer perceptrons. Temporal consistency is enhanced through a sliding-window smoothing strategy. For violent behavior detection, the study systematically evaluates various pretrained architectures and video-level embedding aggregation methods. The proposed approach significantly outperforms existing baselines across all four tasks in the ABAW-10 challenge, achieving substantial gains in robustness for affective and behavioral understanding under complex conditions while maintaining high inference efficiency.

Technology Category

Computer Vision: Video Understanding & Activity AnalysisHumans and AI: Human-Aware Planning and Behavior PredictionCognitive Modeling & Cognitive Systems: Affective Computing

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataResponsible Web: Machine-in-the-loop, human agency and autonomySystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
This article presents our results for the 10th Affective Behavior Analysis in-the-Wild (ABAW) competition. For frame-wise facial emotion understanding tasks (frame-wise facial expression recognition, valence-arousal estimation, action unit detection), we propose a fast approach based on facial embedding extraction with pre-trained EfficientNet-based emotion recognition models. If the latter model's confidence exceeds a threshold, its prediction is used. Otherwise, we feed embeddings into a simple multi-layered perceptron trained on the AffWild2 dataset. Estimated class-level scores are smoothed in a sliding window of fixed size to mitigate noise in frame-wise predictions. For the fine-grained violence detection task, we examine several pre-trained architectures for frame embeddings and their aggregation for video classification. Experimental results on four tasks from the ABAW challenge demonstrate that our approach significantly improves validation metrics over existing baselines.
Problem

Research questions and friction points this paper is trying to address.

Facial Expression Recognition
Valence-Arousal Estimation
Action Unit Detection
Fine-Grained Violence Classification
Affective Behavior Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

facial embedding
EfficientNet
two-stage prediction
sliding window smoothing
video-level aggregation
A
Andrey V. Savchenko
Sber AI Lab; HSE University, Moscow, Russia
K
Kseniia Tsypliakova
HSE University, Moscow, Russia