Hierarchical Action Learning for Weakly-Supervised Action Segmentation

📅 2026-02-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the over-segmentation problem in weakly supervised action segmentation, which arises from coarse-grained annotations, by proposing a hierarchical causal generative model that explicitly captures the multi-scale temporal structure of actions: high-level semantic action variables evolve slowly, while low-level visual features change rapidly. The approach integrates a hierarchical pyramid Transformer architecture, a deterministic temporal alignment mechanism, and sparse transition constraints, grounded in rigorous identifiability theory, to enable stable learning and precise localization of high-level action semantics. Extensive experiments demonstrate that the method significantly outperforms existing approaches across multiple weakly supervised action segmentation benchmarks, confirming its effectiveness and generalization capability in real-world scenarios.

Technology Category

Knowledge Representation and Reasoning: Action, Change, and CausalityComputer Vision: SegmentationMachine Learning: Causal Learning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Humans perceive actions through key transitions that structure actions across multiple abstraction levels, whereas machines, relying on visual features, tend to over-segment. This highlights the difficulty of enabling hierarchical reasoning in video understanding. Interestingly, we observe that lower-level visual and high-level action latent variables evolve at different rates, with low-level visual variables changing rapidly, while high-level action variables evolve more slowly, making them easier to identify. Building on this insight, we propose the Hierarchical Action Learning (\textbf{HAL}) model for weakly-supervised action segmentation. Our approach introduces a hierarchical causal data generation process, where high-level latent action governs the dynamics of low-level visual features. To model these varying timescales effectively, we introduce deterministic processes to align these latent variables over time. The \textbf{HAL} model employs a hierarchical pyramid transformer to capture both visual features and latent variables, and a sparse transition constraint is applied to enforce the slower dynamics of high-level action variables. This mechanism enhances the identification of these latent variables over time. Under mild assumptions, we prove that these latent action variables are strictly identifiable. Experimental results on several benchmarks show that the \textbf{HAL} model significantly outperforms existing methods for weakly-supervised action segmentation, confirming its practical effectiveness in real-world applications.
Problem

Research questions and friction points this paper is trying to address.

weakly-supervised action segmentation
hierarchical reasoning
action segmentation
latent variables
video understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

hierarchical action learning
weakly-supervised action segmentation
latent variable identifiability
multi-timescale modeling
pyramid transformer
🔎 Similar Papers
No similar papers found.