🤖 AI Summary
Existing temporal action segmentation methods face deployment challenges due to architectural complexity, imprecise boundary localization, and weak intra-segment consistency. This work proposes a lightweight dual-loss training framework that enhances performance without modifying the backbone architecture—requiring only an additional output channel and two auxiliary loss terms. The first is a single-channel boundary regression loss designed to improve temporal boundary accuracy, and the second is a segment-level regularization based on the cumulative distribution function (CDF) to strengthen intra-segment consistency. The approach is architecture-agnostic and seamlessly integrates with mainstream models such as MS-TCN, C2F-TCN, and FACT. Evaluated on three benchmark datasets, it consistently achieves significant gains in F1 and Edit scores while maintaining stable frame-wise accuracy, demonstrating the efficacy of a minimalist yet well-designed loss formulation.
📝 Abstract
Recent progress in Temporal Action Segmentation (TAS) has increasingly relied on complex architectures, which can hinder practical deployment. We present a lightweight dual-loss training framework that improves fine-grained segmentation quality with only one additional output channel and two auxiliary loss terms, requiring minimal architectural modification. Our approach combines a boundary-regression loss that promotes accurate temporal localization via a single-channel boundary prediction and a CDF-based segment-level regularization loss that encourages coherent within-segment structure by matching cumulative distributions over predicted and ground-truth segments. The framework is architecture-agnostic and can be integrated into existing TAS models (e.g., MS-TCN, C2F-TCN, FACT) as a training-time loss function. Across three benchmark datasets, the proposed method improves segment-level consistency and boundary quality, yielding higher F1 and Edit scores across three different models. Frame-wise accuracy remains largely unchanged, highlighting that precise segmentation can be achieved through simple loss design rather than heavier architectures or inference-time refinements.