EAAI 2025
Not all samples are equal: Boosting action segmentation via selective incremental learning
Abstract
Temporal action segmentation (TAS) seeks to perform classification for each frame in a video. Existing methods tend to design diverse network architectures, while overlooking the intrinsic characteristics of training samples. Notably, two key issues arise: (1) Frames around action boundaries are more ambiguous and thus pose greater difficulties for training compared to other frames; and (2) beyond the commonly used categorical labels, the total number of action instances within a video may serve as an additional, potentially vital, supervision cue. To address these issues, this paper introduces a novel method that combines a model-agnostic training strategy with an instance number alignment loss, designed to enhance the performance of existing models. Specifically, a selective incremental learning (SIL) strategy is proposed to alleviate the impact of noisy samples by progressively training the model in an easy-to-difficult manner through a dynamic sample selection mechanism. Furthermore, an instance number alignment loss (INAL) is developed to capture both global and local features simultaneously by incorporating a multi-task learning module. Extensive evaluations are conducted on three benchmark datasets, namely 50Salads, Georgia Tech egocentric activities (GTEA), and Breakfast. The experimental results demonstrate that the proposed method achieves substantial performance improvements over state-of-the-art approaches.
Authors
Keywords
Context
- Venue
- Engineering Applications of Artificial Intelligence
- Archive span
- 1988-2026
- Indexed papers
- 13269
- Paper id
- 940360180107957162