🤖 AI Summary
This study addresses the absence of explicit class-temporal correspondence supervision in multi-instance partial label learning for video classification by proposing the PIVOT-MIPL framework. This method jointly models label disambiguation and evidence allocation, achieving non-uniform temporal quality learning through occupancy-regularized spherical matching and bi-marginal KL projection. Furthermore, it demonstrates that the objective function can be decomposed into class-marginal and temporal supervision components, enhancing optimization efficiency via candidate-constrained inference and momentum-based belief refinement. Experimental results indicate that the proposed framework significantly outperforms existing methods across four feature representations on benchmarks such as Breakfast, effectively balancing predictive performance with computational efficiency.
📝 Abstract
Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples label disambiguation with temporal evidence allocation through a joint class--time assignment. Occupancy-regularized spherical matching associates contextualized video features while learning nonuniform temporal mass and discouraging excessive concentration. During training, candidate-restricted inference recomputes the assignment within the candidate label set. A dual-marginal KL projection then constructs a structured teacher that incorporates momentum-refined class beliefs while preserving the proposal's temporal occupancy. A single plan-level KL objective aligns the full-space predictor with this teacher. Our analysis characterizes when candidate re-solving differs from masking and shows that, under the stated construction, the joint objective decomposes into class-marginal and class-conditional temporal supervision. We construct VCMIPL benchmarks from Breakfast, DoTA, and FineAction using model-generated candidate labels and evaluate the method across four feature representations. Extensive experimental results demonstrate that PIVOTMIPL outperforms existing MIPL algorithms in both effectiveness and efficiency.