Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of explicit class-temporal correspondence supervision in multi-instance partial label learning for video classification by proposing the PIVOT-MIPL framework. This method jointly models label disambiguation and evidence allocation, achieving non-uniform temporal quality learning through occupancy-regularized spherical matching and bi-marginal KL projection. Furthermore, it demonstrates that the objective function can be decomposed into class-marginal and temporal supervision components, enhancing optimization efficiency via candidate-constrained inference and momentum-based belief refinement. Experimental results indicate that the proposed framework significantly outperforms existing methods across four feature representations on benchmarks such as Breakfast, effectively balancing predictive performance with computational efficiency.
📝 Abstract
Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples label disambiguation with temporal evidence allocation through a joint class--time assignment. Occupancy-regularized spherical matching associates contextualized video features while learning nonuniform temporal mass and discouraging excessive concentration. During training, candidate-restricted inference recomputes the assignment within the candidate label set. A dual-marginal KL projection then constructs a structured teacher that incorporates momentum-refined class beliefs while preserving the proposal's temporal occupancy. A single plan-level KL objective aligns the full-space predictor with this teacher. Our analysis characterizes when candidate re-solving differs from masking and shows that, under the stated construction, the joint objective decomposes into class-marginal and class-conditional temporal supervision. We construct VCMIPL benchmarks from Breakfast, DoTA, and FineAction using model-generated candidate labels and evaluate the method across four feature representations. Extensive experimental results demonstrate that PIVOTMIPL outperforms existing MIPL algorithms in both effectiveness and efficiency.
Problem

Research questions and friction points this paper is trying to address.

Video Classification
Multi-Instance Partial-Label Learning
Inexact Supervision
Label Disambiguation
Temporal Evidence Allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-instance partial-label learning
Joint class-time assignment
Occupancy-regularized spherical matching
Dual-marginal KL projection
Video classification
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
L
Lingyu Shen
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China
W
Wei Tang
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, United Arab Emirates
F
Fakhri Karray
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, United Arab Emirates
Min-Ling Zhang
Min-Ling Zhang
Professor, School of Computer Science and Engineering, Southeast University, China
Artificial IntelligenceMachine LearningData Mining