🤖 AI Summary
This study addresses the degradation of model performance caused by spurious correlations and label noise under subgroup shifts. To tackle this, we propose POTER, a framework grounded in optimal transport theory. Rather than relying on loss signals susceptible to noise interference, POTER computes sample weights based on the geometric relationships of instance-level transport alignment to assess importance. Notably, it achieves robust learning through a single round of empirical risk minimization (ERM), departing from conventional retraining paradigms. Experimental results demonstrate that POTER attains state-of-the-art worst-group accuracy across standard benchmarks and noisy settings, effectively mitigating the adverse effects of label contamination in minority groups.
📝 Abstract
Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify informative samples, but these signals can become severely distorted under label noise: mislabeled samples may also incur large losses and contaminate subsequent reweighting or retraining. Despite its practical importance, this intersection remains largely underexplored. We propose POTER, a reweighting framework based on optimal transport that derives sample importance from the transport geometry between the training distribution and a reference distribution constructed from limited validation group annotations. By measuring alignment at the individual-sample level rather than relying on loss, POTER downweights mislabeled or strongly bias-aligned samples while assigning higher importance to samples better aligned with the reference distribution. In addition, POTER requires only a single ERM training stage, moving beyond the retraining paradigm common in recent work. Across standard benchmarks and noisy-label settings, POTER achieves state-of-the-art worst-group accuracy, including cases where label corruption is concentrated within minority subgroups.