Reframing Long-Tailed Learning via Loss Landscape Geometry

📅 2026-03-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance trade-off in long-tailed learning, where models tend to overfit on head classes and forget tail classes. Inspired by continual learning, the authors propose a novel framework that leverages the geometric properties of the loss landscape. For the first time, loss landscape flatness is explicitly incorporated into long-tailed learning without relying on external data or pre-trained models. The framework introduces a Grouped Knowledge Preservation module to retain class-group-specific knowledge and a Grouped Sharpness-Aware module to shape the loss landscape, jointly guiding optimization toward a shared, flat minimum beneficial for all classes. Extensive experiments on four benchmark datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches, confirming its effectiveness and generalizability.

Technology Category

Machine Learning: Life-Long and Continual LearningSearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
📝 Abstract
Balancing performance trade-off on long-tail (LT) data distributions remains a long-standing challenge. In this paper, we posit that this dilemma stems from a phenomenon called "tail performance degradation" (the model tends to severely overfit on head classes while quickly forgetting tail classes) and pose a solution from a loss landscape perspective. We observe that different classes possess divergent convergence points in the loss landscape. Besides, this divergence is aggravated when the model settles into sharp and non-robust minima, rather than a shared and flat solution that is beneficial for all classes. In light of this, we propose a continual learning inspired framework to prevent "tail performance degradation". To avoid inefficient per-class parameter preservation, a Grouped Knowledge Preservation module is proposed to memorize group-specific convergence parameters, promoting convergence towards a shared solution. Concurrently, our framework integrates a Grouped Sharpness Aware module to seek flatter minima by explicitly addressing the geometry of the loss landscape. Notably, our framework requires neither external training samples nor pre-trained models, facilitating the broad applicability. Extensive experiments on four benchmarks demonstrate significant performance gains over state-of-the-art methods. The code is available at:https://gkp-gsa.github.io/.
Problem

Research questions and friction points this paper is trying to address.

long-tailed learning
tail performance degradation
loss landscape
class imbalance
model overfitting
Innovation

Methods, ideas, or system contributions that make the work stand out.

loss landscape geometry
long-tailed learning
flat minima
grouped knowledge preservation
sharpness-aware optimization
💼 Related Jobs
No related jobs found.
S
Shenghan Chen
Shandong University
Y
Yiming Liu
Shandong University
Y
Yanzhen Wang
Shandong University
Yujia Wang
Yujia Wang
Zhejiang Sci-Tech University
X
Xiankai Lu
Shandong University