🤖 AI Summary
This work addresses the tendency of neural networks to bias toward majority classes in imbalanced datasets by proposing a novel loss function that, for the first time, incorporates cardinality-based invariants from metric space—such as magnitude and spread—into the training process to enhance effective data diversity. By explicitly quantifying and optimizing the geometric structure of sample distributions, the method significantly improves the model’s ability to recognize minority classes. Experiments on both synthetic and real-world materials science imbalanced datasets demonstrate consistent and substantial gains in both minority-class performance and overall evaluation metrics, thereby validating the effectiveness and generalizability of the proposed approach.
📝 Abstract
Class imbalance is a common and pernicious issue for the training of neural networks. Often, an imbalanced majority class can dominate training to skew classifier performance towards the majority outcome. To address this problem we introduce cardinality augmented loss functions, derived from cardinality-like invariants in modern mathematics literature such as magnitude and the spread. These invariants enrich the concept of cardinality by evaluating the `effective diversity'of a metric space, and as such represent a natural solution to overly homogeneous training data. In this work, we establish a methodology for applying cardinality augmented loss functions in the training of neural networks and report results on both artificially imbalanced datasets as well as a real-world imbalanced material science dataset. We observe significant performance improvement among minority classes, as well as improvement in overall performance metrics.