Score
Designs and analyzes quantitative imbalance metrics and selects appropriate imbalance measures, including constructing monotone metrics that provably decrease under specified balancing moves. Proves existence of balance-improving moves and develops metric properties that enable greedy or local operations to reduce and ultimately restore global balance (with provable restoration guarantees in two-dimensional settings).
This paper addresses data imbalance in regression tasks—where target variables are continuous—a problem extensively studied in classification but lacking systematic investigation in regression. We propose the first taxonomy of resampling methods specifically designed for imbalanced regression. Our framework systematically evaluates oversampling, undersampling, and hybrid strategies across three dimensions: regression models (linear regression, tree-based models, neural networks), learning processes, and specialized evaluation metrics (uM, wMAE). Experimental results demonstrate that judicious resampling significantly improves predictive accuracy in sparse regions of the target space, and that model sensitivity to sampling strategies varies substantially. We uncover mechanistic insights into how resampling operates effectively in continuous output spaces. To foster reproducibility and further research, we publicly release all source code and benchmark datasets. This work establishes a rigorous, extensible analytical framework and practical guidelines for addressing imbalance in regression.
Real-world distance data are often distorted by noise, missing entries, or violations of the triangle inequality, degrading downstream task performance. This work systematically evaluates the effectiveness of metric repair algorithms, explicitly disentangling two critical subproblems: which edges to repair and how to assign their weights. Through large-scale experiments on both real-world and synthetic non-metric graphs, the study empirically demonstrates—for the first time—that repair efficacy depends primarily on the type and proportion of errors, rather than on error magnitude or graph size. Moreover, accurately identifying the set of edges requiring correction is as crucial as appropriately setting their weights. The findings reveal that most existing repair methods fail to recover a true metric and can even degrade performance, thereby challenging prevailing assumptions in the field.
The few-shot and imbalanced (S&I) learning problem suffers from severe generalization degradation and low interpretability due to scarce samples, extreme class imbalance, and ambiguous inter-class feature distributions. This paper proposes the first systematic analytical framework tailored to S&I learning, advocating that quantitative characterization of data properties—such as imbalance ratio and geometric complexity—must precede algorithmic design. The framework unifies multi-dimensional imbalance metrics, data complexity analysis, resampling strategies, classifier adaptation mechanisms, and an interpretable evaluation benchmark. Empirical evaluation on binary and multi-class extreme imbalance benchmarks reveals that classifier selection exerts significantly greater impact on performance than resampling improvements—exposing a fundamental flaw in prevailing heuristic-driven approaches. Our work establishes a theory-guided analytical paradigm and practical design principles for S&I learning, advancing both methodological rigor and empirical reproducibility.
This paper studies the online metric matching problem: given an underlying metric space, $m$ initially idle servers must be matched, in real time and irrevocably, to a sequence of $n$ arriving requests, minimizing total matching distance. Addressing the long-standing open challenge of unbalanced markets ($m eq n$), we deliver the first positive competitive ratio guarantees. Our approach introduces a distributional reduction technique and a prediction-augmented framework, integrating stochastic geometric analysis, prediction calibration mechanisms, and unified competitive-regret analysis. In Euclidean spaces of dimension $d geq 3$, we achieve an $O(1)$ competitive ratio—significantly improving upon the prior best $O((log log log n)^2)$ upper bound. Concurrently, we establish a near-optimal $O(sqrt{n})$ regret bound and ensure robustness under worst-case inputs.
This study addresses the challenge posed by severe class imbalance in anomaly detection, which complicates the interpretation and comparison of common evaluation metrics. The authors systematically analyze the behavior of AUROC, AUPR, F1-score, and Matthews Correlation Coefficient (MCC) across varying anomaly ratios and introduce a novel "metric landscape" visualization technique. This approach reveals, for the first time, each metric’s inherent preference for true positive rate versus true negative rate and how their stability varies with imbalance levels. By modeling the relationship between metrics and anomaly prevalence, the work delineates clear applicability boundaries for each metric, thereby providing a principled, interpretable foundation for reliable metric selection in highly imbalanced anomaly detection scenarios.
This work addresses the misalignment between offline evaluation metrics and online performance objectives in industrial applications by establishing a unified theoretical framework that systematically quantifies the relationships among diverse evaluation metrics for the first time. By introducing the concepts of Bayes-optimal sets and regret transfer mechanisms, the study reveals structural asymmetries among metrics and provides a principled classification and relational modeling of metrics with varying mathematical forms. Theoretically characterizing metric consistency and transferability, this research offers novel insights and a methodological foundation for designing offline evaluation systems that are aligned with online objectives and backed by rigorous theoretical guarantees.
This study addresses the problem of dynamic fair allocation of indivisible goods (or tasks) in an online setting, where items arrive sequentially and must be allocated immediately upon arrival. Under a broad range of models—including normalized and non-normalized utilities as well as identical and general additive utility functions—the work designs constructive online algorithms within the competitive analysis framework, targeting multiple fairness criteria such as EF1 and PROP1. For most settings considered, the paper not only presents algorithms achieving optimal competitive ratios but also establishes matching theoretical upper bounds, thereby substantially expanding the theoretical foundations of online fair division.
This work proposes the p-AVL tree, a novel variant of binary search trees that introduces a tunable probability parameter \( p \) to govern the likelihood of performing rotation-based rebalancing at unbalanced nodes. By varying \( p \) from 0 to 1, the structure continuously interpolates between an ordinary binary search tree (when \( p = 0 \)) and a strictly balanced AVL tree (when \( p = 1 \)). This approach is the first to integrate probabilistic mechanisms into AVL balancing strategies, thereby bridging the behavioral gap between deterministic balancing and completely unbalanced structures. Empirical evaluations on randomly generated insertion sequences systematically assess structural properties—such as height, average node depth, and global imbalance metric \( \sigma \)—across different \( p \) values. The results demonstrate that even minimal non-zero values of \( p \) yield substantial improvements in balance, highlighting the sensitivity and efficacy of the p-AVL family in structural evolution.