Score
Designs and evaluates algorithms, datasets, and experimental protocols that measure and improve model performance under distribution shift across domains; builds cross-domain benchmarks, domain sampling strategies, multi-domain experiments, and robustness metrics to assess and compare generalization to unseen environments.
Existing synthetic data evaluation lacks unified, transferable quantitative metrics. This paper proposes a novel evaluation framework grounded in generalized cross-validation (GCV) and domain transfer learning. It constructs a cross-dataset performance matrix and defines two core metrics: *fidelity*, quantifying distributional similarity between synthetic and real data; and *generalization coverage*, measuring the task-transfer capability of synthetic data across diverse real-world source domains. The framework is model-agnostic and enables normalized, comparative evaluation of detectors such as YOLOv5s across heterogeneous datasets—including Virtual KITTI, KITTI, and BDD100K. Experiments demonstrate that the method effectively quantifies synthetic data quality, significantly enhancing evaluation generality, comparability, and utility for model optimization. It establishes a scalable, reproducible, and standardized evaluation paradigm for synthetic data development.
Training large language models on heterogeneous multi-source data (e.g., Wikipedia, GitHub) suffers from sampling imbalance and loss imbalance, leading to high gradient variance and degraded generalization. Method: This paper systematically analyzes the complementary roles of sampling weights and loss weights in suppressing gradient variance and narrowing the generalization gap. We propose a joint optimization framework grounded in linear regression theory and SGD dynamics, deriving principled co-design criteria for both weight types through theoretical analysis and empirical validation. Contribution/Results: Our key insight is the first formal characterization that sampling and loss weights are not independently tunable but must be jointly configured to simultaneously ensure gradient stability and strong cross-domain generalization. Experiments demonstrate that our method significantly reduces training variance, accelerates convergence, and improves out-of-distribution generalization across diverse domains.
This work addresses the challenging problem of model performance estimation under distribution shift when ground-truth labels are unavailable. The authors propose Fused Reference Alignment Prediction (FRAP), a novel approach that synergistically combines the generalization capability of external foundation models with the domain-specific expertise of the target task model. FRAP calibrates prediction distributions via temperature scaling, aligns them by minimizing KL divergence, and generates a robust, domain-adaptive pseudo-label reference distribution through confidence-weighted fusion. Extensive experiments demonstrate that FRAP consistently and substantially outperforms existing performance estimation methods across diverse datasets and model architectures, achieving stable and significant improvements.
This study addresses the lack of generality and interpretability in applicability domain (AD) estimation for machine learning models. We propose a unified AD assessment framework based on kernel density estimation (KDE), which quantifies the distance of a query sample from the training data distribution in feature space and establishes a quantitative relationship among distance, prediction error, and uncertainty. Chemical prior knowledge is incorporated to calibrate the AD decision threshold. To our knowledge, this is the first method enabling consistent, cross-model and cross-task AD evaluation across diverse models—including random forests (RF), gradient-boosted decision trees (GBDT), and graph neural networks (GNN)—and heterogeneous materials datasets (crystals, molecules, alloys). Experiments demonstrate that large KDE-derived distances strongly correlate with high prediction residuals and elevated uncertainty estimates. An open-source toolkit enables automated in-domain/out-of-domain classification. The implementation and documentation are publicly available.
This work addresses the critical challenge of evaluating model generalization in high-stakes scenarios with scarce labels, where existing methods lack reliable, label-free metrics for pre-deployment model selection and post-deployment performance monitoring. To bridge this gap, the study introduces, for the first time, the internal causal circuit mechanisms of Vision Transformers into generalization assessment, proposing two novel unsupervised metrics: Dependency Depth Bias and Circuit Shift Score. The former quantifies depth-wise biases in representational dependency structures, while the latter measures changes in circuit stability under distribution shifts. Extensive experiments across diverse tasks demonstrate that these metrics achieve substantially higher correlations with true generalization performance—improving by 13.4% and 34.1% on average over current approaches—thereby significantly enhancing the reliability of generalization prediction without requiring ground-truth labels.
This paper addresses the problem of evaluating and ranking model generalization under distribution shift in the absence of test-set labels—covering both dataset-centric (evaluating a single model across multiple test sets) and model-centric (ranking multiple models on a single test set) deployment scenarios. We propose a hybrid unsupervised evaluation metric that jointly leverages prediction confidence and inter-class dispersion, and introduce the nuclear norm as a novel, efficient, and robust unified measure computed directly from the model’s output probability distributions. Unlike prior approaches, our method requires no ground-truth labels and imposes no architectural assumptions. Extensive experiments demonstrate that it consistently outperforms confidence-only or dispersion-only baselines across diverse settings—including multi-task learning, various distribution shifts, class imbalance, and real-world datasets—achieving superior generalizability and practical utility.
This work addresses two key limitations in multi-domain algorithm performance evaluation: (i) the neglect of user preferences in assessment, and (ii) the masking of domain-specific performance disparities by conventional arithmetic averaging. To this end, we propose a user-preference-parameterized weighted scoring framework. Methodologically, we introduce, for the first time, a continuous family of scoring functions to model performance distributions; integrate probability measures with normalized confusion matrices; rigorously define four critical domain types—easiest, hardest, dominant, and bottleneck—and prove that only specific scoring functions preserve weighted mean consistency. Our contributions include: (i) establishing a general theoretical foundation for multi-domain performance analysis; (ii) developing a visualization toolkit tailored to binary classification tasks; and (iii) enabling fine-grained, interpretable performance decomposition—thereby substantially enhancing transparency and practical utility in cross-domain evaluation. (149 words)
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
Diffusion models suffer from poor generalization in small-target-domain transfer learning, while test-time guidance methods incur high computational overhead and compromise sample diversity. To address these challenges, we propose DogFit, a domain-guided fine-tuning framework. Its core innovation lies in internalizing test-time guidance into the fine-tuning stage: lightweight conditional encoders dynamically inject domain-aware guidance offsets, and two scheduling strategies—late-start and truncation—implicitly balance fidelity and diversity during training. Built upon DiT/SiT architectures, DogFit leverages the strong marginal estimation capability of pretrained unconditional source-domain models and enables controllable generation with a single forward pass. Experiments across six target domains demonstrate that DogFit significantly outperforms existing guidance methods, achieving substantial improvements in FID and FDDINOV2 scores, reducing sampling TFLOPS by up to 2×, and incurring zero additional inference-time computation.