Score
Designs and implements methods to record, preprocess, and summarize per-example loss values across training iterations or epochs, producing features, metrics, and visualizations of each sample's loss sequence. Analyzes those loss trajectories to detect anomalous or non-convergent dynamics, identify learning or labeling issues, and rank or select samples for relabeling, review, or further investigation.
To address scalability challenges of deep learning models under resource constraints, this paper proposes an efficient coreset selection method based on loss trajectory alignment. The core innovation is the Loss Trajectory Correlation (LTC) metric—a novel measure quantifying the dynamic relationship between individual training samples and the validation loss evolution during training—thereby transforming coreset construction into a lightweight byproduct computation task. Crucially, LTC requires no additional gradient computations or model retraining. It achieves, for the first time, robust cross-architecture transferability (across ResNet, VGG, DenseNet, and Swin) and enables interpretable analysis of training dynamics. On CIFAR-100 and ImageNet-1K, models trained on subsets selected by LTC match or surpass state-of-the-art methods in accuracy (within <1% gap), exhibit minimal architecture generalization degradation (<2%), and incur substantially lower computational overhead.
To address the degradation of prediction probability calibration in deployed image classification models due to concept drift, this paper proposes an online calibration monitoring method that requires no access to model internals—only predicted probabilities and ground-truth labels. Our approach introduces, for the first time, a Cumulative Sum (CUSUM) control chart with dynamic control limits into calibration monitoring. It computes cumulative deviations of calibration error over time and adaptively adjusts detection thresholds to enable early warning of calibration loss. Compared to static-threshold methods, our framework significantly enhances sensitivity to temporal distribution shifts and accelerates response to emerging miscalibration. We validate its effectiveness and robustness across multiple image classification benchmarks under diverse concept drift scenarios. The proposed method establishes a scalable, black-box-compatible paradigm for trustworthy model deployment, enabling continuous, lightweight calibration assessment without architectural or training modifications.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
This paper addresses the lack of systematic guidance for selecting loss functions and evaluation metrics in deep learning. Methodologically, it conducts a comprehensive analysis of mathematical properties, gradient behaviors, scale sensitivities, and optimization stability of canonical losses and metrics—including cross-entropy, MSE, IoU, BLEU, F1, and Dice—across 12 mainstream task categories (e.g., regression, classification, CV, NLP). Based on this analysis, it constructs a task-driven “loss–metric alignment matrix.” Its key contribution is the first cross-task, interpretable selection framework that formally bridges the gap between theoretical design principles and empirical engineering practice. The framework provides principled, semantics-aware criteria for matching losses to metrics according to task objectives and optimization dynamics. It has been widely adopted in industry as a standard reference for model development, significantly reducing trial-and-error overhead and improving evaluation consistency across teams and applications.
This work addresses the mismatch between conventional machine learning practices and the specific performance requirements of clinical tasks in healthcare settings. Traditional approaches rely on differentiable validation losses for model optimization, which often fail to align with clinically meaningful outcomes. To bridge this gap, the paper proposes replacing standard loss functions with non-differentiable yet clinically interpretable custom metrics to guide critical optimization decisions—such as hyperparameter selection and training termination—thereby redefining the model validation pipeline. In two controlled experiments, models optimized using this framework demonstrated significantly superior performance on key clinical tasks compared to those guided by conventional differentiable validation losses. This approach overcomes the inherent limitation of relying solely on differentiable objectives and better aligns medical AI development with real-world clinical goals.
This work addresses the detrimental impact of annotation errors—such as mislabeling and temporal misalignment—in video datasets on the performance of models for temporally sensitive tasks. To tackle this issue, the authors propose a model-agnostic approach based on dynamic loss trajectory analysis: by tracking the average loss of each frame across multiple training checkpoints, they construct cumulative sample loss (CSL) trajectories that serve as frame-level learnability fingerprints. This method enables the identification of hard-to-learn samples without requiring ground-truth error labels. Experiments on the EgoPER and Cholec80 datasets demonstrate that the proposed technique effectively detects subtle annotation inaccuracies, exhibiting strong generalization capability and practical utility in real-world scenarios.
This study addresses the coupled challenges of class imbalance and class overlap in software defect prediction, which jointly impair model training dynamics and performance. The authors propose the first interaction-aware protocol for analyzing training dynamics under these intertwined data quality issues. By training a fixed multilayer perceptron (MLP) under three conditions—imbalance only, overlap only, and their coupling—the protocol systematically records training trajectories. Integrating effect size analysis, sensitivity analysis, and rule-based classification, it constructs a taxonomy of training dynamic patterns. The work uncovers distinctive neural network behaviors specific to the coupled scenario, offering empirical insights and novel diagnostic tools to enhance the understanding, evaluation, and refinement of defect prediction models.
研究通过分析Jupyter笔记本中的反馈语句,采用CodeBERT嵌入和质性研究方法,提出了一种针对机器学习中潜在静默错误的检查机制。
Dynamic data pruning struggles to efficiently estimate per-sample loss under complex models or loss functions, limiting its practical applicability. This work proposes the Batch Loss Score (BLS), which treats batch-level losses as noisy observations of individual sample losses and employs exponential moving average (EMA) to filter out noise induced by varying batch compositions, thereby assigning each sample an importance score. Requiring only three lines of code for integration and a single line for proxy adaptation, BLS is theoretically grounded in a low-pass filtering perspective that ensures its effectiveness. Evaluated across 14 datasets, 11 tasks, and 18 models, BLS achieves lossless pruning of 20%–50% of training samples, substantially improving training efficiency.
This study addresses the absence of effective evaluation mechanisms for learning signals during data sampling in time series foundation model pretraining. To this end, it proposes a static data selection framework that employs a reference predictor to score samples and retains those within an intermediate interval. By establishing a theoretical connection between loss and gradient norms, and integrating stratified dataset selection with local Jacobian conditioning analysis, the approach achieves efficient data filtering across varying scales and architectures. This method substantially reduces the number of pretraining windows required while significantly improving MASE and CRPS metrics, outperforming baselines trained on full datasets. Consequently, it establishes an efficient new paradigm for data selection in time series foundation models.