Score
Designs, implements, or evaluates methods that determine whether a particular example (or paired examples across different modalities) was included in a model’s training set by comparing model outputs and candidate samples; practitioners build attacks or analyses that embed outputs and samples in a shared space, align distributions across modalities, and compute membership scores (e.g., likelihood ratios) to make membership decisions.
This work identifies an evaluation bias introduced by data augmentation (e.g., SMOTE, mutation-based augmentation) in scarce-data scenarios—particularly flaky test classification—where augmented samples inadvertently contaminate the test set, severely compromising fairness and reliability assessments. To address this, the authors first empirically identify and validate the critical phenomenon that “augmented data participation in testing” induces systematic evaluation distortion. They then propose a detection framework capable of disentangling training-induced bias from evaluation-induced bias, and design a bias-calibrated evaluation protocol. Experiments across multiple flaky-test benchmark datasets demonstrate that test sets containing augmented samples inflate accuracy by up to 23.7% and introduce F1-score deviations exceeding 0.15. This study establishes both theoretical foundations and practical guidelines for trustworthy model evaluation under data augmentation.
This work addresses the challenge of quantifying prediction uncertainty in generative biomolecular design, where feedback covariate shift undermines conventional uncertainty estimation. We propose the first conformal prediction framework tailored to closed-loop design paradigms. Departing from standard i.i.d. assumptions, our method imposes no structural constraints on either the design algorithm or the regression model, delivering finite-sample statistically valid confidence sets for arbitrary black-box design pipelines. Key innovations include quantile-regression-driven adaptive conformal prediction, explicit modeling of feedback-induced distributional shift, and robust error calibration. Evaluated on protein and small-molecule design tasks, our approach achieves ≥94.8% empirical coverage at the 95% nominal confidence level—substantially outperforming standard conformal methods (which drop to as low as 72%)—while maintaining high predictive accuracy.
To address copyright infringement and transparency concerns arising from unauthorized use of third-party data in machine learning model training, this paper proposes the first general-purpose, task-agnostic data usage auditing framework for black-box models. Methodologically, it innovatively integrates arbitrary black-box membership inference techniques with a custom sequential probability ratio test (SPRT), enabling zero assumptions about downstream tasks, strict control over false positive rates (tunable within 0.5%–5%), and cross-model generalization. The framework features a model-agnostic interface, supporting heterogeneous architectures including image classifiers and multimodal large language models. Extensive experiments on ImageNet classifiers and multimodal foundation models demonstrate an average detection accuracy exceeding 92%, with false positive rates consistently meeting user-specified thresholds. This work significantly enhances the quantifiability and reliability of training data provenance auditing.
In A/B testing, rigorously evaluating novel estimation algorithms—when the true treatment effect is unobserved—remains a fundamental methodological challenge. This paper establishes, for the first time, a comprehensive theoretical framework for estimation and inference based on sample splitting: it derives the asymptotic distribution of sample-split estimators and characterizes their bias structure relative to full-sample performance; introduces a bias–variance trade-off analytical paradigm and proposes a correction-based confidence interval construction method. Leveraging statistical inference, asymptotic theory, Monte Carlo simulation, and empirical validation, the framework enables robust, production-grade evaluation of new algorithms within industrial A/B testing platforms. Theoretical results are thoroughly validated via simulation studies. The proposed infrastructure enhances A/B testing methodology by delivering an interpretable, reproducible, and deployable evaluation system.
To address computational redundancy and high annotation costs in large-scale data training, this paper proposes an influence-function-based data subset selection method—the first systematic application of influence function theory to efficient training set pruning. The method models each training sample’s impact on model parameters via logistic regression, enabling principled ranking and selection of the most representative subset. On binary classification tasks, the selected subset achieves full-training accuracy using only 10% of the original data; remarkably, with 60% of the data, it surpasses the full-training baseline in accuracy. This approach substantially reduces computational overhead while preserving model performance. Crucially, it offers a novel, interpretable, and scalable paradigm for small-sample efficient training—grounded in theoretically justified influence estimation rather than heuristic sampling.
Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction following, or pairwise preference, and these scores are often used to support claims about deployed alignment. This paper argues that deployment-relevant alignment cannot be inferred from model-level evaluation alone. Alignment claims should instead be indexed to the level at which evidence is collected: model-level, response-level, interaction-level, or deployment-level. Two studies support this position. First, a structured audit of eleven alignment benchmarks, extended to a sixteen-benchmark corpus, dual-coded against an eight-dimension rubric with Cohen's kappa = 0.87, finds that user-facing verification support is absent across every benchmark examined, while process steerability is nearly absent. The few interactional benchmarks identified, including tau-bench, CURATe, Rifts, and Common Ground, remain fragmented in coverage, and benchmark construction rather than data source determines what is measured. Second, a blinded cross-model stress test using 180 transcripts across three frontier models and four scaffolds finds that the same verification scaffold raises one model's verification support to ceiling while leaving another categorically unchanged. This shows that scaffold efficacy is model-dependent and that the gap identified by the audit cannot be closed at the model level alone. We propose a system-level evaluation agenda: alignment profiles instead of single scores, fixed-scaffolding protocols for comparable interactional evaluation, and reporting templates that make the inferential distance between evaluation evidence and deployment claims explicit.
This study addresses the limited diversity and perceived relevance of problem domains in software modeling instruction, which often undermine student motivation and inclusivity. Through parallel surveys of 90 students and 22 instructors, combined with quantitative and qualitative analyses, it reveals a significant mismatch between instructor assumptions and student preferences: learners prioritize socially relevant problem contexts and value autonomy in topic selection. Furthermore, their sense of engagement markedly increases when they perceive their feedback has been explicitly incorporated. The work proposes a learner-centered strategy for selecting problem domains that foregrounds social relevance and autonomy as critical enablers of inclusive learning. It also highlights how seemingly minor instructional design choices can inadvertently foster exclusion, offering empirical insights and actionable guidance for improving pedagogical practice.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
This study reveals that large language models struggle to effectively assess the veracity of statistical evidence when integrating multi-source information, exhibiting a tendency to rely on superficial stylistic cues in methodological text rather than numerical plausibility when judging source credibility. The work identifies a previously undocumented “cognitive alignment” bias—where models prefer sources with strong analytical register over those with content consistency. Employing interpretable techniques including causal tracing, linear probing (AUC: 0.83–0.92), and component-level attribution, the authors replicate this blind spot across five mainstream models through cross-model and cross-domain experiments. Further analysis localizes the issue to a methodology-register gating mechanism and demonstrates that neither prompt engineering nor post-training interventions adequately mitigate the bias, instead raising concerns about model generalization.