Score
Designs and evaluates choices of learning objectives and target formulations for training systems, including selection among target variables, loss/score functions, velocity or noise adjustments, and posterior-matching strategies. Builds analyses that compare utility versus risk tradeoffs, quantify effects on downstream performance and disclosure risk, and recommend objective choices or default objectives under operational constraints.
Traditional financial decision-making relies on fixed optimization objectives, rendering it ill-suited for dynamic markets. Existing approaches that switch among latent regimes are often sensitive to noise, exhibit latency, and incur high turnover. To address these limitations, this work proposes the DOSS framework, which eschews latent state modeling and instead formulates objective selection as a sequential classification task. DOSS dynamically selects the optimal objective—among candidates such as return maximization, loss aversion, and risk-adjusted performance—based on interpretable statistical summaries of recent returns. By integrating confidence-aware gating with a rule-driven large language model (LLM) supervision mechanism, DOSS enables forward-looking, time-leakage-free objective switching. This approach significantly reduces portfolio turnover and operational instability while enhancing overall decision-making performance.
This work addresses the mismatch between post-training objectives and the best-of-N evaluation strategy used during deployment, particularly in realistic scenarios where the training sampling budget is substantially lower than that at test time. To bridge this gap, the authors propose a family of tail-extrapolated estimators—TEA and Prefix-TEA—that leverage structural assumptions about the upper tail of the reward distribution. By extrapolating tail statistics from limited samples and integrating moment-based debiasing with advantage function construction, these estimators effectively approximate the policy gradient of the best-of-N objective. Extensive experiments across diverse language models, reward models, and datasets demonstrate that the proposed approach significantly improves best-of-N performance under various training–testing budget configurations, achieving, for the first time, effective alignment between post-training objectives and deployment performance under low training budgets.
Evaluating fairness in machine learning systems faces challenges including ambiguous metric definitions and the difficulty of quantifying trade-offs between utility and fairness. This paper proposes a model-agnostic, scalable multi-objective evaluation framework that unifies the characterization of utility–fairness trade-offs across multiple fairness dimensions—such as group and individual fairness—and supports model comparison and decision-making under single or multiple constraints. Innovatively integrating convergence, system capacity, and diversity, the framework introduces a joint visualization paradigm combining radar charts and calibrated metric scales. It also features a modular evaluation interface compatible with mainstream ML models. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that the framework significantly enhances interpretability of evaluation outcomes, cross-model comparability, and decision-support capability.
This work addresses the critical challenge of accurately modeling preference functions that aggregate multidimensional criteria into holistic judgments in settings such as admissions and medical diagnosis. Departing from conventional assumptions of linearity or strong structural forms, the paper proposes the first robust nonparametric learning algorithm that achieves optimal performance without requiring any prior knowledge of the preference structure, assuming only monotonic non-decreasing behavior across each criterion. Theoretical analysis demonstrates the severe consequences of common model misspecifications, while experiments on both synthetic and real-world data confirm that the method maintains statistical efficiency under linear preferences and reliably recovers true evaluator preferences in general cases. Notably, the approach effectively uncovers key behavioral differences between human evaluators and large language models in their assessment strategies.
When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.
This study addresses the lack of a systematic framework for identifying critical input variables and conducting sensitivity analysis under uncertainty in complex simulations, particularly in military decision-making contexts. The authors propose a unified sensitivity analysis framework that integrates local and global methods—including variance-based, derivative-based, screening, and uncertainty quantification techniques—and strategically maps these approaches to specific decision objectives such as factor prioritization, fixing, variance reduction, and mapping. Innovatively, the framework introduces a “sensitivity audit” mechanism to enhance traceability of model assumptions and promote responsible model usage. By providing a structured guide for high-dimensional, complex simulation systems, this work significantly improves model interpretability, transparency, and the credibility of decisions derived from such models.
This work addresses the lack of interpretability in existing methods regarding the influence of hyperparameters in multi-objective optimization. The authors propose a novel game-theoretic framework that, for the first time, integrates Shapley effects with the Pareto front to enable objective-aware global sensitivity analysis, thereby uncovering key hyperparameters and their interactions under different optimization objectives. This approach not only identifies efficient hyperparameter configurations and substantially reduces the search space but also facilitates early-stage model performance estimation. The effectiveness and generalizability of the framework are empirically validated across three distinct neural network architectures and tasks.
This study addresses the fundamental trade-off in large language model evaluation among evaluator coupling (γ), policy diversity (measured by entropy H), and few-shot reliability (quantified by the coefficient of variation CV). Extending empirical conditions from five to eleven, the work systematically quantifies the interplay among these three factors and introduces the first standardized benchmark dataset for evaluation. Results reveal a strong negative correlation between γ and H (r = −0.989), indicating that low coupling is accompanied by high measurement noise. Notably, no experimental setting simultaneously achieves γ < 0.2 and CV(N=5) < 0.3, highlighting an inherent tension among these desiderata. The analysis also uncovers anomalous patterns linked to version drift in GPT-4o, offering empirical grounding for the design of more robust and reliable evaluation frameworks.