Score
Designs, implements, or evaluates reranking procedures that select among multiple model outputs by estimating expected utility (minimum Bayes risk) and combining model likelihoods with MBR utility scores to choose candidates that maximize task-specific accuracy; this includes preferring holistic or long-form candidates over segmented ones and recovering semantic accuracy lost when relying solely on likelihood.
To address the challenge of balancing computational cost and translation quality in machine translation re-ranking, this paper pioneers a Bayesian optimization formulation for candidate translation selection. We propose a multi-fidelity re-ranking framework: lightweight noisy surrogate models perform rapid initial filtering, while a high-accuracy distilled scorer—guided by COMET-Kiwi—conducts critical validation. The method integrates Bayesian optimization, multi-fidelity modeling, and COMET-Kiwi–guided scorer distillation. Experiments on the WMT22 benchmark demonstrate that our approach achieves comparable COMET-Kiwi scores to a baseline requiring 180 precise evaluations using only 70 such evaluations, substantially improving the cost–effectiveness trade-off. Our core contributions are twofold: (1) the first formulation of MT re-ranking as an efficient Bayesian optimization problem, and (2) the design of a scalable, multi-fidelity scoring coordination mechanism that jointly leverages surrogates and distilled scorers.
This work addresses the limitation of traditional Minimum Bayes Risk (MBR) decoding, which overlooks the bidirectional relationship between hypotheses and reference texts when employing asymmetric evaluation metrics. The authors introduce, for the first time, a noisy channel model into the MBR framework, decomposing the risk into four components: the likelihoods of “hypothesis→reference” and “reference→hypothesis” directions along with their respective priors. This formulation explicitly models metric asymmetry and provides a unified interpretation of existing MBR variants. By integrating Bayesian inference, pseudo-reference sampling, and expectation computation over metrics such as BLEU and COMET, the method achieves channel-level interpretability and flexible weighting. Experiments demonstrate consistent contributions across channels in diverse tasks, with properly weighted combinations significantly outperforming standard MBR decoding.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
This paper addresses the lack of theoretical foundation for Minimum Bayes Risk (MBR) decoding in large language model generation. We introduce, for the first time, a bias–diversity decomposition framework: bias quantifies the alignment between a utility function and human evaluation, while diversity measures estimation disagreement across utility functions. Based on this, we design a pseudo-bias metric and propose Metric-augmented MBR (MAMBR), which dynamically weights utility functions to enhance diversity without requiring pseudo-references. Extensive experiments across summarization, translation, and dialogue generation demonstrate that both bias and diversity strongly correlate with generation quality. MAMBR consistently improves standard automatic metrics—including BLEU and BERTScore—across tasks. The implementation is publicly available.
This paper addresses selective classification in high-stakes settings, aiming to minimize the rejection (indecision) rate under a user-specified misclassification constraint—potentially stricter than the Bayes optimal error. We propose a threshold-adaptive framework grounded in statistical learning theory and risk-controlling optimization. By constructing tight confidence sets and dynamically adjusting decision boundaries, we establish, for the first time, theoretical guarantees on achieving the optimal rejection rate under stringent misclassification constraints. Our work challenges the conventional belief that the Bayes error is an insurmountable lower bound, and instead derives the fundamental trade-off between misclassification rate and rejection cost. Experiments demonstrate substantial reductions in misclassification—approaching zero—on hard classification tasks, while incurring only negligible rejection rates; the gain in misclassification reduction far outweighs the cost introduced by rejection.
This work addresses the critical challenge of accurately modeling preference functions that aggregate multidimensional criteria into holistic judgments in settings such as admissions and medical diagnosis. Departing from conventional assumptions of linearity or strong structural forms, the paper proposes the first robust nonparametric learning algorithm that achieves optimal performance without requiring any prior knowledge of the preference structure, assuming only monotonic non-decreasing behavior across each criterion. Theoretical analysis demonstrates the severe consequences of common model misspecifications, while experiments on both synthetic and real-world data confirm that the method maintains statistical efficiency under linear preferences and reliably recovers true evaluator preferences in general cases. Notably, the approach effectively uncovers key behavioral differences between human evaluators and large language models in their assessment strategies.
Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.
该研究提出Quit策略,通过提前终止候选生成和重排序过程来减少神经机器翻译中的计算瓶颈,提高效率同时保持翻译质量。
本文提出一种基于置信区域的筛选框架,用于解决模拟系统可接受性问题,保证高概率筛选出所有或每个可接受系统,并支持并行化。
This study addresses the time-consuming quantification of conditional probability tables in Bayesian networks for software decision-making and the lack of empirical comparisons among available methods. Within a software development organization, we systematically compare two semi-automated quantification approaches: the weighted sum algorithm (WSA) and the ranked nodes method (RNM). Employing expert-driven modeling and model walkthrough techniques, their performance is evaluated in feature selection and UI design decisions. Results indicate that although both methods yield similar rankings, their uncertainty distributions differ significantly. This work reveals the critical impact of method selection on uncertainty representation, demonstrating that shortlists alone are insufficient to establish model equivalence. It further emphasizes the necessity of calibrating complete output distributions in conjunction with node semantics.