mbr reranking

Designs, implements, or evaluates reranking procedures that select among multiple model outputs by estimating expected utility (minimum Bayes risk) and combining model likelihoods with MBR utility scores to choose candidates that maximize task-specific accuracy; this includes preferring holistic or long-form candidates over segmented ones and recovering semantic accuracy lost when relying solely on likelihood.

mbrreranking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Bayesian Optimization Approach to Machine Translation Reranking

Nov 14, 2024
JC
Julius Cheng
🏛️ University of Cambridge | Karlsruhe Institute of Technology | ETH Zürich

To address the challenge of balancing computational cost and translation quality in machine translation re-ranking, this paper pioneers a Bayesian optimization formulation for candidate translation selection. We propose a multi-fidelity re-ranking framework: lightweight noisy surrogate models perform rapid initial filtering, while a high-accuracy distilled scorer—guided by COMET-Kiwi—conducts critical validation. The method integrates Bayesian optimization, multi-fidelity modeling, and COMET-Kiwi–guided scorer distillation. Experiments on the WMT22 benchmark demonstrate that our approach achieves comparable COMET-Kiwi scores to a baseline requiring 180 precise evaluations using only 70 such evaluations, substantially improving the cost–effectiveness trade-off. Our core contributions are twofold: (1) the first formulation of MT re-ranking as an efficient Bayesian optimization problem, and (2) the design of a scalable, multi-fidelity scoring coordination mechanism that jointly leverages surrogates and distilled scorers.

Improving cost-performance with multi-fidelity proxy scorersOptimizing machine translation reranking using Bayesian methodsReducing computational costs in translation scoring models

This work addresses the limitation of traditional Minimum Bayes Risk (MBR) decoding, which overlooks the bidirectional relationship between hypotheses and reference texts when employing asymmetric evaluation metrics. The authors introduce, for the first time, a noisy channel model into the MBR framework, decomposing the risk into four components: the likelihoods of “hypothesis→reference” and “reference→hypothesis” directions along with their respective priors. This formulation explicitly models metric asymmetry and provides a unified interpretation of existing MBR variants. By integrating Bayesian inference, pseudo-reference sampling, and expectation computation over metrics such as BLEU and COMET, the method achieves channel-level interpretability and flexible weighting. Experiments demonstrate consistent contributions across channels in diverse tasks, with properly weighted combinations significantly outperforming standard MBR decoding.

Asymmetric Evaluation MetricsDecodingMinimum Bayes Risk

Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.

Bayesian decision proceduresexperimental designgeneralized posteriors

Theoretical Aspects of Bias and Diversity in Minimum Bayes Risk Decoding

Oct 19, 2024
HK
Hidetaka Kamigaito
🏛️ Nara Institute of Science and Technology | The University of Tokyo

This paper addresses the lack of theoretical foundation for Minimum Bayes Risk (MBR) decoding in large language model generation. We introduce, for the first time, a bias–diversity decomposition framework: bias quantifies the alignment between a utility function and human evaluation, while diversity measures estimation disagreement across utility functions. Based on this, we design a pseudo-bias metric and propose Metric-augmented MBR (MAMBR), which dynamically weights utility functions to enhance diversity without requiring pseudo-references. Extensive experiments across summarization, translation, and dialogue generation demonstrate that both bias and diversity strongly correlate with generation quality. MAMBR consistently improves standard automatic metrics—including BLEU and BERTScore—across tasks. The implementation is publicly available.

Bias-diversity trade-off in quality estimation of hypothesesDiversity's role in explaining inference scaling lawsTheoretical understanding of MBR decoding's performance improvements

Ask for More Than Bayes Optimal: A Theory of Indecisions for Classification

Dec 17, 2024
MN
Mohamed Ndaoud
🏛️ ESSEC Business School | University of Sydney

This paper addresses selective classification in high-stakes settings, aiming to minimize the rejection (indecision) rate under a user-specified misclassification constraint—potentially stricter than the Bayes optimal error. We propose a threshold-adaptive framework grounded in statistical learning theory and risk-controlling optimization. By constructing tight confidence sets and dynamically adjusting decision boundaries, we establish, for the first time, theoretical guarantees on achieving the optimal rejection rate under stringent misclassification constraints. Our work challenges the conventional belief that the Bayes error is an insurmountable lower bound, and instead derives the fundamental trade-off between misclassification rate and rejection cost. Experiments demonstrate substantial reductions in misclassification—approaching zero—on hard classification tasks, while incurring only negligible rejection rates; the gain in misclassification reduction far outweighs the cost introduced by rejection.

Control misclassification rate below Bayes optimal errorExtend selective classification principles to hypothesis testingMinimize indecisions while achieving target classification accuracy

Latest Papers

What's happening recently
View more

This work addresses the critical challenge of accurately modeling preference functions that aggregate multidimensional criteria into holistic judgments in settings such as admissions and medical diagnosis. Departing from conventional assumptions of linearity or strong structural forms, the paper proposes the first robust nonparametric learning algorithm that achieves optimal performance without requiring any prior knowledge of the preference structure, assuming only monotonic non-decreasing behavior across each criterion. Theoretical analysis demonstrates the severe consequences of common model misspecifications, while experiments on both synthetic and real-world data confirm that the method maintains statistical efficiency under linear preferences and reliably recovers true evaluator preferences in general cases. Notably, the approach effectively uncovers key behavioral differences between human evaluators and large language models in their assessment strategies.

evaluator preferencesmodel mismatchmulti-criteria evaluation

Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.

computational costfixed-size benchmarksmodel evaluation

This study addresses the time-consuming quantification of conditional probability tables in Bayesian networks for software decision-making and the lack of empirical comparisons among available methods. Within a software development organization, we systematically compare two semi-automated quantification approaches: the weighted sum algorithm (WSA) and the ranked nodes method (RNM). Employing expert-driven modeling and model walkthrough techniques, their performance is evaluated in feature selection and UI design decisions. Results indicate that although both methods yield similar rankings, their uncertainty distributions differ significantly. This work reveals the critical impact of method selection on uncertainty representation, demonstrating that shortlists alone are insufficient to establish model equivalence. It further emphasizes the necessity of calibrating complete output distributions in conjunction with node semantics.

Bayesian NetworksConditional Probability TablesExpert Elicitation