Score
Design and implement decoders and rerankers that select output hypotheses by computing and minimizing expected loss (minimum Bayes risk) over candidate sets such as n‑best lists, lattices, or sampled outputs; the expected‑risk computation can use error metrics (e.g., CER/WER) or candidate‑similarity measures and may be approximated with auxiliary scorers (e.g., pseudo‑log‑likelihood models). Build noisy‑channel formulations and decompositions that combine multiple posterior sources (for example acoustic and language‑model posteriors), implement MBR‑CER and MBR‑reranking procedures, and analyze their performance and robustness across noise conditions and domain shifts.
This work addresses the limitation of traditional Minimum Bayes Risk (MBR) decoding, which overlooks the bidirectional relationship between hypotheses and reference texts when employing asymmetric evaluation metrics. The authors introduce, for the first time, a noisy channel model into the MBR framework, decomposing the risk into four components: the likelihoods of “hypothesis→reference” and “reference→hypothesis” directions along with their respective priors. This formulation explicitly models metric asymmetry and provides a unified interpretation of existing MBR variants. By integrating Bayesian inference, pseudo-reference sampling, and expectation computation over metrics such as BLEU and COMET, the method achieves channel-level interpretability and flexible weighting. Experiments demonstrate consistent contributions across channels in diverse tasks, with properly weighted combinations significantly outperforming standard MBR decoding.
This work addresses the gap between the empirical success and theoretical underpinning of Minimum Bayes Risk (MBR) decoding. We establish, for the first time, a convergence theory for MBR decoding under a finite reference hypothesis set. Leveraging statistical learning theory, the law of large numbers, and the Bayesian decision framework, we rigorously prove that MBR decoding converges to the optimal solution at rate $O(n^{-1/2})$ with high probability. Moreover, under identical assumptions, MBR strictly outperforms Maximum A Posteriori (MAP) decoding. By quantifying the expected utility gap between MBR and MAP, we identify the fundamental source of MBR’s robustness advantage. Our analysis provides the first probabilistically guaranteed convergence rate for MBR decoding, thereby bridging a critical theoretical void and formally justifying its practical efficacy.
This paper addresses the lack of theoretical foundation for Minimum Bayes Risk (MBR) decoding in large language model generation. We introduce, for the first time, a bias–diversity decomposition framework: bias quantifies the alignment between a utility function and human evaluation, while diversity measures estimation disagreement across utility functions. Based on this, we design a pseudo-bias metric and propose Metric-augmented MBR (MAMBR), which dynamically weights utility functions to enhance diversity without requiring pseudo-references. Extensive experiments across summarization, translation, and dialogue generation demonstrate that both bias and diversity strongly correlate with generation quality. MAMBR consistently improves standard automatic metrics—including BLEU and BERTScore—across tasks. The implementation is publicly available.
To address the challenge of balancing test-time performance and inference efficiency in instruction-following large language models (LLMs), this paper proposes a lightweight test-time optimization framework. It integrates Minimum Bayes Risk (MBR) candidate re-ranking decoding with a small-parameter LLM judge (as small as 1.5B) for quality assessment of outputs from a 70B model, coupled with iterative self-training via Direct Preference Optimization (DPO). The key contribution is the first systematic validation that small judges can effectively supervise ultra-large models, establishing a synergistic paradigm between MBR decoding and DPO self-training—eliminating test-time overhead while permanently embedding performance gains. Experiments demonstrate significant improvements over greedy decoding, Best-of-N, and existing MBR baselines on AlpacaEval and MT-Bench. After self-training, the optimized model achieves superior performance using only greedy decoding—outperforming the original 70B model’s MBR results.
Large language models (LLMs) frequently suffer from hallucinations and unreliable outputs. Method: This paper proposes the first information-theoretic re-ranking framework for LLM generation, rigorously modeling multi-candidate generation as redundant message transmission over parallel, dependent noisy channels—establishing a formal theoretical analogy between generative re-ranking and noisy-channel coding. Under realistic constraints—including imperfect rerankers and channel statistics with inter-output dependencies—we derive sufficient conditions for asymptotically zero decoding error and distill universal re-ranking principles. Contribution/Results: Integrating Mallows and Zipf–Mandelbrot ranking models with statistical channel analysis, we validate the framework on DeepSeek-Coder-7B and TowerInstruct-13B across code generation and medical translation tasks, achieving significant improvements in correctness. Empirical results confirm both the theoretical soundness and robustness of the approach under practical deployment conditions.
This work addresses the performance limitations of non-autoregressive speech recognition, which stem from the absence of contextual modeling over previously generated tokens. The study introduces, for the first time, the Minimum Bayes Risk (MBR) criterion into this paradigm, proposing an efficient decoding framework based on parallel sampling and expected utility maximization. By leveraging the inherent parallelism of non-autoregressive models, the method generates multiple output hypotheses in a single forward pass to estimate risk and optimize the final prediction. Evaluated on LibriSpeech, Switchboard, AMI, and web-based lecture datasets, the approach consistently outperforms existing non-autoregressive methods while achieving faster decoding speeds than autoregressive counterparts.
This study addresses the miscalibration of softmax confidence during decoding in masked diffusion language models by proposing BayesER, a framework that leverages Bayesian predictive entropy to guide token commitment. Specifically, it constructs a lightweight posterior via training-free LoRA adapters and Laplace approximation, thereby optimizing position ranking and selection within denoising steps to enable uncertainty-aware decoding with cross-dataset transferability. Experimental results demonstrate that BayesER significantly reduces sequence-level calibration error while maintaining or improving accuracy on tasks such as code generation. Overall, this work establishes a reliable Bayesian inference decoding paradigm for diffusion language models.
This study addresses how the likelihood ranking of candidate noise sequences in list decoding influences the posterior probability that the first-ranked entry is correctly decoded. Leveraging Soft-Output Guessing Random Additive Noise Decoding (SOGRAND) and a random codebook model, we construct a finite-blocklength conditional probability framework to analyze how subsequent list entries affect the reliability of the first entry. We prove that the asymptotic reliability of the first entry depends solely on the first, second, and last rankings. Furthermore, we derive the decision functions and exponential rates governing its convergence to either one or zero, revealing a “rescue” effect whereby a late-emerging second-ranked entry can substantially enhance confidence in the first entry.
为解决最小贝叶斯风险解码中的度量过拟合问题,本文通过奇异值分解法对效用矩阵进行降噪处理,从而提高文本生成质量。
该研究提出Quit策略,通过提前终止候选生成和重排序过程来减少神经机器翻译中的计算瓶颈,提高效率同时保持翻译质量。