π€ AI Summary
This study addresses the issues of sample homogenization in multi-sampling for diffusion language models and the degradation of single-sample accuracy caused by high-temperature sampling. To this end, we propose the self-repelling sampler, which introduces a peer-commitment-based counting penalty mechanism that penalizes duplicate tokens during the denoising phase, enabling diverse reasoning paths without additional training or backpropagation. By integrating batched forward passes with a sequential commitment strategy, our method achieves deterministic diverse sampling at zero temperature. Experimental results demonstrate that ten-way majority voting on GSM8K reaches 80.38% accuracy, significantly outperforming greedy decoding, while evaluations on benchmarks such as MATH further validate the effective improvement in correct coverage rate.
π Abstract
Sampling several responses and voting over their answers can improve a language model's accuracy, but repeated answers limit the benefit of additional samples. Raising temperature increases diversity at a potential cost to per-sample accuracy. We introduce Self-Repulsion (SR), a sampler for masked diffusion language models that uses peer commitments to diversify the pool. At each penalized denoising step, each path lowers a token's logit according to how many peers have committed that token at the same position. Paths share a batched forward pass and then commit in sequence, so later paths observe choices made earlier in the same step. This coupling requires no training or additional forward or backward pass and can produce distinct paths even at temperature zero. When all paths commit a position together from identical logits, the update exactly maximizes total logit minus a convex duplication cost. On LLaDA-8B-Instruct with ten paths and 128 denoising steps, deterministic SR reaches 80.38% plurality accuracy on GSM8K, compared with 70.17% for the unpenalized greedy decoder. At temperature 0.6 and matched model-evaluation budgets, the count penalty improves over self-consistency by 2.06 percentage points in blocks of 32 and 14.50 under pure diffusion. Experiments on GSM8K, MATH and TruthfulQA show that voting gains arise mainly from higher coverage of correct answers, with gains that vary by benchmark and decoding regime.