Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underperformance of Process Reward Model (PRM)-guided inference in discrete diffusion models compared to an independent sampling-and-reranking baseline. We identify two underlying failure mechanisms: weak pruning signals and poor terminal-state discrimination. To overcome these limitations, this work proposes a novel guidance paradigm that preserves early correct solutions while deferring final selection to a terminal-state verifier. Furthermore, leveraging Sequential Monte Carlo sampling, we release a labeled corpus of denoising states. By quantifying the significant performance gap between PRMs and the baseline, this research provides critical empirical insights for optimizing inference guidance in discrete diffusion language models.
📝 Abstract
Discrete diffusion language models (dLLMs) expose a denoised solution at every step, which makes process reward model (PRM) guidance look like a way to spend compute at test time. We show that once denoising, PRM scoring, and outcome reward model (ORM) scoring are charged in the same budget of forward passes, its deterministic form loses to a much simpler baseline. Our PRMs score intermediate denoising states and are trained on the correctness of the final answer. On Dream-v0-Instruct-7B with 8 candidates per GSM8K problem, keeping the candidate with the highest PRM score at every scoring step reaches 65.18%, while independent sampling plus an ORM reranker trained for the task reaches 75.13%. The gap grows to 12.69 percentage points (pp) with 32 candidates, and is 9.85 pp on MATH and 12.16 pp on MBPP. We trace it to two separable failures. First, guidance prunes on a weak signal: on GSM8K, PRM ROC-AUC falls from 0.77 to 0.54 as the mask ratio rises, a decay that persists when states are relabeled with fresh rollouts, and pruning lowers the best accuracy reachable from the candidate pool from 81.05% for independent samples to 67.30%. Second, on GSM8K and MATH, the PRM is a poor final judge: a sequential Monte Carlo sampler at the same budget restores that ceiling to 77.89%, yet selecting with the PRM gives 65.48%, on par with deterministic guidance, while a PRM retrained on final states matches the ORM on identical candidates. MBPP separates the two: there the PRM reaches 65.47% when reranking finished programs, on par with the ORM, but 50.88% when it guides denoising. The results point to two targets for dLLM guidance: keep correct partial solutions alive through early denoising, and leave the final choice to a verifier trained on final states. We release the corpus of denoising states with outcome labels and evaluation toolkit for reproducible comparisons at matched compute.
Problem

Research questions and friction points this paper is trying to address.

Discrete diffusion language models
Process reward model
Deterministic guidance
Reasoning
Test-time compute
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete Diffusion Language Models
Process Reward Model
Deterministic Guidance
Outcome Reward Model
Sequential Monte Carlo
🔎 Similar Papers
No similar papers found.
Y
Yan Zhan
Peking University
Shaobo Liu
Shaobo Liu
Unknown affiliation
Z
Zhijun Gao
Peking University