LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究针对扩散语言模型中的稳定但错误锁定问题,提出LOCKR方法,通过隐藏状态轨迹指导规划器进行检测和修复,提高模型推理准确性。
📝 Abstract
Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level decoding signals such as confidence, entropy, margin, and answer stability are insufficient to reliably distinguish correct from erroneous lock-in. We formulate selective reasoning repair as a lightweight test-time planning problem and propose LOCKR, a hidden-state trajectory-guided planner that decides when to allocate additional computation, expands a structured set of targeted repair branches, and selects the most promising continuation using trajectory-aware verification. Across two diffusion language models and three mathematical reasoning benchmarks, hidden-state trajectories consistently outperform surface signals and single hidden snapshots for both wrong-lock-in detection and repair selection. On natural evaluation distributions, LOCKR yields absolute accuracy gains of 2.21--5.37 percentage points across all five evaluated settings, with repair rates ranging from 22% to 41%. These results establish hidden diffusion trajectories as actionable signals for selective test-time reasoning repair.
Problem

Research questions and friction points this paper is trying to address.

diffusion language models
stable-but-wrong lock-in
reasoning failure
denoising
surface-level decoding signals
Innovation

Methods, ideas, or system contributions that make the work stand out.

hidden-state trajectory
test-time planning
selective reasoning repair
diffusion language models
stable-but-wrong lock-in
🔎 Similar Papers
No similar papers found.