🤖 AI Summary
Existing process supervision data lacks controllability in error location, type, and trajectory consistency, hindering effective training of process reward models. This work proposes a controllable and verifiable synthetic framework that injects template-aware errors into intermediate steps of correct symbolic reasoning chains, recomputes subsequent steps via state propagation, and ensures the first error cannot be logically derived from prior context, thereby generating natural language process pairs that are internally consistent and precisely localize the first mistake. For the first time, this approach enables fine-grained control over error position, type, and trajectory coherence while incorporating a verification mechanism to guarantee supervision reliability. Experiments demonstrate that the synthesized data significantly improves Best-of-8 reranking performance on logical reasoning tasks and generalizes effectively to mathematical reasoning, further revealing that pinpointing the first error is substantially more challenging than classifying entire step sequences—highlighting the necessity of fine-grained process supervision.
📝 Abstract
Process reward models (PRMs) rely on high-quality process supervision data, yet existing construction methods often provide limited control over error location, error type, and trajectory consistency. We propose a controllable and verifiable framework for synthesizing process supervision data for PRMs. Our framework first constructs a correct symbolic reasoning chain, injects a template-aware error into an intermediate step, recomputes subsequent steps under the corrupted state, and verifies that the injected step is not derivable from its prefix. The resulting paired trajectories are prefix-invalid at the first error while remaining trajectory-consistent after symbolic recomputation, and are translated into aligned natural-language processes for PRM training and evaluation. Experiments show that the synthesized data improve Best-of-8 reranking on logical reasoning benchmarks and transfer to mathematical reasoning. Step-level evaluation further shows that first-error localization remains substantially more challenging than overall step classification, highlighting the need for fine-grained and verifiable process supervision.