🤖 AI Summary
This study addresses the lack of unified strategies in existing fact-checking systems and the difficulty of optimizing sequential decisions for retrieval and verification. We propose a unified verification strategy framework that formulates fact-checking as a sequential decision-making process. Built upon the Qwen3-8B model, our approach introduces a complementary mechanism integrating SCCS cold-start initialization with VA-RL reinforcement learning. Through supervised fine-tuning, budget-aware tool rewards, and advantage reweighting techniques, the method achieves end-to-end control over reasoning, evidence retrieval, and termination, generating reliable trajectories and optimizing policies without gold labels. Experimental results demonstrate that our approach attains an average accuracy of 70.39% and a macro-F1 score of 63.30% across six benchmarks, significantly outperforming existing methods.
📝 Abstract
Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipelines or separately prompted modules rather than learning them as a unified task-specific policy. We introduce \textbf{OpenFC}, a unified verification-policy training framework that post-trains Qwen3-8B as a compact next-action controller over reasoning, evidence acquisition, and stopping. OpenFC learns this policy in two stages. \textbf{Stepwise-Calibrated Cold Start (SCCS)} uses a strong training-time supervisor to review post-initial reasoning, tool-use, and stopping proposals before execution, producing reliable trajectories for supervised fine-tuning without access to gold verdicts. \textbf{Verification-Aware Reinforcement Learning (VA-RL)} then improves the cold-start policy on unresolved claims through budget-aware tool rewards, label-aware advantage reweighting, and localized response masking. Across six fact-checking benchmarks, OpenFC achieves 70.39\% average accuracy and 63.30\% macro-F1, the highest overall averages among the evaluated methods. Stage-wise ablations further show that SCCS and VA-RL provide complementary gains, supporting the design of the two-stage training framework. These results position OpenFC as a strong and effective framework for open-search fact-checking. We will open-source our code and release the model checkpoints to support reproducibility.