Score
Designs formulations and models that treat chains of reasoning as bags of latent instances (multiple-instance learning), representing each reasoning step or subproof as a potential instance whose correctness is unobserved. Builds training objectives, learnable pooling/aggregation functions, and inference procedures that assign credit to steps and enable learning from only final-answer supervision by modeling step correctness as latent variables.
Existing chain-of-thought (CoT) methods exhibit limited generalization for logic-intensive domain-specific reasoning tasks—such as legal reasoning—that require deep, structured domain knowledge; meanwhile, Monte Carlo Tree Search (MCTS) lacks adaptability to professional reasoning contexts. Method: This paper pioneers the integration of MCTS into domain-specific reasoning via a step-level supervision framework: (i) a knowledge-guided MCTS search space constraint mechanism that explicitly aligns domain rules (e.g., statutory provisions, legal elements) with reasoning steps; and (ii) a learnable reflection-path preference model that enhances self-monitoring and correction of erroneous reasoning trajectories. Contribution/Results: Our approach achieves significant improvements over state-of-the-art baselines across multiple legal reasoning benchmarks. Further analysis reveals a strong positive correlation between fine-grained domain knowledge representation—such as statutory hierarchy and precise legal-element decomposition—and reasoning accuracy, establishing a novel paradigm for expert-level AI reasoning.
This work addresses the computational challenge of learning from multiple annotators who provide correct yet stylistically diverse chain-of-thought (CoT) rationales. The study is the first to formally characterize the computational complexity of this setting and introduces an efficient active learning algorithm that overcomes the limitations of passive learning. The proposed method requires only a fixed number of CoT examples per annotator, O(log(1/ε) log log(1/ε)) annotators, and Õ(1/ε) supervisory labels on final answers to achieve ε-accurate learning. Notably, its annotation efficiency is independent of the target accuracy ε, substantially enhancing the scalability of learning under heterogeneous CoT supervision.
This work addresses the challenge of overfitting and poor generalization in multiple instance learning (MIL) under label-scarce conditions by proposing a context-based, fine-tuning-free approach. The method leverages a Perceiver architecture pretrained on diverse synthetic bag-structured datasets, integrating complementary inductive biases from varied generation strategies. This enables the model to perform accurate classification on new MIL tasks through a single forward pass with only a few labeled bags, without requiring any gradient-based adaptation. Evaluated across twelve established MIL benchmarks, the proposed approach consistently outperforms supervised baselines that rely on task-specific training, demonstrating substantially improved generalization and practical utility in few-shot MIL scenarios.
This study addresses the limitation of reasoning models in solving complex problems, particularly their difficulty in effectively extracting reusable solution strategies from known solutions. To overcome this, it proposes a hindsight hierarchical framework coupled with a self-improving loop mechanism that jointly trains three tasks: answer prediction, reverse engineering, and thought-based solving. By reverse-engineering known solutions to generate additional supervisory signals, the approach leverages hindsight to enable closed-loop self-enhancement, which is subsequently applied to the Lean interactive theorem prover. The primary contribution lies in providing formal specifications and concrete instantiations of the core methodology, establishing a novel paradigm for the autonomous iterative optimization of reasoning models, while empirical evaluation remains to be conducted in future work.
This study challenges the prevailing view that supervised fine-tuning (SFT) merely memorizes without generalizing, systematically investigating the cross-domain generalization capabilities of long chain-of-thought (CoT) SFT in reasoning tasks. Through cross-domain evaluation, solution trajectory analysis, comparisons across base models of varying scales, and training dynamics tracking, the work reveals that SFT generalization is conditional: stronger models internalize transferable procedural reasoning strategies from simple tasks, exhibiting a non-monotonic “dip-and-recover” training dynamic, whereas weaker models only mimic superficial patterns. The study further demonstrates that high-quality long CoT data substantially enhances out-of-domain performance, yet this gain in reasoning capability may come at the cost of reduced safety, highlighting a potential trade-off inherent in generalization.
Existing continuous chain-of-thought (Continuous CoT) methods rely on slow autoregressive generation and suffer significant performance degradation on tasks requiring long reasoning trajectories. This work proposes C-MTP, a novel approach that, for the first time, directly supervises hidden states using the mean of corresponding chain-of-thought embeddings, thereby employing embedding averages as supervision signals to simplify training and eliminate the need for autoregressive decoding. The method outperforms existing direct supervision approaches on short reasoning tasks and matches the performance of indirect supervision methods. However, on long reasoning trajectories spanning hundreds of tokens, all current methods—including C-MTP—experience a performance drop of approximately 65%, revealing a fundamental limitation of contemporary Continuous CoT frameworks in long-horizon reasoning.
This work addresses the challenge that large language models struggle to learn from sparse correct answers in complex reasoning tasks, as conventional outcome-based rewards fail to leverage partial progress from failed attempts. To overcome this, the authors propose the SCRL framework, which constructs a curriculum of verifiable subproblems derived from reference reasoning chains, treats the original problem as the final subproblem, and applies reward normalization at the subproblem level to enable fine-grained credit assignment. SCRL is the first approach to integrate verifiable subproblem curricula with subproblem-level reward normalization, transforming partial reasoning progress into effective learning signals without external scoring and mitigating gradient vanishing in hard problems. Evaluated on seven mathematical reasoning benchmarks, SCRL significantly outperforms strong baselines, achieving an average accuracy gain of 4.1 points on Qwen3-4B-Base and improvements of 3.7 and 4.6 points in pass@1 and pass@64 on AIME24/25 and IMO-Bench, respectively.
This study addresses a critical data degradation issue in reasoning distillation for large language models, where answer-conditioned chain-of-thought generation induces models to favor post-hoc rationalization over genuine forward reasoning. Crucially, this degradation persists even after filtering by posterior correctness. The work systematically uncovers this mechanism for the first time and demonstrates its universality across multiple mainstream large language models through controlled ablation studies, cross-model transfer experiments, prompt ablations, and unsupervised evaluation protocols. Empirical results reveal that training with answer-conditioned chains reduces verifiable reasoning accuracy by up to 27 percentage points on the most challenging competition-level problems, with degradation severity intensifying significantly as problem difficulty increases.
本文针对潜变量推理中的表示崩溃和信息分布不均问题,提出了一种基于原型介导的过程监督方法PMPS,并通过实验验证了其有效性。
Implicit chain-of-thought reasoning relying solely on outcome supervision is prone to semantic drift and gradient vanishing, hindering robust inference. This work addresses these limitations by reframing process supervision through an information-theoretic lens, decoupling it into trajectory and spatial supervision. Rather than enforcing geometric compression, the proposed approach preserves informational fidelity in the reasoning space via mutual information maximization. The authors introduce a “dual collapse” mechanism to elucidate the root causes of failure and develop a dual-dimensional supervision framework that jointly governs trajectory and spatial aspects. Semantic structure is retained through generative reconstruction instead of rigid geometric constraints. Experiments demonstrate a strong correlation between reasoning accuracy and the fidelity of latent trajectory information, establishing an “information–performance binding” principle that offers principled guidance for supervising implicit reasoning systems.