Score
Designs and executes stepwise constructions or proofs that start from a base case and extend via a clearly specified inductive step to produce mathematical or algorithmic objects. Controls and analyzes parameter growth at each step to maintain specified bounds and thereby guarantee existence and required properties of the constructed object.
本文提出STEP-KTODER框架,通过定义模块级函数为步骤并自动生成单元测试来优化代码偏好,解决了代码生成中过程监督不明确的问题。
Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.
This work addresses the challenge in functional programming pedagogy where evaluation traces generated by algebraic steppers are often cluttered with trivial reduction steps that obscure core concepts and hinder learning. To remedy this, the authors propose an algebraic stepper equipped with a scoping-aware filter mechanism that enables users to selectively show or hide reduction steps via lightweight pattern expressions. Inner filters can override outer ones, facilitating focused, hierarchical trace exploration. The approach introduces the first composable, lexically scoped filtering framework for steppers, accompanied by a formal semantics proven consistent with the original language semantics. Notably, it supports stepping through programs containing holes or type errors—a capability not previously available. The system’s meta-theory, including preservation, progress, and simulation theorems, is mechanically verified in Agda and integrated into the Hazel live programming environment. Classroom deployment demonstrates that students intuitively adopt the tool, while instructors effectively construct concise, pedagogically optimized traces.
本文通过改进循环不变式、增加循环变体等方法,优化了VDM操作的证明义务生成过程,确保模型内部一致性。
This paper addresses the lack of a unified metatheoretic characterization for program logics handling multi-branching effects—such as nondeterminism and probabilism. We propose a novel program logic framework centered on algebraic choice structures. Methodologically, we are the first to embed algebraic effects modeling directly into the core of Hoare logic, integrating modal semantics with a relatively complete proof system that supports general loops and uniform reasoning across effect types (e.g., nondeterministic and probabilistic). Our main contributions are: (1) the first relatively complete proof system for Hoare logic strictly extending it to cover multiple branching effects; (2) a unified metatheoretic account of multi-result programs; and (3) formal support for cross-model reuse of proof fragments—enabling verification transfer between distinct semantic models (e.g., relational, probabilistic, or game-based interpretations).
This study addresses the theoretical drift caused by model modifications and the challenges of automated construction in the formalization of stochastic optimization algorithms. We propose a fully automated formalization framework driven by large language model (LLM) agents. This framework employs proof obligations to guide the automatic construction of Lean models and supporting theories, introduces signature contracts alongside independent auditing mechanisms to prevent assumption weakening, and establishes a reusable verification library, SOptLib, to enable cumulative verification cycles. Experimental results demonstrate that the system achieves an average score of 6.3 out of 7 across 15 tasks, generates 490,000 lines of `sorry`-free code, and identifies 28 formula errors and proof gaps in published literature, thereby realizing highly reliable automated formalization of research-grade algorithms.
This study addresses the significant challenge of verifying termination in real-world C/C++ programs, where loop interactions and nondeterministic inputs complicate analysis. The authors propose a lightweight, tool-agnostic, source-level preprocessing approach that isolates loop obligations via loop slicing and enhances termination analysis by generating input-driven concrete variants tailored to specific scenarios. An empirical evaluation integrating six termination analyzers on 117 real programs demonstrates that slicing conservatively achieves structural isolation, while concretization improves detectability in targeted scenarios at the cost of reduced semantic coverage. Crucially, the combined effect of these techniques is non-additive, indicating that preprocessing should complement—rather than replace—analysis of the original program. The work further reveals substantial variation in how different analyzers respond to preprocessing, offering practical guidance for adaptive usage by developers.
本文通过元编程解决归纳类型扩展性问题,提出组合算法及Lean证明助手的语法扩展,实现类型和函数定义的模块化复用与扩展。
该研究通过消除几何学框架探讨局部最优对象能否由共享部署规则实现,分析信息、架构等因素对缺陷修复的影响。
Symbolic execution often struggles to adequately explore program paths due to resource constraints. To address this limitation, this work proposes Agolic, a novel system that introduces agent-based planning into the symbolic execution workflow. Without altering the underlying exploration logic, Agolic dynamically configures multiple rounds of bounded symbolic execution through cross-round, high-level reasoning. The approach synergistically integrates large language models, source code analysis, coverage replay, and goal-directed strategies to substantially enhance path coverage. Experimental results demonstrate that Agolic achieves, on average, more than three times the branch coverage of continuous symbolic execution across several C/C++ programs and uncovers previously unexplored branches in six out of seven benchmarks—branches missed by a combination of fuzzing and compiler-assisted concrete execution.