Score
Designs and builds methods that parse and ground high-level or natural-language intents into explicit, step-level plans or executable pipelines, producing alternative decompositions and a compiled plan space of candidate step sequences. Analyzes and validates the resulting plans and pipelines for feasibility, required resources and dependencies, and correctness of intent-to-step mappings.
Large language models (LLMs) excel at natural language understanding but lack robustness in long-horizon automated planning tasks requiring structured, symbolic reasoning. This paper pioneers the conceptualization of LLMs as “planning modelers,” systematically surveying their role in generating formal symbolic planning models—such as PDDL—from natural language specifications, thereby establishing a unified analytical framework bridging NLP and automated planning. Methodologically, we integrate prompt engineering, few-shot learning, programmatic validation, and neuro-symbolic interfaces, while incorporating formal verification and off-the-shelf planners. Through a taxonomy-driven evaluation of over 100 works, we identify five fundamental bottlenecks—including semantic gaps and insufficient action generalization—and propose scalable solutions such as a verifiable modeling pipeline. Our approach substantially enhances the interpretability, reliability, and automation capability of symbolic planners.
This work addresses the prevalent issue in large language models (LLMs) of introducing control-flow, type, or I/O errors during code translation due to neglect of program intent. To mitigate this, the paper proposes the first systematic use of a language-agnostic, structured intermediate specification that preserves semantic fidelity through an intermediate representation, structured generation, and automated test-based validation. Evaluated on the Avatar and CodeNet datasets with five state-of-the-art LLMs, the approach significantly improves translation accuracy, raising the micro-averaged accuracy from 67.7% to 78.5%. It completely eliminates lexical errors and substantially reduces errors related to structure, declarations, and runtime dependencies.
Large language models (LLMs) exhibit significantly degraded function completion performance when source code lacks explicit documentation (e.g., docstrings), hindering accurate intent understanding. Method: This paper proposes a three-stage, intention-driven approach: (1) context-aware intention encoding via reasoning over code context; (2) interactive intention refinement to enhance precision; and (3) target function generation conditioned on the clarified intention. Contribution/Results: The method innovatively shifts intention inference to the earliest stage and constructs a high-quality, 40K-instance dataset featuring intermediate reasoning traces. It introduces an optional interactive alignment mechanism and integrates chain-of-thought prompting, context signal extraction/synthesis, and multi-stage conditional generation. Evaluated on DevEval and ComplexCodeEval, our approach yields average improvements exceeding 20% across mainstream LLMs; the interactive component delivers further substantial gains.
Large language models often fail in multi-step structured workflows due to error propagation, particularly when state manipulation and validation are required. This work proposes a novel compiler architecture that decouples planning from execution, introducing for the first time a deterministic compilation paradigm. By leveraging a typed node registry, static graph validation, and structured JSON plan generation, the approach reliably translates LLM outputs into executable Python code, integrated with SQLite for persistent state management. Evaluated on a benchmark of 300 tasks, the method achieves 278 first-pass successes—substantially outperforming GPT-4.1 and Claude Sonnet baselines—while reducing cost to \$0.356 per task and maintaining competitive end-to-end latency, thereby clearly delineating the current capability boundary of such approaches.
This work addresses the “intent gap” between user expectations and program behavior in AI-generated code by proposing intent formalization as a central pathway to transform informal requirements into verifiable formal specifications. We systematically identify intent formalization as a critical challenge for reliable coding in the AI era and introduce an end-to-end verifiable coding framework that integrates formal methods, test-driven development, AI-generated postconditions, domain-specific languages, and human-AI collaboration. The framework supports a spectrum of approaches ranging from lightweight testing to fully automated correctness-preserving synthesis. Preliminary experiments demonstrate that interactive, test-driven formalization effectively enhances program correctness, that AI-generated postconditions can uncover real-world bugs, and that provably correct code can be automatically synthesized from informal specifications.
To address the generalization failure of large language models (LLMs) in generating generalized plans for PDDL domains—often caused by erroneous initial strategies—this paper proposes a novel method integrating pseudocode-based strategy modeling, automated debugging, and reflective multi-program mutation selection. Our approach features three key contributions: (1) explicit strategy representation via executable pseudocode to proactively rectify logical errors; (2) a natural language inference–driven reflective prompting mechanism to enhance strategy consistency; and (3) systematic program mutation generation coupled with formal verification to select the optimal generalized plan. Evaluated on 17 standard benchmark domains, our method achieves significant improvements in planning accuracy and robustness, with zero performance degradation. Notably, for 12 domains, the best generated program solves all automatically instantiated tasks within the domain, demonstrating full-domain generalization capability.
Automatically constructing high-quality, reusable skills from heterogeneous, fragmented interaction traces—often missing critical security behaviors—is highly challenging. This work proposes the W2S framework, which introduces a novel intermediate representation called RWSA to decouple skills into workflow structure, execution semantics, and runtime attachments, thereby enabling task decomposition, control-flow modeling, verification, rollback, and state management. W2S achieves efficient skill construction through trajectory segmentation, local skill draft generation, structural alignment, branch fusion, redundancy compression, and confidence-aware retention. Experimental evaluation across 70 skills demonstrates that W2S improves behavioral replay consistency by 10.5% compared to baseline approaches based on summarization and prompting.
This work addresses the challenge of ensuring program correctness in natural language-to-code generation, which is often hindered by the absence of high-quality formal specifications. The authors propose VeriSpecGen, a framework that decomposes natural language requirements into atomic clauses through a traceable refinement mechanism, generates requirement-driven tests with explicit traceability mappings, and synthesizes formal specifications aligned with user intent by localizing and repairing faulty clauses upon verification failure. Integrating large language models (e.g., Claude Opus 4.5) with the Lean proof assistant, the approach leverages refinement trajectories to generate 343K training samples, substantially enhancing model generalization and reasoning capabilities. Evaluated on the VERINA SpecGen benchmark, VeriSpecGen achieves an accuracy of 86.6%, outperforming the best baseline by up to 31.8 percentage points and demonstrating a relative improvement of 62–106% in specification synthesis performance.
This work addresses the unreliability of large language models (LLMs) in executing structured workflows specified through natural language. To overcome this limitation, the authors propose RunAgent, a multi-agent platform that integrates the expressiveness of natural language with the determinism of programmatic execution through a novel agent language. RunAgent introduces a constraint-guided stepwise execution mechanism augmented with explicit control structures, dynamic selection of reasoning strategies, and context filtering to ensure robust task execution. The framework automatically derives verifiable constraints and supports a hybrid paradigm combining tool invocation, Python code generation, and execution. Evaluated on the NaturalPlan and SciBench benchmarks, RunAgent substantially outperforms both baseline LLMs and the current state-of-the-art PlanGEN method.
In multi-step LLM workflows, accumulated context induces hallucinations, intermediate output confusion, and loss of task constraints. To mitigate context contamination, we propose NormCode—a semi-formal language enforcing strict data isolation and stepwise reasoning: each step accepts only explicit inputs, decoupling semantic reasoning (non-deterministic LLM inference) from syntactic operations (deterministic data manipulation). NormCode introduces a novel tri-format isomorphism (.ncds/.ncd/.ncn) enabling progressive formalization—from human-authored drafts to machine-executable code to human-verified artifacts. It integrates dependency-aware scheduling, SQLite-based checkpointing, and native loop management, ensuring end-to-end auditability. Evaluated on arbitrary-length X-ary addition, NormCode achieves 100% accuracy; it further successfully self-hosts a five-stage compiler. This work establishes a traceable, hallucination-resistant AI planning infrastructure for high-stakes domains including law, healthcare, and finance.
This work addresses the frequent lack of semantic validity in code generated by large language models (LLMs) for software engineering tasks. To this end, it introduces a projection decoding framework that, for the first time, treats graph-based representations as first-class citizens alongside textual sequences during generation. The framework incrementally constructs partial graph structures in parallel with token prediction, directly embedding domain-specific semantics into the decoding process. This integration enables uncertainty modeling, incremental semantic validation, and provable correctness guarantees, thereby establishing a verifiable foundation for LLM-driven software engineering automation. Experimental results demonstrate that the proposed approach significantly improves the semantic validity of generated artifacts in program synthesis tasks.