Score
Designs methods and tools that take recorded execution traces and derive a context-free grammar which compactly models the traces’ control-flow and repeated I/O patterns; the resulting CFG preserves access sequences to enable replay and generation of observed behaviors.
This work addresses the problem of automatically synthesizing complete, functionally correct programs that reproduce a target behavior from partial execution traces containing only side-effecting function calls—requiring joint modeling of side-effecting functions, pure functions, and control flow, with positive-only examples and no negative ones. To solve it, we propose the first hierarchical framework integrating syntax-guided program synthesis (SyGuS) with optimization-based rewriting: (1) side-effect-aware modeling and trace abstraction ensure semantic-preserving generalization; (2) a lightweight cost metric mitigates overgeneralization; and (3) rewriting rules combined with syntactic constraints guarantee correctness-by-construction. Evaluated on real-world API and system-call benchmarks, our approach efficiently generates functionally complete programs, significantly advancing synthesis capability for scenarios involving mixed side effects, pure functions, and complex control flow.
Real-world microservice call-graph traces are scarce, hindering system behavior analysis and resource management. Method: This paper pioneers the use of large language models (LLMs) for synthetic workload trace generation, proposing a two-stage paradigm: (1) recursive sequential generation to explicitly model the hierarchical structure and dynamic evolution of call graphs; and (2) instruction-tuning guided by implicit structural constraints to ensure topological consistency and coverage of rare scenarios. The approach integrates graph-structured modeling, LLM fine-tuning, and a synthetic-data-driven evaluation framework. Contribution/Results: Experiments demonstrate that the generated traces significantly outperform state-of-the-art methods in diversity, topological fidelity, and downstream utility—including feature prediction and missing-trace completion. The synthetic traces effectively substitute real traces for microservice resource optimization and system-level analysis.
This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.
Traditional loop invariant generation tools exhibit limited precision and applicability on real-world programs where complex data structures intertwine with intricate control flow. Method: This paper proposes ACInv, the first static-analysis-driven, LLM-augmented framework for invariant synthesis. It extracts loop semantic features to construct structured prompts for LLM-based candidate invariant generation, and introduces an LLM-powered semantic evaluator that dynamically refines candidates via strengthening, weakening, or rejection. Contribution/Results: ACInv is the first approach to support template-level invariant generation for user-defined data structures. Evaluated on benchmarks containing complex data structures, ACInv achieves a 21% higher overall solving rate than AutoSpec, while matching its performance on purely numeric programs. Moreover, the generated invariants are reusable and significantly improve practicality for industrial-scale program verification.
Automated verification of interactive console I/O programs in Haskell education remains challenging due to the dynamic, history-dependent nature of student implementations. Method: We propose a lightweight, formal behavioral specification language that uniquely integrates global state and execution history, expressed via regex-like syntax; its trace-based semantics enable probabilistic testing and scalable verification through *sampleable validity*. Contribution/Results: Our system automatically validates student submissions against behavioral specifications and supports pedagogical closed-loop applications—including real-time feedback generation, example solution synthesis, and exercise randomization. Empirical evaluation demonstrates substantial improvements in test coverage and pedagogical adaptability while preserving formal rigor. To our knowledge, this is the first framework for verifying interactive behaviors in functional programming education that simultaneously achieves theoretical soundness and practical deployability.
本文提出记录和利用执行轨迹来持久化和分析操作的内部行为,解决VDM-SL规范中操作内部行为观察问题。
This work addresses the challenge of efficiently testing black-box systems with side effects by proposing a test generation approach that integrates under-approximate typing with effect systems. The method employs symbolic traces to capture data and control dependencies of side-effecting operations, preserving essential constraints to guide test case synthesis. Precise coverage is achieved through an integration of property-based testing and model checking. The implemented tool, Clouseau, demonstrates substantial improvements over default strategies in frameworks such as QCheck and P, achieving test effectiveness comparable to state-of-the-art hand-crafted test suites. These results validate both the efficacy and practicality of the proposed methodology.
This work addresses the challenge of efficiently storing and querying long-sequence trajectory data—comprising nested branches, state transitions, and textual payloads—under strict token budget constraints. The paper proposes the Budgeted Dynamic Trajectory Structure (BDTS), a novel framework that unifies, for the first time, mechanisms including state-filtered reachability, cursor-based pagination, soft-bounded recent logs, reference-counted observation keys, incremental overwrites, bounded caching, and summary-suffix compression. BDTS maintains a rooted trajectory graph and an append-only history while rigorously preserving formal invariants within hard token limits. Experimental results demonstrate that BDTS achieves sub-3-millisecond construction and query latency on trajectories with tens of thousands of nodes, compressing raw inputs from 350K–2.71M tokens down to 1,048–4,120 tokens—reducing model input size by nearly 90%.
This work addresses the challenge that large language models struggle to effectively detect property violations in program verification, particularly exhibiting significant performance degradation on long programs. The authors propose a novel approach that leverages error traces generated by the symbolic execution engine Soteria as continued pretraining data for the Qwen3-8B model, combined with chain-of-thought reasoning at inference time to enhance semantic understanding of programs. Using only approximately 3,000 such traces, this method improves violation detection accuracy by over 17 percentage points, enabling the 8B-parameter model to surpass a 32B-parameter counterpart without chain-of-thought reasoning. The approach demonstrates balanced performance across five SV-COMP property categories and generalizes to unseen property types, validating the superadditive effect arising from the synergy among error trace semantics, formatting, and chain-of-thought reasoning.
Traditional control flow graph (CFG) generation methods rely on syntactically complete code and language-specific tooling, making them ill-suited for handling erroneous or incomplete code fragments and lacking unified support across multiple programming languages. This work proposes the first approach to leverage lightweight large language models—such as CodeLlama and QwenCoder—for robust CFG construction from such low-quality inputs. By employing instruction fine-tuning, a unified serialization format, and an automatically curated, error-augmented dataset derived from LeetCode, the method achieves strong parsing performance even on malformed or partial code. Notably, it not only excels on languages seen during training but also demonstrates cross-lingual generalization to unseen programming languages, establishing a new paradigm for static analysis of multilingual, low-quality code.