Score
Designs and implements algorithms, evolutionary loops, and frameworks that generate, mutate, recombine, and select symbolic communication protocols or compact languages. Builds scoring and evaluation pipelines to measure correctness, cost, and other performance metrics and iteratively optimize candidate protocols to discover compact, high‑accuracy symbolic protocols (language symbolism frameworks and evolution loops).
This paper addresses the stagnation in performance, code redundancy, and stylistic rigidity commonly observed when large language models (LLMs) iteratively generate algorithms within evolutionary computation frameworks. To model the dynamic iterative trajectory of LLM-generated code, we propose the **Code Evolution Graph**—the first formal graph-based representation of such evolution. Leveraging static analysis, graph representation learning, and cross-model behavioral comparison across three benchmark task categories, we empirically reveal: (i) iterative generation often increases code complexity while degrading performance; (ii) generated code exhibits significant heterogeneity and stylistic isolation across models; and (iii) repeated prompting induces redundant overcomplication. Building on these insights, we introduce a **multi-LLM co-evolution paradigm**, which demonstrably mitigates degeneration and improves solution quality. Our approach provides an interpretable, controllable pathway for LLM-driven automated algorithm design.
Traditional chain-of-thought (CoT) prompting generates verbose reasoning traces in complex tasks, struggling to balance efficiency and accuracy. This work proposes the Communicative Language Symbolism Routing (CLSR) framework, which treats linguistic symbol systems as reusable symbolic protocols. During inference, multiple large language model agents autonomously invent, evolve, and share compact symbolic languages, dynamically composed via a latent-variable-free router to trade off accuracy against computational cost. Integrating multi-agent collaboration, evolutionary optimization, and information-theoretic analysis, CLSR reduces reasoning token consumption by 3–6× across multiple benchmarks while matching standard CoT accuracy. The approach further establishes a theoretical connection between emergent symbolic protocols and program execution pipelines.
Symbolic execution frequently terminates prematurely due to unanalyzable external functions—such as native methods or third-party library calls—for which no source-level implementation is available. To address this, we propose the first genetic programming–based approach for automatically synthesizing symbolic stubs: it collects input-output samples via random testing, then evolves algebraic expressions to approximate the target function’s behavior—without requiring manual modeling or auxiliary contextual information. Our method tightly integrates symbolic execution, randomized testing, and SMT-based constraint solving to uncover language-specific semantics and identify boundary conditions. Experimental evaluation demonstrates that 55% of the tested functions achieve over 90% behavioral approximation accuracy; moreover, our synthesized stubs successfully recover critical execution paths inaccessible to conventional stubbing techniques, yielding substantial improvements in both path coverage and testing depth.
Addressing core challenges in legacy system modernization—including implicit contract identification, performance degradation, and integration-evolution misalignment—this paper proposes a self-evolving software system framework based on typed directed graphs. The framework uniformly models source code, build scripts, documentation, and issue tickets as evolvable graph structures. It innovatively integrates a lightweight domain-specific language model–driven graph mutation mechanism with a multi-objective fitness selection strategy to enable cross-artifact co-evolution. Evaluated on three benchmarks, the system automatically repairs 83% of security vulnerabilities, achieves 93% functional equivalence in COBOL-to-Java translation, reduces documentation update latency to ≤2 minutes, and shortens feature delivery cycles by 7×. This work delivers the first verifiable, end-to-end autonomous evolution framework for Software 3.0.
This work addresses the limitations of current large language models in generating Verilog code that, while functionally correct, is often unsynthesizable and lacks timing awareness, thus failing to meet practical RTL design requirements. To overcome this, the authors propose a feedback-driven iterative refinement framework that employs a multi-round generate–evaluate–promote mechanism. This framework integrates multidimensional feedback from functional simulation, Yosys synthesis, an ABC-based timing proxy, and GEMM-specific metrics to progressively enhance code quality. Novel strategies—including versioned refinement, cross-session skill evolution, and verification-gated skill updates—enable modular skill retrieval and history-aware decisions for skill creation, improvement, or skipping. Evaluated on VerilogEval and mixed-precision GEMM tasks, the approach significantly improves functional correctness, version promotion stability, and downstream hardware performance, achieving state-of-the-art pass rates on GEMM preservation sets and superior synthesis scores.
This study investigates whether large language models (LLMs) genuinely explore novel program structures or fall into repetitive cycles when iteratively mutating code in the absence of selection pressure. By constructing LLM-driven mutation chains without selective constraints and analyzing program structure, detecting loops, and comparing against classical genetic programming’s subtree mutation as a baseline, the work reveals a strong structural convergence bias in LLM-based mutation: in 87% of mutation chains, over 93% of generated programs reuse existing structures, with variation largely confined to terminal symbol substitutions. Short cycles and self-loops are prevalent across diverse prompt designs, model families, and random replication conditions, demonstrating that LLMs intrinsically gravitate toward limited attractor regions—a behavior markedly distinct from traditional genetic programming.
This study addresses the limitations of traditional handcrafted approaches to link prediction in complex networks, which often suffer from suboptimal performance and poor generalization. To overcome these challenges, the authors propose a novel code evolution framework that integrates large language models with a genetic algorithm to automatically search for and optimize the program structure of link prediction algorithms. The method innovatively explores adaptive combinations of node and link features. Extensive experiments on 580 real-world networks demonstrate that the proposed approach achieves an average AUC of 0.915, substantially outperforming existing hand-designed methods (AUC = 0.783). Furthermore, it exhibits high computational efficiency and scales effectively to networks with millions of links.
This work addresses the challenge of consistently generating effective and adaptive executable trading strategies in noisy, non-stationary, and highly discontinuous algorithmic trading environments. The authors propose a two-level evolutionary framework: at the inner level, a large language model (LLM) serves as a semantic mutation operator to iteratively generate and refine Python-based trading strategies; at the outer level, a meta-evolutionary mechanism automatically optimizes prompting instructions, autonomously discovering program synthesis heuristics that outperform human-designed ones. This approach represents the first integration of LLMs into the strategy evolution process for algorithmic trading, combining evolutionary algorithms with meta-learning. Rigorous backtesting demonstrates that the system adaptively responds to market regimes, dynamically switches trading logic, significantly reduces zero-trade failures, and consistently outperforms baseline strategies guided by manually crafted prompts.
Symbolic execution often struggles to adequately explore program paths due to resource constraints. To address this limitation, this work proposes Agolic, a novel system that introduces agent-based planning into the symbolic execution workflow. Without altering the underlying exploration logic, Agolic dynamically configures multiple rounds of bounded symbolic execution through cross-round, high-level reasoning. The approach synergistically integrates large language models, source code analysis, coverage replay, and goal-directed strategies to substantially enhance path coverage. Experimental results demonstrate that Agolic achieves, on average, more than three times the branch coverage of continuous symbolic execution across several C/C++ programs and uncovers previously unexplored branches in six out of seven benchmarks—branches missed by a combination of fuzzing and compiler-assisted concrete execution.
This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.