Score
Designs and implements algorithms that extract semantically structured gradient signals from agent execution trajectories and use those signals to compute policy updates or to distill behavior into updated proposal ranking and selection mechanisms. Builds methods for continually or offline converting past outcomes into semantic-gradient updates that sharpen agent behavior and improve policy optimization and proposal selection over time.
This paper addresses the limitations of large language model (LLM)-based agents in long-horizon planning, dynamic interaction, and complex decision-making within intricate environments. Methodologically, it introduces the first systematic optimization survey, proposing a unified classification framework that dichotomizes optimization strategies into parameter-driven approaches (e.g., supervised fine-tuning, PPO, DPO) and parameter-agnostic techniques (e.g., prompt engineering, retrieval-augmented generation, reward shaping, trajectory construction). It further analyzes critical integrative aspects—such as hybrid optimization—and synthesizes evaluation benchmarks and representative applications. Contributions include: (1) a structured, comprehensive review encompassing over 100 works; (2) an open-source, standardized reference library hosted on GitHub; and (3) a clear articulation of open challenges and actionable research directions. Collectively, this work provides both theoretical foundations and a reproducible toolchain for efficient LLM agent optimization.
本文针对多智能体系统中的提示优化问题,提出AgentGrad框架,通过顺序干预和语义文本梯度抽象方法改进现有文本梯度方法的局限性。
This work addresses the instability and performance degradation in AI agents caused by ambiguity in natural language prompts. To mitigate this issue, the authors propose an automated prompt refinement mechanism grounded in semantic trajectory analysis. By monitoring agent execution logs and extracting semantic features to detect undesirable behaviors, the system dynamically injects corrective instructions, enabling incremental, data-driven optimization of system prompts. Integrated into the open-source Agent Mentor library, this approach synergistically combines large language models, log analysis, and dynamic prompt engineering. Empirical evaluations across diverse agent configurations and benchmark tasks demonstrate significant improvements in accuracy, with particularly pronounced gains in scenarios where initial prompts exhibit semantic ambiguity.
To address the poor adaptability of semantic communication to dynamic environments and resource constraints, as well as its low collaborative efficiency in AI agent-to-agent communication, this work proposes a semantics-driven lightweight cooperative communication framework. Methodologically, it innovatively integrates semantic-adaptive transmission, semantic encoding via model fine-tuning and generative sample adaptation, lightweight transmission via pruning-quantization and perception-driven sampling, and a distributed hierarchical self-evolving control mechanism—enabling end-to-end co-optimization across semantic representation, transmission, and decision-making. Simulation results demonstrate that, compared with conventional approaches, the framework achieves a 37% faster convergence rate, reduces communication overhead by 52%, and significantly enhances robustness under time-varying channels and heterogeneous node conditions. It thus establishes a scalable, adaptive semantic collaboration paradigm for AI-native networks.
Existing ML engineering agents rely solely on large language model (LLM) prompting and lack experience-driven, continual optimization capabilities. This paper introduces the first reinforcement learning (RL) agent framework specifically designed for ML engineering tasks, breaking away from conventional prompting-based paradigms. Our method features: (1) a duration-aware gradient update mechanism to mitigate delayed reward propagation in long-action sequences; (2) static LLM-based environment instrumentation, enabling automatic logging injection to generate fine-grained, partially rewardable execution feedback; and (3) a distributed asynchronous RL training architecture for scalable and efficient policy optimization. Evaluated on 12 Kaggle tasks from MLEBench, our RL-trained Qwen2.5-3B agent achieves an average 22% performance gain over the Claude-3.5-Sonnet prompting baseline—demonstrating that lightweight models can attain superior engineering intelligence through experience-driven evolution.
This work addresses the challenge of semantic drift in multi-agent scientific computing, which often leads to inconsistencies between policy selection and execution outcomes, thereby undermining causal traceability and adaptive learning. To mitigate this issue, the paper proposes a multi-agent system that integrates contextual bandits, a structured semantic communication protocol, and a semantic checkpointing mechanism, embedding the principle of empowerment into the decision-making process to preserve semantic consistency along action–outcome chains. The framework leverages large language model–driven specialized agents, code-generation verification, and a self-repairing execution loop. Evaluated on sensitivity analysis and uncertainty quantification tasks, the approach significantly enhances policy convergence, robustness, and generalization to novel scenarios, effectively suppressing semantic drift.
This study addresses the limitation of existing structured policy generation methods, which rely on manual or static knowledge and struggle to align with expert demonstrations. To overcome this, we propose a closed-loop iterative framework leveraging large language models (LLMs). The approach semanticizes rollout data into tabular formats, enabling LLMs to automatically diagnose and rectify structural deficiencies in policies. By utilizing tabularized rollout analysis as a feedback signal, the framework achieves automated alignment of policy structures without human intervention. Furthermore, this work integrates imitation learning with automated code generation techniques. Experimental results demonstrate that the proposed method improves performance by 15% while reducing computational costs by 75%, significantly enhancing overall policy generation efficiency.
This study addresses the limitation of conventional diffusion policies in offline reinforcement learning, where KL penalties suppress high-value yet low-density actions and existing methods struggle to model multimodal behaviors. To overcome these challenges, this work proposes PReFlow, a framework integrating critic-based proposal selection with conditional refinement flows. The method introduces a proposal-centric Gaussian reference distribution to regularize action variations and employs a simulation-free, closed-form adjoint matching objective to circumvent backward adjoint computation, while optimizing the Gibbs policy via KL regularization. Evaluated across 50 OGBench tasks, PReFlow demonstrates superior performance, achieving the highest aggregate score following online fine-tuning and reaching a 91% success rate within 500,000 interaction steps.
Traditional optimization approaches rely on external controllers—such as evolutionary algorithms or multi-armed bandits—to handle prompts, programs, and machine learning workflows separately, lacking a unified, autonomous decision-making mechanism. This work proposes ReASearch, a novel framework that unifies these three optimization tasks within a single reasoning-capable agent architecture. The agent autonomously evaluates objectives, diagnoses failures, edits solutions, and manages validation and restart strategies. Integrating tool invocation, persistent memory, domain-specific tool integration, and self-directed budget allocation, ReASearch matches or surpasses existing specialized systems across 14 tasks, achieving performance gains of 2%–40% and, in some cases, discovering solutions superior to the best known human-designed ones.
Existing activation intervention methods struggle to effectively steer agent behavior even in simple decision-making tasks. This work formulates behavioral intervention as a reinforcement learning problem for the first time, constructing removable, composable, and reversible task vectors by accumulating policy gradients toward temporary behavioral objectives over a small number of trajectories. The resulting framework enables dynamic behavioral modulation and supports cross-task customization and composition of behaviors. Empirical validation demonstrates calibrated and reversible interventions in grid-world environments, flexible composition of tactical objectives in chess, and successful modification of team-specific behaviors in a football simulation setting, with effective generalization across diverse opponents.
This study addresses the performance bottlenecks of fixed-parameter language agents caused by localized decision errors and the limited reusability of their procedural experience. To overcome these limitations, this work proposes an external strategy memory approach that introduces a novel natural language policy gradient mechanism requiring no modifications to model parameters or program structures. By leveraging execution trajectory diagnosis, module graph feedback propagation, and strategy aggregation, the method transforms failures into localized natural-language corrections, enabling continuous evolution and interpretable experience reuse for frozen agents. Evaluated across six benchmarks, the proposed approach outperforms the strongest baseline by an average of 8.71 percentage points, demonstrating its effectiveness in enhancing the capabilities of fixed-parameter language agents.