Score
Designs and implements mechanisms that adapt a model’s input prompt in real time—by appending notes, running reflective or self‑augmentation loops, or applying accept/reject gates—to change model behavior without updating model weights or using gradient-based training. Builds and evaluates online, gradient‑free prompt modification pipelines that detect anomalous outputs, isolate strategy improvements, and measure how incremental prompt edits affect downstream behavior.
This paper addresses the lack of a systematic theoretical framework for automated prompt engineering by proposing the first unified optimization-theoretic survey across modalities. It formalizes discrete, continuous, and hybrid prompt variables—including instructions, soft prompts, and in-context examples—as constrained optimization problems, applicable to text, vision, and multimodal tasks. Methodologically, it integrates gradient-based optimization, evolutionary algorithms, and reinforcement learning to enable end-to-end, objective-driven prompt generation. Key contributions are: (1) the first optimization-theory-driven taxonomic framework for cross-modal prompt engineering; (2) the explicit identification and delineation of two frontier directions—constrained optimization and agent-oriented prompt design; and (3) a comprehensive knowledge system spanning the full technical spectrum and application domains of automated prompt engineering, providing a scalable theoretical foundation and methodological guidance for both research and practice.
Prompt engineering for large language models (LLMs) faces key bottlenecks: heavy reliance on manual design, static updates, coarse-grained editing, and poor reusability of empirical knowledge. To address these, this paper proposes PromptFlow—a modular, TensorFlow-inspired framework for end-to-end prompt optimization. Methodologically, it introduces (1) differentiable, fine-grained prompt editing; (2) a hybrid optimization strategy integrating gradient-based meta-learning and reinforcement learning to enable dynamic policy selection and cross-task LLM experience transfer; and (3) a unified computational graph comprising meta-prompts, prompt operators, optimization modules, and evaluation components. Empirically validated across diverse NLP tasks, PromptFlow achieves substantial downstream performance gains using only minimal labeled data—demonstrating clear superiority over conventional static prompt engineering approaches.
Efficient, lightweight downstream adaptation of large language models remains challenging due to fragmented methodologies and lack of unifying principles. Method: This paper proposes a unified framework—“neural reprogrammability”—modeling parameter-free adaptation paradigms—including model reprogramming, prompt tuning, and prompt instruction—as targeted manipulations of information flow at interfaces such as input, intermediate layers, or context. Contribution/Results: We introduce the first cross-modal, architecture-agnostic four-dimensional taxonomy (format, location, operator, output alignment), revealing intrinsic unity among in-context learning, chain-of-thought, and related methods. By systematically integrating existing interface perturbation techniques—including input perturbation, token insertion, and example injection—we empirically validate their generality across multimodal foundation models. Our framework establishes foundational principles and provides actionable guidelines for lightweight, controllable, and interpretable model adaptation.
This work investigates the mechanism by which prompts induce behavioral switching in fixed-weight Transformers. Method: We propose the “Prompt-as-Program” theoretical framework, formalizing prompts as executable programs that leverage attention for memory routing, conditionally activate arithmetic operations in feed-forward networks (FFNs), and achieve multi-step compositional computation via depth-wise stacking—all atop a single frozen backbone. Using simplified Transformer modeling, mechanistic decomposition, and constructive existence proofs—grounded in computability and expressive power theory—we rigorously characterize capability boundaries and structural limits under constraints on prompt length and precision. Contribution/Results: This is the first unified formal analysis foundation for prompt engineering, transcending empirical practice and establishing a rigorous theoretical basis for prompt-driven computation in fixed-weight Transformers.
This work addresses two key limitations in black-box large language model (LLM) prompt optimization: underutilization of correct prediction signals and poor cross-model transferability. To this end, we propose an enhanced feedback-driven prompt optimization framework. Methodologically, it introduces a dual-track reinforcement mechanism—retaining effective prompt components via both positive and negative signals—integrates text-gradient reconstruction, multi-signal feedback aggregation, and noise filtering, and incorporates an explicit prompt transfer strategy. Our key contribution is the first systematic integration of positive reinforcement learning into automated prompt optimization, enabling active exploitation of correct prediction information. Experiments demonstrate that our method consistently outperforms strong baselines on both standard prompt optimization and cross-model/cross-API transfer tasks, achieving simultaneous improvements in accuracy, convergence speed, and computational efficiency.
Under rapid generative AI model iteration, users’ ability to adapt to evolving models critically determines the translation of technological advancement into economic value. Method: This paper introduces *prompt adaptation*—users’ deliberate refinement of input prompts—as a dynamic complementarity mechanism in generative AI. We systematically quantify its impact via online controlled experiments, analysis of over 18,000 real-world human-AI interactions, and image similarity-based evaluation. Contribution/Results: Approximately 49% of DALL·E 3’s performance gain stems from user-initiated prompt adjustments; replacing such human adaptation with automated prompt rewriting incurs a 58% loss in upgrade benefits. These findings demonstrate that user adaptive behavior accounts for nearly half of the value generated by model upgrades, underscoring its pivotal role in human-AI co-creation. The study provides empirical grounding for AI deployment strategies and human-centered interface design.
This study addresses the limitations of fixed architectures and instruction redundancy caused by repetitive revisions in large language model self-correction. To overcome these challenges, we propose WDA, a framework that jointly evolves stage-wise instructions and structures to generate task-adaptive self-correction pipelines. Methodologically, WDA introduces a SPLIT mechanism to disperse redundant instructions, integrating three-example reflection, local filtering, and Pareto admission strategies. It further combines prompt optimization, workflow design agents, and evolutionary search to achieve efficient calibration. Experimental results demonstrate that this framework improves average scores by 8.63 and 5.62 percentage points on Qwen3.5-9B and GPT-4.1-mini, respectively, significantly enhancing the self-correction capabilities of large language models.
Existing neural network editing methods rely on task-specific handcrafted algorithms, which are costly and exhibit poor generalization. This work proposes a unified, learnable framework by formulating model editing as a reinforcement learning problem for the first time. An agent learns to edit model parameters autonomously within two environments—MaskWorld (multiplicative mask scaling) and ShiftWorld (additive weight shifting)—guided by a multi-objective reward function that balances task-specific objectives with overall model performance preservation. Experiments demonstrate that the approach effectively reduces accuracy on forget sets to nearly 0% while maintaining over 90% accuracy on retain sets in machine unlearning tasks. In bias mitigation scenarios, it improves fairness metrics by more than 5% without compromising classification utility.
This study addresses the difficulty language model agents face in efficiently adapting execution frameworks to diverse tasks at test time. To this end, this work proposes "framework learning," which formulates framework revision as meta-learning over executable programs. Specifically, a proposer model is trained via reinforcement learning to iteratively refine a solver's code framework using execution feedback, thereby enabling test-time adaptation without parameter updates. By integrating large language model agents with program synthesis and automated repair techniques, this approach endows agents with the capacity to continuously generalize and improve from experience. Experimental results demonstrate significant performance gains on reasoning and multi-hop question answering tasks, validating that such test-time adaptation capabilities transfer effectively to unseen tasks.
This work addresses the challenge of continuously optimizing user-defined reasoning scaffolds at low cost within a framework of co-evolution between models and inference scaffolds, thereby enhancing the quality of agent execution trajectories. It proposes a recursive scaffold self-improvement mechanism that unifies scaffold optimization with model training data generation, representing scaffolds via prompt-level specifications and iteratively refining them through pairwise preference feedback derived from their own revision history. Rather than extending reasoning chains, the method emphasizes task context management to achieve efficient information flow control with minimal inference overhead. Evaluated on 30 cross-domain synthetic tasks, the approach significantly outperforms high-overhead baselines within just a few iterations, reducing inference costs by up to 60%, thus demonstrating the critical role of effective context management in performance improvement.
Existing prompt optimization methods struggle to proactively and error-free adapt to continuously evolving constraints in the absence of test-time feedback. To address this gap, this work introduces the first evaluation framework for continual prompt adaptation under active adaptation scenarios, strictly adhering to a “adapt-then-test” protocol to systematically assess methods’ continual learning capabilities—specifically, their susceptibility to forgetting, regression, and forward transfer—under dynamic constraints. We benchmark six prominent approaches across four large language models and three constraint evolution schedules, revealing that current methods yield no significant performance gains while incurring higher latency, thereby demonstrating their inadequacy for this paradigm. This study fills a critical void in evaluating prompt adaptation under dynamic constraints and zero-feedback settings.