Score
Design and implement methods that compute a model loss and its gradients with respect to prompt tokens or their embeddings, use those gradient signals to identify high-impact token or span-level regions, and generate candidate local edits. Iteratively apply, evaluate, and select gradient-informed prompt edits (with stopping criteria and search strategies) to optimize the prompt for a specified objective such as accuracy or calibration.
This paper addresses the lack of a systematic theoretical framework for automated prompt engineering by proposing the first unified optimization-theoretic survey across modalities. It formalizes discrete, continuous, and hybrid prompt variables—including instructions, soft prompts, and in-context examples—as constrained optimization problems, applicable to text, vision, and multimodal tasks. Methodologically, it integrates gradient-based optimization, evolutionary algorithms, and reinforcement learning to enable end-to-end, objective-driven prompt generation. Key contributions are: (1) the first optimization-theory-driven taxonomic framework for cross-modal prompt engineering; (2) the explicit identification and delineation of two frontier directions—constrained optimization and agent-oriented prompt design; and (3) a comprehensive knowledge system spanning the full technical spectrum and application domains of automated prompt engineering, providing a scalable theoretical foundation and methodological guidance for both research and practice.
This paper challenges the theoretical foundations and explanatory power of “text gradient”-based automated prompt optimization methods, which metaphorically equate discrete text updates with continuous, differentiable gradient descent. Method: Through systematic LLM prompt fine-tuning experiments, multi-task comparative analysis, ablation studies, and behavioral attribution, we rigorously examine whether these methods operate as genuine gradient-based optimizers. Contribution/Results: We demonstrate that performance gains are not attributable to gradient update logic; instead, “text gradients” function merely as empirical heuristics without theoretical grounding in differentiable optimization. First, we formally establish their non-gradient nature. Second, we propose a novel conceptual framework for prompt optimization explicitly tailored to discrete text spaces. Third, we advocate shifting prompt engineering from analogical transfer (e.g., borrowing optimization metaphors from continuous domains) toward intrinsic, ontology-aware modeling. These findings call for a fundamental methodological rethinking of prompt optimization.
This work addresses the high sensitivity of text-to-image diffusion models to input prompts, which typically necessitates extensive manual tuning. The authors propose a model-agnostic, automated prompt optimization method that directly evolves prompt embeddings in CLIP latent space via a genetic algorithm, eliminating the need for textual rewriting. The fitness function integrates aesthetic quality, measured using LAION-Aesthetics V2, and image-text alignment, quantified by CLIPScore. Evaluated on 36 prompts from the P2 dataset, the approach significantly outperforms both Promptist and random search, achieving up to a 23.93% improvement in fitness scores. The framework is modular and scalable, offering a general-purpose solution for prompt refinement in text-to-image generation.
This work proposes a hierarchical attribution-based prompt optimization framework to address the limitations of existing methods, which often suffer from prompt drift that degrades performance on historical tasks and lack interpretability when generating prompts from scratch. The framework employs a dynamic attribution mechanism to precisely identify error-inducing patterns, integrates semantic-unit-level editing to preserve the functional structure of prompts, and introduces a multimodal-friendly, end-to-end optimization pipeline. Evaluated on benchmarks such as OCR-V2 and BBH, the approach significantly outperforms current automatic prompt optimization techniques, achieving superior efficiency while enhancing both interpretability and scalability. This study thus establishes a novel paradigm for prompt engineering that balances performance, transparency, and adaptability across diverse tasks.
Soft prompt tuning often suffers from catastrophic forgetting of general-purpose knowledge in vision-language models (e.g., CLIP) under few-shot settings, leading to performance worse than zero-shot inference. Method: We propose Gradient Alignment (GA), a novel optimization mechanism that constrains prompt gradient updates to align with the direction of zero-shot predictions derived from predefined prompts—thereby explicitly preserving task-agnostic, pre-trained knowledge without requiring additional data, regularization, or architectural modifications. Contribution/Results: GA effectively mitigates overfitting and inter-class interference. It consistently outperforms state-of-the-art prompt-tuning methods across diverse transfer scenarios—including few-shot learning, domain generalization, base-to-novel class adaptation, and cross-dataset transfer—delivering substantial improvements in both generalization stability and accuracy.
To address the lack of interpretability and controllability in large language model (LLM) prompt optimization, this paper proposes the Gradient-inspired Prompt Optimizer (GPO). GPO establishes, for the first time, a systematic analogy between prompt optimization and gradient descent, designing an interpretable and controllable iterative mechanism along two dimensions: (i) *update direction*, determined via retrieval over prompt trajectories, and (ii) *update method*, combining generative refinement with cosine-decayed constraint on edit distance. Crucially, GPO operates entirely at the prompt level—requiring no model parameter fine-tuning—via meta-optimization over prompts. Evaluated on Big-Bench Hard and MMLU, GPO achieves absolute improvements of 56.8% and 62.6% over strong baselines, respectively, substantially enhancing zero-shot and few-shot reasoning capabilities. This work introduces a novel paradigm for prompt engineering grounded in principled, gradient-motivated optimization.
Existing image editing quality assessment methods rely on handcrafted heuristic prompts, which struggle to generalize across diverse editing effects and fail to account for the continuity of the scoring space. To address these limitations, this work proposes DS-IEQA, a unified framework that adaptively learns evaluation criteria through a feedback-driven prompt optimization mechanism (FDMPO) and models the continuous structure of quality scores via a token-disentangled distance regression loss (TDRL). Leveraging a multimodal large language model, the proposed method achieves competitive performance without requiring additional training data, securing fourth place in Track 2 of the NTIRE 2026 X-AIGC Quality Assessment Challenge. This result demonstrates the framework’s effectiveness and strong generalization capability.
Prompt engineering for large language models (LLMs) faces key bottlenecks: heavy reliance on manual design, static updates, coarse-grained editing, and poor reusability of empirical knowledge. To address these, this paper proposes PromptFlow—a modular, TensorFlow-inspired framework for end-to-end prompt optimization. Methodologically, it introduces (1) differentiable, fine-grained prompt editing; (2) a hybrid optimization strategy integrating gradient-based meta-learning and reinforcement learning to enable dynamic policy selection and cross-task LLM experience transfer; and (3) a unified computational graph comprising meta-prompts, prompt operators, optimization modules, and evaluation components. Empirically validated across diverse NLP tasks, PromptFlow achieves substantial downstream performance gains using only minimal labeled data—demonstrating clear superiority over conventional static prompt engineering approaches.
为解决文本梯度方法在提示优化中的不稳定性问题,通过错误驱动精炼和正则化验证两种机制提出STEVE框架,提高优化稳定性和效果。
This study addresses the limited generalizability of existing automatic prompt optimization methods across tasks and models. Adopting a causal inference perspective, it systematically investigates the interplay between prompt-editing behaviors and task characteristics across diverse optimization frameworks, large language models, and NLP benchmarks. By integrating complementary approaches—including propensity score adjustment, cognitive load annotations, surface textual features, and editing motifs—the work uncovers the structural causes underlying prompt optimization failures: edits that increase complexity or introduce meta-instructions impair performance on mathematical and multi-hop reasoning tasks, whereas step-by-step guidance and metacognitive edits substantially enhance logical and sequential reasoning. These findings demonstrate consistent and generalizable patterns across multiple experimental frameworks.
Existing prompt optimization methods treat prompts as monolithic units, making it difficult to localize errors, preserve critical instructions, or control prompt inflation—limitations that particularly hinder the reasoning performance of small open-source models. This work proposes Modular Prompt Optimization (MPO), a novel framework that introduces segment-wise local optimization based on fixed semantic blocks such as system roles and task descriptions. MPO leverages a critic model to generate text gradients at the segment level, enabling independent refinement of each module followed by deduplicated fusion, thereby achieving interpretable, robust, and efficient optimization while preserving the original prompt structure. Evaluated on the ARC-Challenge and MMLU benchmarks, MPO substantially outperforms both raw prompts and the TextGrad baseline, significantly boosting the reasoning accuracy of LLaMA-3-8B-Instruct and Mistral-7B-Instruct.