Score
Designs, implements, and evaluates methods that create, organize, and apply hierarchical prompt structures and prompt memories—including non‑parametric, probabilistic, and mixture‑of‑prompts approaches—to condition model internals and outputs. Builds prompt‑based conditioning and context‑aware prompt tuning workflows that modulate encoder/decoder features and adapt model behavior across varying contextual levels to improve robustness and generalization.
This paper presents a systematic survey of prompt tuning—a parameter-efficient paradigm for adapting pre-trained language models—focusing specifically on the setting where the backbone model is frozen and only continuous, prefix-based prompt embeddings are optimized. Addressing key challenges including computational inefficiency and training instability, the work introduces the first unified taxonomy encompassing encoder-based, low-rank decomposition, and mixture-of-experts prompt tuning methods, rigorously distinguishing direct prompt learning from transferable prompt learning. Through methodological analysis and visualized performance comparisons across diverse benchmarks, it characterizes fundamental trade-offs among parameter count, optimization convergence, and generalization capability. The study provides both theoretical insights and practical guidelines for enhancing training robustness and extending prompt tuning to multi-task and low-resource scenarios.
Existing prompt-based continual learning methods suffer from catastrophic forgetting due to layer-wise independent prompt updates. To address this, we propose Hierarchical Grouped Prompt Tuning (HGPT): model layers are partitioned into groups, with prompts shared within each group; positional encoding is incorporated to preserve feature structural stability. Furthermore, a root-prompt generation mechanism is introduced, wherein a hierarchical network dynamically derives child prompts from a root prompt, enhancing inter-prompt coordination and generalization. HGPT achieves efficient task-specific prompt allocation and cross-task feature alignment without modifying any pre-trained parameters. Evaluated on four standard continual learning benchmarks, HGPT consistently outperforms state-of-the-art prompt-tuning approaches, striking a superior balance between adaptation to new tasks and retention of knowledge from previous tasks.
This work challenges the conventional assumption in fine-tuning that semantically equivalent training prompts yield comparable performance, highlighting their critical yet overlooked impact on cross-task forgetting and generalization. To address this, the authors propose State-Adaptive Prompt Optimization (SAPO), a lightweight mechanism that dynamically adjusts the form of training prompts during learning. SAPO identifies high-quality prompts in real time based on pre-trained task losses and adaptively updates them according to the model’s current training state. This approach substantially mitigates catastrophic forgetting and enhances generalization, consistently outperforming state-of-the-art methods across multiple benchmarks. The results underscore the pivotal role of training prompt design in learning dynamics and demonstrate its inherent optimizability.
This work proposes a hierarchical attribution-based prompt optimization framework to address the limitations of existing methods, which often suffer from prompt drift that degrades performance on historical tasks and lack interpretability when generating prompts from scratch. The framework employs a dynamic attribution mechanism to precisely identify error-inducing patterns, integrates semantic-unit-level editing to preserve the functional structure of prompts, and introduces a multimodal-friendly, end-to-end optimization pipeline. Evaluated on benchmarks such as OCR-V2 and BBH, the approach significantly outperforms current automatic prompt optimization techniques, achieving superior efficiency while enhancing both interpretability and scalability. This study thus establishes a novel paradigm for prompt engineering that balances performance, transparency, and adaptability across diverse tasks.
Prompt engineering for large language models (LLMs) faces key bottlenecks: heavy reliance on manual design, static updates, coarse-grained editing, and poor reusability of empirical knowledge. To address these, this paper proposes PromptFlow—a modular, TensorFlow-inspired framework for end-to-end prompt optimization. Methodologically, it introduces (1) differentiable, fine-grained prompt editing; (2) a hybrid optimization strategy integrating gradient-based meta-learning and reinforcement learning to enable dynamic policy selection and cross-task LLM experience transfer; and (3) a unified computational graph comprising meta-prompts, prompt operators, optimization modules, and evaluation components. Empirically validated across diverse NLP tasks, PromptFlow achieves substantial downstream performance gains using only minimal labeled data—demonstrating clear superiority over conventional static prompt engineering approaches.
Current prompt engineering research lacks systematic taxonomies and comparable evaluation protocols. Method: This paper introduces the first comprehensive taxonomy spanning large language models (LLMs) and multimodal models, categorizing over one hundred prompt techniques by application scenario and uniformly specifying their supported models, benchmark datasets, and boundary conditions. We integrate bibliometric analysis, cross-model/cross-dataset empirical comparison, methodological abstraction, and taxonomy construction techniques; further proposing a standardized evaluation framework and an interactive knowledge graph to clarify strengths, limitations, and open challenges of each technique. Contribution/Results: We deliver a structured technical survey, a complete classification table, and reusable evaluation dimensions—establishing the first authoritative benchmark and research navigation toolkit for prompt engineering.
This work addresses the limitations of existing in-context learning approaches, which treat prompts merely as semantic cues and thus fail to enable task-adaptive dynamic computation, resulting in shallow and uninterpretable reasoning. To overcome this, the authors propose PromptPath, a novel framework that directly integrates prompt information into the model’s inference architecture. PromptPath employs a prompt-conditioned routing mechanism to dynamically activate and compose lightweight low-rank expert modules, thereby constructing task-specific computational pathways. This approach achieves dynamic adaptation at the computational level, significantly outperforming current methods on both 3D point cloud and 2D visual recognition benchmarks while demonstrating strong cross-domain and cross-task generalization capabilities.
This study addresses the computational overhead, inference latency, and accuracy degradation caused by prompt redundancy in large language models (LLMs) by proposing a novel prompt minimization paradigm. Methodologically, we construct an LLM-based multi-version prompt optimization and evaluation framework that employs three strategies to identify the most concise, high-density inputs. This work reveals substantial redundancy within the input space, establishes new evaluation criteria for minimalist prompts, and redefines the theoretical boundaries of efficient prompt engineering. Experimental results demonstrate that minimal prompts maintain output fidelity comparable to their longer counterparts while significantly reducing inference costs and enhancing overall system efficiency.
This work addresses the high cost and sensitivity to phrasing inherent in manually crafted prompts, as well as the limited ability of existing automated optimization methods to systematically identify and correct failure patterns. To overcome these challenges, the authors propose Reflective Prompt Tuning (RPT), a novel framework that introduces a reflection mechanism leveraging large language model function calling. RPT employs a diagnostic function to analyze failure modes on an optimization set, generates structured reports, and iteratively refines prompts by integrating historical memory with confidence calibration. Experimental results demonstrate that RPT achieves performance gains of up to 12.9 points across three reasoning tasks, with particularly pronounced improvements in multi-hop and mathematical reasoning, while also enhancing the calibration of model output confidence.
This work addresses the limitations of existing automatic prompt optimization methods, which treat prompts as monolithic strings and thus struggle to model reusable sub-behaviors, resulting in fragile updates and poor adaptability across inputs. To overcome this, the authors propose Prompt Codebooks (PCO), a novel framework that introduces, for the first time, a discrete codebook-based compositional prompt optimization mechanism. PCO reformulates prompt construction as dynamic selection and composition from a finite set of natural language “instinct” units. It employs an LLM-driven encoder–generator–critic architecture to jointly train the codebook and routing policy while keeping the target model frozen, leveraging a linguistic value minimax objective and textual gradient decomposition to enable instance-specific customization. Evaluated on Qwen3-8B and LLaMA-3.1-8B, PCO achieves gains up to 30.36 points across six benchmarks over the strongest baseline, GEPA, while compressing prompt length to 1/14.1 of MIPROv2 and 1/3.0 of GEPA using only 16 atomic units.
为解决现有提示调优方法的局限性,提出HiVe框架,通过层次结构和垂直混合专家机制实现输入依赖的提示专业化,提高多任务学习性能。