🤖 AI Summary
This study addresses the computational overhead, inference latency, and accuracy degradation caused by prompt redundancy in large language models (LLMs) by proposing a novel prompt minimization paradigm. Methodologically, we construct an LLM-based multi-version prompt optimization and evaluation framework that employs three strategies to identify the most concise, high-density inputs. This work reveals substantial redundancy within the input space, establishes new evaluation criteria for minimalist prompts, and redefines the theoretical boundaries of efficient prompt engineering. Experimental results demonstrate that minimal prompts maintain output fidelity comparable to their longer counterparts while significantly reducing inference costs and enhancing overall system efficiency.
📝 Abstract
Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while preserving output fidelity. Practically, shorter prompts reduce computational overhead and inference latency, especially when large contexts, such as entire documents or codebases, are included unnecessarily. Further, longer prompts can damage LLM reasoning and accuracy. Theoretically, the existence of multiple prompts yielding equivalent outputs suggests a high degree of redundancy in the input space, raising fundamental questions about what information is essential to elicit specific model behaviors. We propose three variant frameworks to identify and evaluate minimal prompts and demonstrate that minimal prompts often produce outputs comparable to those of their longer counterparts. These findings suggest new directions for efficient prompt engineering and deepen our understanding of input compression in LLMs.