🤖 AI Summary
This study addresses the instability and lack of reliability analysis in large language model outputs caused by opaque and redundant prompt structures. We propose a few-shot prompt minimization framework from a black-box perspective. By integrating black-box optimization-based ablation experiments, character-level statistical analysis, and logical identifier preservation strategies, the method distills prompts to their causally essential subsets while revealing divergent model preferences when functioning as universal encoders versus decoders. Experiments demonstrate that the framework achieves an average 65.3% reduction in character count while fully preserving propositional output fidelity. These findings confirm that models preferentially rely on logical identifiers over natural language descriptions, which can be safely discarded. This work establishes a novel paradigm for efficient prompt compression and structural interpretation.
📝 Abstract
Prompts are the primary mechanism for directing the behavior of large language models (LLMs). Yet the internal structure and causal hierarchy of prompts remain poorly understood: which parts are causally necessary and which are redundant is an open question. This opacity can have severe consequences. Subtle prompt variations can silently shift model outputs in critical software systems, and engineers lack techniques to reason about prompt reliability.
We present \framework, a blackbox prompt-minimization framework that reduces few-shot prompts to their necessary minimal subset. We use a case study to apply \framework to a few-shot learning system and demonstrate the insights that this framework can provide.
Our experiments show that few-shot exemplars can be reduced by a mean of 65.3\%~$\pm$~15.8\% in character count while fully preserving propositional output fidelity. The models preferentially retain logical identifiers and constraint declarations while discarding natural language prose and cross-prompt relational annotations.
Our analysis also shows that some models are universal encoders, able to produce highly legible yet minimized prompts, while others are universal decoders, able to interpret minimized prompts from most other models.
By identifying which components are indispensable, \framework provides a principled basis for prompt compression and structural analysis of few-shot exemplars.