PC-SubMax: Efficient Prompt Compression via Regularized Submodular Maximization

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference costs, latency, and "lost-in-the-middle" phenomena induced by lengthy prompts, as well as the limitation of existing methods that overlook inter-sentence redundancy. To this end, we propose a selective prompt compression framework grounded in submodular optimization. Specifically, this work formulates compression as regularized monotone submodular maximization subject to knapsack constraints, constructing a utility function that integrates information coverage, query relevance, and diversity. Furthermore, we design the RGM algorithm coupled with encoder-based feature extraction to achieve efficient optimization with a theoretically guaranteed approximation ratio, thereby circumventing the overhead of autoregressive scoring. Extensive evaluations across seven benchmarks demonstrate that the proposed approach substantially reduces compression overhead while preserving highly competitive downstream performance.
📝 Abstract
While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle''phenomenon. Selective prompt compression offers a model-agnostic approach to alleviating these issues. However, methods based on fixed token- or sentence-level importance scores may overlook how content contributions change with the selected subset, limiting their ability to account for inter-sentence redundancy. Compression procedures that rely on autoregressive LLM scoring can also introduce substantial overhead. We propose PC-SubMax, a theoretically grounded framework that formulates selective prompt compression as regularized monotone submodular maximization under a knapsack constraint. The objective is $U(S)-\ell(S)$, where the monotone submodular utility $U$ combines information coverage, query relevance, and log-determinant diversity, and the non-negative modular penalty $\ell$ captures token cost. Through diminishing marginal returns, the objective evaluates each sentence's contribution relative to the selected content. To optimize this objective, we develop the Regularized Greedy+Max (RGM) algorithm, which deterministically returns a feasible set $Q$ satisfying $U(Q)-\ell(Q)\geq \frac{1}{2}U(O)-\ell(O)$, where $O$ is an optimal feasible solution to the regularized problem. RGM uses $O(n\kappa)$ value-oracle queries, where $n$ is the number of candidate sentences and $\kappa$ is the maximum feasible subset size. PC-SubMax uses encoder representations and avoids autoregressive LLM scoring during compression. Experiments across seven diverse benchmarks demonstrate competitive downstream performance with low compression overhead.
Problem

Research questions and friction points this paper is trying to address.

Prompt Compression
Large Language Models
Long-context Scenarios
Lost-in-the-middle
Inference Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Compression
Submodular Maximization
Regularized Optimization
Large Language Models
Knapsack Constraint
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.