Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of critical semantics being diluted by redundant information during mean pooling over long documents by proposing the SCSP framework. This work introduces a novel training-free selective pooling mechanism based on semantic compression, which evaluates token importance via semantic compression prompts. By integrating sentence-aware chunking with prompt-isolated attention masks, the method selectively aggregates intermediate-layer features from large language models to generate high-quality embeddings. The proposed framework operates in a plug-and-play manner and achieves consistent performance improvements across long-context benchmarks for both zero-shot and fine-tuned models.
📝 Abstract
Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long documents, such mean pooling can dilute salient semantic information with abundant redundant or weakly informative content. To this end, we propose SCSP, a training-free framework that leverages semantic compression for informative token selection in long-context embedding. Specifically, SCSP first partitions a document into sentence-aware chunks and appends a semantic compression prompt to each chunk. A prompt-isolated attention mask preserves information flow among document tokens while restricting each prompt to its corresponding local context. We then use the attention patterns elicited by these prompts to estimate token importance, select informative tokens, and aggregate their intermediate-layer representations into the final embedding. Extensive experiments on long-context embedding benchmarks demonstrate that SCSP can be integrated into both zero-shot and fine-tuned models in a plug-and-play manner, consistently improving their performance.
Problem

Research questions and friction points this paper is trying to address.

long-context embeddings
mean pooling
semantic dilution
informative token selection
redundancy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Compression
Selective Pooling
Long-Context Embeddings
Training-free Framework
Prompt-Isolated Attention Mask
🔎 Similar Papers
No similar papers found.