🤖 AI Summary
This study addresses the vulnerability of prompt-guided conditional text embeddings in large language models (LLMs) to generic semantic interference and their suboptimal representation quality. To this end, we propose an inference-time self-contrastive guidance method. Specifically, unconditional embeddings are constructed by modifying attention masks and positional encodings, after which vector-level contrastive correction is performed by intervening in multi-head self-attention computations, thereby refining conditional embeddings to focus on target semantics. This approach requires no training or additional data and functions as a plug-and-play module with only a single additional multi-head attention computation. Extensive evaluations on clustering, semantic textual similarity (STS), and triplet alignment tasks demonstrate that our method significantly enhances the performance of diverse LLM baselines, exhibiting both computational efficiency and broad generalizability.
📝 Abstract
Extracting conditional text embeddings from large language models (LLMs) is a promising paradigm, as it requires neither additional data nor fine-tuning. Existing methods incorporate conditions into prompts to guide LLMs to focus on specific aspects and elicit conditional text embeddings. However, relying solely on prompts often fails to produce high-quality conditional text embeddings, as they remain entangled with general text embeddings, ultimately degrading their quality. To this end, we propose an inference-time, plug-and-play Self-Contrastive Steering (SCS) method that constructs unconditional general text embeddings and uses them to refine conditional text embeddings, making them more focused on the target condition. Specifically, we modify the attention mask and positional encodings to mask the condition, thereby obtaining unconditional text embeddings and intervening in the multi-head self-attention computation process. Notably, our method is highly efficient, requiring only a single additional multi-head self-attention computation at inference time. Extensive experiments on clustering, Semantic Textual Similarity, and triplet alignment datasets demonstrate that our method can seamlessly improve the performance of existing prompt-based methods across different LLMs in a training-free and plug-and-play manner. Our code will be released at https://github.com/zifengcheng/SCS