🤖 AI Summary
This study addresses the limited instruction-awareness of text embedding models by proposing Attention Relay, a training-free cross-model attention transfer mechanism. This method directly transfers attention weights from large language models (LLMs), such as Qwen3 and Llama 3.1, to general-purpose Transformer-based embedding models, endowing them with instruction-following capabilities without requiring additional training. Experimental results demonstrate that this mechanism is consistently effective across diverse combinations of LLMs and embedding architectures. It significantly enhances the ability of embedding representations to focus on instruction-relevant content. Consequently, this work establishes an efficient, training-free paradigm for constructing instruction-aware embedding models, offering a practical solution to bridge the functional gap between generative LLMs and representation learning systems.
📝 Abstract
Text embedding models trained with contrastive learning learn to follow task instructions from instruction-paired data, while instruction-tuned LLMs already know how to follow them. We show that this instruction-following ability can carry over from an LLM to a Transformer-based embedder without any training. We propose Attention Relay, which passes the attention weights an LLM produces to the embedder's own attention. Across six instruction-tuned LLMs from the Qwen3, Llama 3.1 and OLMo 3 families and ten widely used embedding models that differ in tokenizer, size and pooling type, Attention Relay makes nearly every combination instruction-aware. Experiments that break the method down into its parts show that the LLM's attention weights track the instruction in its later layers and come largely from instruction tuning. They also show that relaying these weights selects which content in the text matters: it makes the aspect of the text that the instruction asks about dominant in the embedding, or restores that aspect where averaging had diluted it.