Reasoning Quality Matters: Combating Reasoning Collapse in LLM-based Embedding Learning

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对LLM嵌入学习中的推理崩溃问题,提出CoFree框架,通过两阶段优化保留推理质量并提升嵌入效果。
📝 Abstract
Large Language Models (LLMs) have recently shown strong potential for producing context-rich text embeddings for retrieval. Most existing methods either treat embedding learning as passive feature extraction or exploit LLM reasoning through instruction following for better embedding optimization. However, specialization toward embedding objectives can suppress useful reasoning generation or produce retrieval-irrelevant text. We refer to these two forms of degradation as reasoning collapse. To address this issue, we propose CoFree (Collapse-Free Reasoning Embedding), a two-stage framework that progressively integrates LLM reasoning into query and document embedding optimization while preserving reasoning quality. At the first stage, CoFree applies reference-guided supervised fine-tuning to restore the reasoning ability and retain representational strength of the foundation embedding model. At the second stage, we introduce dual rewards, an embedding-oriented reward and a reasoning-oriented reward, to guarantee fine-grained reasoning of the relevance toward the embedding goal in reinforcement learning. This endpoint-coupled optimization transforms embedding learning from static alignment into a high-quality reasoning-guided search process for retrieval. Extensive experiments demonstrate the effectiveness of CoFree, with CoFree-4B achieving an average absolute improvement of 2.8 nDCG@10 points over Qwen3-Embedding-4B across 22 datasets from MTEB and BRIGHT. Online experiments in a real-world retrieval system further show consistent gains. Code, RTED, and model checkpoints will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

Reasoning Collapse
Embedding Learning
Large Language Models
Retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Collapse-Free Reasoning
Dual Rewards
Supervised Fine-Tuning
Reasoning Quality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.