๐ค AI Summary
Fine-tuning language models for generative retrieval often leads to the degradation of their natural language capabilities. To address this issue, this work proposes SpeakGR, a dual-objective framework that enables models to acquire semantic identifiers while preserving language generation proficiency. The method introduces an adaptive mechanism to dynamically modulate preservation intensity and integrates supervised learning, online policy distillation, and forward KL divergence constraints to effectively balance retrieval performance against language drift. Experimental results demonstrate that SpeakGR substantially mitigates language distribution shift while maintaining competitive retrieval effectiveness across multiple benchmark datasets.
๐ Abstract
Generative retrieval (GR) enables end-to-end retrieval by generating document semantic identifiers (SIDs). However, retrieval-only fine-tuning can over-specialize pretrained language models to SID prediction, substantially distorting their natural-language distribution and limiting their suitability for interactive systems that must both retrieve documents and generate natural-language responses. We introduce SpeakGR, a dual-objective framework that learns SIDs while preserving language generation. It combines supervised SID learning with speak-preserving regularization: an on-policy distillation objective that aligns the current model with a frozen copy of the original model on student-generated prefixes using forward KL over the original text vocabulary. We further propose Adaptive SpeakGR, which dynamically adjusts the preservation strength based on observed language drift. Compared with SFT-only, SpeakGR reduces WikiText-2 forward KL by 81.3-93.8% on MS MARCO and 81.2-85.2% on Natural Questions (NQ) while retaining effective retrieval across three different LLMs. Adaptive SpeakGR further improves retrieval over SpeakGR in most settings while maintaining substantially lower language drift than SFT-only.