arXiv · 2609.35430
Can Generative Retrievers Learn Semantic IDs Without Forgetting How to Speak?
Abstract
Generative retrieval (GR) enables end-to-end retrieval by generating document semantic identifiers (SIDs). However, retrieval-only fine-tuning can over-specialize pretrained language models to SID prediction, substantially distorting their natural-language distribution and limiting their suitability for interactive systems that must both retrieve documents and generate natural-language responses. We introduce SpeakGR, a dual-objective framework that learns SIDs while preserving language generation. It combines supervised SID learning with speak-preserving regularization: an on-policy distillation objective that aligns the current model with a frozen copy of the original model on student-generated prefixes using forward KL over the original text vocabulary. We further propose Adaptive SpeakGR, which dynamically adjusts the preservation strength based on observed language drift. Compared with SFT-only, SpeakGR reduces WikiText-2 forward KL by 81.3-93.8% on MS MARCO and 81.2-85.2% on Natural Questions (NQ) while retaining effective retrieval across three different LLMs. Adaptive SpeakGR further improves retrieval over SpeakGR in most settings while maintaining substantially lower language drift than SFT-only.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junchen Fu, Kleomenis Katevas, Vandana Rajan, Sofía Celi, Hamed Haddadi. 2026-09-28. Can Generative Retrievers Learn Semantic IDs Without Forgetting How to Speak?. https://arxiv.org/abs/2609.35430
Cite the original work for its findings. Save a collection to share your selection of sources.