arXiv Science⌕ Search

arXiv · 2610.01139

Do Multilingual Encoders Produce Language-Consistent Semantic IDs?

Abstract

Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Abhinav Bohra, Anuj Bohra. 2026-10-01. Do Multilingual Encoders Produce Language-Consistent Semantic IDs?. https://arxiv.org/abs/2610.01139

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Route What Remains: A Meta-Modal Agent for Missing-Modality Candidate Reranking in Recommender Systems

Missing-modality recommenders usually reconstruct absent representations, although the observed evidence may not determine the missing content. We formulate candidate reranking as budgeted sequential evidence acquisition. A policy queries text, image, and interaction-graph tools, incorporates \texttt{Null} returns into its observation history, and sparsely rescores a retrieved candidate pool. Our \textbf{Meta-Modal Agent} (MMA) uses PPO to optimize terminal NDCG and tool cost without explicit access to the route-availability mask or target identity. When only one evidence route is available, MMA-Auto improves NDCG@10 by $10.0$\% over the strongest completion baseline and by $9.5$\% over a fixed router with the same Llama scorer. It obtains the highest result in all nine reported combinations of dataset and available route against these comparators. MMA-Auto also reduces failed calls by 17.8 percentage points and uses 1.1 fewer turns than the fixed router. On the fixed candidate pools produced by full-catalog retrieval, MMA-Auto improves NDCG@10 by $19.7$\%. These results associate adaptive evidence routing with improved reranking under severe, constructed missingness. The code is available at: https://anonymous.4open.science/r/WSDM2027-MMA-C381.

cs.IR↗

TAGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories

Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories. Localized graph configurations traverse a common base graph; AdaptiveGraph adds chronological edges and Personalized PageRank diffusion. We also evaluate BM25 over the same extracted notes and OpenClaw as a raw-input external reference. Retrieval rankings vary across memory settings. On LongMemEval-S, AdaptiveGraph is the strongest graph configuration at 0.844 MRR, but BM25 reaches 0.867 and OpenClaw 0.880. On ATANT Core, localized graph traversal outperforms diffusion and BM25, whereas BM25 leads the stress rounds. Reducing LongMemEval-S within the tested range does not reproduce the ATANT diffusion penalty, but the smallest tested store remains larger than ATANT Core, so store size cannot be ruled out. The penalty also persists under a permissive content-match criterion. Vocabulary normalization and extraction quality substantially affect graph retrieval, and missing extraction tags are common among top-five misses. Retrieval strategies should therefore be evaluated jointly with the memory setting and against strong lexical baselines.

cs.IR↗

Exploring Forum Post Retrieval with Generative Modeling

Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a standalone application for medium-to-heavy users of Facebook Groups. Because Forum is a new surface, its own interaction data are too sparse to train a GR model from scratch. We address this with transfer along two axes: we train on a broader corpus of Facebook Groups engagements rather than Forum sessions alone, and we reuse hierarchical, prefix-based semantic IDs (SIDs) learned from cross-platform Facebook Feed data instead of fitting a Forum-specific tokenizer. A 3B-parameter instruction-tuned language model is then supervised-fine-tuned to generate SIDs directly from user context. We systematically ablate the design choices that matter most in practice, including SID construction, the composition and length of user history, and the inclusion of user-profile features. Our results show that cross-platform SIDs transfer to a new recommendation surface, and offer practical guidance for teams deploying GR on real-world social platforms.

cs.IR↗