arXiv · 2610.05573
Better Retrieval, Limited Clustering Gains: A Controlled Study of Multilingual Company Entity Resolution
Abstract
Improved name retrieval may have little effect on company clusters when the pair classifier remains unchanged. We examine this dependency by adapting multilingual E5 encoders under fixed candidate budgets and downstream decision rules. Random-negative and hard-negative training use identical positive schedules. Checkpoints are selected before collecting a new GLEIF sample of 3,633 names, 2,880 source identities and 882 silver-positive pairs. At 72,660 candidate edges, adaptation with a multi-view selector increases direct pair recall from 53.74% to 76.98%. The primary matcher adds only seven correct and two incorrect co-cluster pairs: cluster recall rises from 32.54% to 33.33%, while precision falls from 95.99% to 95.45%. Of 208 newly retrieved silver-positive pairs, 202 fall below its decision threshold. Random-negative and hard-negative training produce identical final partitions. An AI-assisted, single-reviewer audit of 137 pairs supports the observed pattern, although its predominantly LEI-derived evidence does not establish independent gold labels. The results locate the immediate loss of retrieval gains at the existing confirmation stage and show why encoder evaluation must also measure final cluster quality.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yijiashun Qi, Yuxuan Li, Hanzhe Guo. 2026-10-04. Better Retrieval, Limited Clustering Gains: A Controlled Study of Multilingual Company Entity Resolution. https://arxiv.org/abs/2610.05573
Cite the original work for its findings. Save a collection to share your selection of sources.