arXiv · 2403.06434
BoostER: Leveraging Large Language Models for Enhancing Entity Resolution
Abstract
Entity resolution, which involves identifying and merging records that refer to the same real-world entity, is a crucial task in areas like Web data integration. This importance is underscored by the presence of numerous duplicated and multi-version data resources on the Web. However, achieving high-quality entity resolution typically demands significant effort. The advent of Large Language Models (LLMs) like GPT-4 has demonstrated advanced linguistic capabilities, which can be a new paradigm for this task. In this paper, we propose a demonstration system named BoostER that examines the possibility of leveraging LLMs in the entity resolution process, revealing advantages in both easy deployment and low cost. Our approach optimally selects a set of matching questions and poses them to LLMs for verification, then refines the distribution of entity resolution results with the response of LLMs. This offers promising prospects to achieve a high-quality entity resolution result for real-world applications, especially to individuals or small companies without the need for extensive model training or significant financial investment.
Explore related subjects
Keep this discovery
Huahang Li, Shuangyin Li, Fei Hao, Chen Jason Zhang, Yuanfeng Song, Lei Chen. 2024-03-11. BoostER: Leveraging Large Language Models for Enhancing Entity Resolution. https://doi.org/10.1145/3589335.3651245%2010.1145%2F3589335.3651245%2010.1145%2F3589335.3651245
Cite the original work for its findings. Save a collection to share your selection of sources.