arXiv Science⌕ Search

arXiv · 2610.05690

Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts

Abstract

Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LLM platforms in retrieving references related to original articles from five leading environmental science journals (Energy and Environmental Science, Nature Sustainability, Nature Climate Change, Lancet Planetary Health, and Environmental Science and Technology) published in 2024 to 2025. Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini were used as the LLM platforms. LLMs retrieved 10 references for each of the 50 randomly selected original article using either the article's abstract or its full-text as prompt. The retrieved references were subject to a multimetric score ratio combining validity of bibliographic data, Google Scholar link, digital object identifier, Scopus Electronic Identifier and relevance score (cited by or being the index paper), and the proportion of complete fabrication that failed all metrics. Abstract-only prompt yielded significantly higher accuracy than full-text one. This advantage was confirmed in multilevel mixed-effect multivariable regression after adjusting for journal, platform, and output order. Source journal and the position of a reference within the output list were also independently associated with retrieval accuracy, with lower-listed references associated with lower accuracy. These findings suggest that LLM assisted literature retrieval in environmental science remains moderately accurate and overall inconsistent, varying significantly by platform, journal, prompt type, and output position. Abstract-based prompting, as task-aligned information compression, may outperform full-text one in literature retrieval. Caution should be used when generalizing our findings.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yanjun Chen, Yongfeng Zhang, Lanjing Zhang. 2026-10-05. Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts. https://arxiv.org/abs/2610.05690

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Challenges of PROTAC Permeability Prediction

Cell permeability is a key bottleneck for PROTAC development, and public data available to model it is scarce and inconsistent. We adapt an expert-in-the-loop LLM extraction workflow to mine PAMPA measurements from the primary literature, recovering image-only structures by optical chemical structure recognition and hand-verifying every record, expanding the public record from 31 PROTACs to 87. Ridge models trained on PROTAC-DB 3.0 reach $R^2 = 0.67$ within that resource but collapse on the newly extracted chemistry ($ρ= 0.12$), while models trained on the new compounds transfer back successfully ($ρ= 0.80$). We conclude that the current composition of the published records, and not dataset size, is limiting the construction of more generalizable models, and we outline what would need to change in reporting practices for better data-driven permeability models.

cs.DL↗

Large scientific teams are more, not less, disruptive

Modern science has moved decisively toward larger and more collaborative research teams. Yet an influential 2019 study reported that small teams are more disruptive than large teams, a conclusion difficult to reconcile with both the well-known trend toward increasing collaboration and the disproportionately large teams behind widely recognized disruptive research, such as work that has been awarded Nobel Prizes. Here we show that the reported disruptive advantage of small teams is largely an artifact of measurement rather than a fact about teams. Analyzing about 60 million scientific publications, we find that larger teams cite more references and that the standard disruption index declines mechanically as reference lists lengthen. Equalizing reference-count distributions across team sizes reverses the negative gradient within the index's own framework. Evaluating each paper against one reference at a time yields a reference-robust measure that distinguishes recognized breakthroughs from ordinary papers more sharply and more consistently than the standard disruption index across six independently curated benchmarks, including the Nobel Prize. Under this validated measurement, scientific disruption increases with team size across four scientific domains and in every decade from the 1970s to the 2010s. For scientific publications, we conclude that larger teams are more, not less, disruptive.

cs.DL↗

Citations Are Late: Reading epistemic instability from what papers believe, years before the citation graph catches up

Paradigm shifts in science are visible in what researchers assert and contest before they are visible in the citation graph. We ask whether a cheap, content-level signal of epistemic instability, derived from the changing distribution of stated modelling beliefs in paper abstracts, can anticipate a paradigm shift earlier than the dominant citation-based disruption index (CD5). On the displacement of recurrent networks by Transformers in NLP, a pre-registered content signal crosses its detection threshold in 2016-Q1, whereas a real-time CD5 monitor cannot even observe the 2017 breakthrough until 2022-Q2, since CD5 needs a five-year forward-citation window: a lead of about 25 quarters. The flat citation baseline is not an artifact of one index, as our CD5, a reference-normalised variant, and an authoritative precomputed index all sit near zero across the shift. The lead is also not a faster proxy: at equal latency the signal beats the content competitors tested, including a learned CD-from-text model and an embedding disruption measure, and the decomposition sees what a keyword cannot (keyword-blind AUC of about 0.9, reproduced on human labels). The lead over CD5 is an observability lead. Which signal carries it depends on the shift: belief adoption in NLP, contestation and applicability stress in vision. Measured against the breakthroughs themselves, the signal leads by five quarters in NLP and is contemporaneous in computer vision.

cs.DL↗