arXiv · 2610.09446
Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science
Abstract
Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at https://github.com/BenWilcox8/arctic-qa.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Benjamin Wilcox, Dawei Gao, Pradeeban Kathiravelu, Douglas Causey, Kewei Sha, Yunhe Feng. 2026-10-07. Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science. https://arxiv.org/abs/2610.09446
Cite the original work for its findings. Save a collection to share your selection of sources.