arXiv · 2609.17644
Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models
Abstract
Domain-specialized language models are widely used for scientific question answering, but stronger general-purpose systems raise a sharper question: when does domain-specific fine-tuning remain valuable for open-ended scientific reasoning? We study this in astronomy with a curated QA benchmark from publicly available 2017--2026 Olympiad-style materials. The free-response subset contains 300 questions, including 204 text-only and 96 image-linked examples. We compare open-weight and API-served general-purpose, multimodal, and astronomy-specialized models using judge-based correctness and complementary reference metrics. Strong general-purpose models establish the highest correctness baseline in this testbed, while analyses of metric agreement, judge sensitivity, benchmark composition, and modality reveal variation not captured by a single leaderboard. These results motivate treating domain specialization as a task- and deployment-dependent property and highlight the role of domain-specific evaluation in determining which models, capabilities, and evaluation criteria are appropriate for scientific workflows.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Vanessa Lama, Sanjay Das, Emily Herron, Yuan-Sen Ting, Tijmen de Haan, Junqi Yin, Tirthankar Ghosal, Feiyi Wang. 2026-09-15. Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models. https://arxiv.org/abs/2609.17644
Cite the original work for its findings. Save a collection to share your selection of sources.