arXiv · 2609.30897
From annotation to reasoning: Culture in language models
Abstract
How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpretation offers a useful setting for studying these capabilities. We focus on cultural referencing and reuse: how texts invoke, repeat, and transform earlier expressions across historical and linguistic contexts. Our central claim is that literary scholars can disagree about an interpretation while recognizing the quality of its support. We propose linking evidence-centered benchmarks, evaluation that preserves scholarly disagreement, and model-development experiments on literary data, contextual resources, and scholarly feedback. Danish literature provides a concrete starting point, with implications for other languages and domains. The aim is to develop alternative evaluation strategies that go beyond conventional benchmark metrics and guide model development toward cultural robustness in AI systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daniel Hershcovich, Alexander Conroy, Jens Bjerring-Hansen. 2026-09-25. From annotation to reasoning: Culture in language models. https://arxiv.org/abs/2609.30897
Cite the original work for its findings. Save a collection to share your selection of sources.