arXiv · 2609.23178
Chronologic: Measuring Language Models' Ability to Represent the Past
Abstract
Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct answers. We use historical texts to develop a benchmark for a model's representation of English-language contexts 1831-1930, relying on pairwise comparisons to multiple ground truths and strong distractors to score the hardest questions in an appropriately graduated way. We find that generative tasks are harder than discriminative ones; in fact, reasoning models can typically discern the weakness of their own generated answers. While models pretrained exclusively on historical text lead the pack when evaluated by answer likelihood, they cannot compete with commercial models in free generation. None of the models we tested represent historical contexts in a fully persuasive way yet, but progress toward that goal is evident.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Edwin Roland, Wenyi Shang, Matthew Wilkens. 2026-09-19. Chronologic: Measuring Language Models' Ability to Represent the Past. https://arxiv.org/abs/2609.23178
Cite the original work for its findings. Save a collection to share your selection of sources.