arXiv · 2610.08208
STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty
Abstract
We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nina Nusbaumer, Iria de-Dios-Flores, Corentin Bel, Christophe Pallier, Guillaume Wisniewski, Benoît Crabbé. 2026-10-06. STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty. https://arxiv.org/abs/2610.08208
Cite the original work for its findings. Save a collection to share your selection of sources.