arXiv · 2609.26539
A retrospective analysis on the use of LLMs to study infant syntax learning
Abstract
Large language models (LLMs) have increasingly been used to investigate how children acquire syntax at an early stage of development. This is notably the central scientific goal of the BabyLM challenge, a community-wide effort to develop models that achieve human-level syntactic performance while being trained on developmentally realistic corpora. In this paper, we reflect on the use of LLMs in the study of infant syntax learning by providing an epistemological assessment of several studies from this research program. We discuss how datasets are built, which models are implemented, how they are trained and syntactically evaluated. We observe significant assumptions in the methodology of BabyLM and related studies, thus mitigating their theoretical scope. We additionally observe that using developmentally-realistic corpora have limited effects on models performance on commonly-used benchmarks, which suggest important computational differences between LLMs and the infant syntax learner.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hélie Bazin, Anouk Barberousse, François Yvon. 2026-09-22. A retrospective analysis on the use of LLMs to study infant syntax learning. https://arxiv.org/abs/2609.26539
Cite the original work for its findings. Save a collection to share your selection of sources.