arXiv · 2609.32991
What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use
Abstract
What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data principle connects them: localize the missing operation, preserve its causal relation, vary shortcut-bearing context, and re-audit the residual. Formation-sensitive selection and prerequisite ordering accelerate a binding-matching-transport path; a brief early prefix from the same training multiset retains a validation advantage through 100B tokens. Availability counterfactuals then distinguish writing content from invoking available memory, while paired supervision and context-opportunity ranking improve matched route decisions and long-context answer likelihood. A continuous 350M-model experiment connects the three interventions on the same facts: early circuit training improves subsequent learning, and the complete sequence outperforms stage-replacement controls on facts withheld from Use teaching. Independent query surfaces and opposed-source decisions expose conditional arbitration as the remaining frontier. Together, these results show why a change in the limiting operation calls for a change in supervision, not merely a new ranking of difficult examples.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yixiao Chen, Ke Cheng, Jiangtao Guan, Shuo Huang, Yue Liu, Jun Zhang, Yuhong Liu, Jie Jiang. 2026-09-26. What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use. https://arxiv.org/abs/2609.32991
Cite the original work for its findings. Save a collection to share your selection of sources.