AI-Research Agents in the Wild. From GitHub and arXiv to Regularities and Gaps
AI-research agents, or autoresearch systems, combine language models with tools, search, evaluation, and iterative modification of research artifacts. Their public software ecology is hard to compare because repositories, papers, benchmarks, libraries, and companion artifacts are often counted as one population. We connect two agent-drafted, human-adjudicated registries frozen on 10 June 2026: 139 canonical public repositories and 101 papers on AI-assisted research systems, each record carrying an evidence card anchored to its primary sources. Nine promoted design lineages contain 59 memberships among 52 repositories, seven of them assigned to two lineages. Against this evidence we examine six time-stamped candidate design regularities. A prospective audit of 25 newly ingested repositories observed zero of six specified trigger events, supplying bounded support for R1-R4; broader corpus evidence narrowed the original scope of R5, and a failed directional prediction leaves R6 exploratory. The paper-repository graph contains 212 curator-assigned relatedness edges: 64 of 101 papers carry at least one edge, the edges reach 23 of 139 repositories, and the six most-linked repositories carry 120 of the 212 edges (56.6%). A full-text check of seven papers shows the relation measures our curation, not the papers' citation behavior. An identity audit adjudicated against commit histories and public contributor records confirms at least 18 paper authors who also author a canonical registry repository (20 of 63 candidate author slots in 13 of 101 papers); a byline-matched tier reaches 58 identities but is reported separately and more weakly. The study contributes a source-grounded theory under prospective test, a curator-assigned map of paper-repository relatedness, and explicit limits on claims about public visibility and community structure.