arXiv Science⌕ Search

arXiv · 2610.10468

A Society of Researchers: Designing Institutions for Populations of Autonomous Research Agents

Abstract

Deployments of research agents are moving to populations of thousands that share one pool of compute, while most current systems organize one project at a time or leave the population unorganized. We argue that such a population will acquire an organization whether or not its designers provide one, so designers should provide it explicitly, and that the multi-agent systems community holds the tools to do so. We propose a society of agents, a population of persistent agents under explicit institutions, and develop it for science as a society of researchers built on six principles. Principal investigators compete for compute through requests for proposals, independent review, and grants; a human governor, the mayor, allocates resources and assigns no tasks. In a running society of ten thousand researchers, asked only to improve the pretraining of language models, one lab reported a way to reach the same quality with about 30% less compute, a result the labs that tested it do not yet agree on. We close with six open problems for the agents community.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ali Asaria, Deep Gandhi, Tony Salomone. 2026-10-07. A Society of Researchers: Designing Institutions for Populations of Autonomous Research Agents. https://arxiv.org/abs/2610.10468

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Towards Strategy-Level RSI for Skill-Augmented Agents: Learning When to Reuse Skills from Execution Feedback

Long-running agents accumulate reusable Skills, but a Skill that is semantically relevant to a task is not necessarily worth loading in the current state. We study the applicability question that arises once a candidate Skill is known: should it be loaded in the current state? We propose SkillApt, which uses matched WITH/WITHOUT Skill executions on the same task state as persistent evidence, estimates the conditional marginal utility of the Skill, and chooses LOAD or ABSTAIN accordingly. The base model, agent architecture, and Skill contents stay fixed; only the external deployment policy changes. We call this constrained setting strategy-level recursive self-improvement (Strategy-Level RSI). On 20 Skills and 160 held-out states, as paired evidence accumulates, SkillApt's task success rises from 81.9% under a cold start to 91.3%, matching a strong zero-shot LLM controller; yet SkillApt activates Skills on only 26.3% of states, versus 98.8% for the zero-shot controller. A hard-candidate study shows that non-optimal Skills mostly leave correctness unchanged while raising execution cost, and occasionally cause correctness harm. An ablation shows that a history recording only WITH success makes the policy load almost everywhere, whereas paired evidence substantially improves selectivity. These results indicate that relevance is not applicability: the main effect of execution evidence is not to make the model stronger but to change how existing Skills are deployed, moving the system from near-always loading to selective reuse.

cs.MA↗

Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions

Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study ($n{=}5$), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing

cs.MA↗

ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration

Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.

cs.MA↗