arXiv Science⌕ Search

arXiv · 2609.35719

AI Agent Swarms as Researchers: Progress, Challenges, and Open Questions

Abstract

Artificial intelligence (AI) agents, language models connected to tools and run in a loop, can now carry out long, multi-step tasks with little supervision. We gave swarms of off-the-shelf coding agents a short statement of scope, from a narrow topic to a whole field, access to the literature and to computing tools, and one standing instruction: make real, correct, useful progress, and do not stop. We supplied no scientific ideas. Within weeks, the agents produced a large body of research notes, paper-length drafts, and formal proofs in five areas of optimization theory and physical science, and proposed untested laboratory experiments in a sixth. We do not claim that all of it is correct or new, but it is not noise: in what we have checked so far, we found no major scientific error, and several results are proved in a proof assistant. The agents produced results faster than we could review them; we estimate that a full review would take us months. Together with two widely discussed 2026 results in mathematics obtained with swarms, our runs suggest that agents can already do a large part of routine theoretical research, at least in areas that we experimented with. This raises questions we cannot yet answer: how to trust results when review, not production, is the scarce resource; what credit and publication counts mean when the human input is a prompt, and why institutions would pay researchers rather than buy computing time; and how people can learn a field, add to what agents do, and stay in control of research they cannot keep up with. Research institutions are not ready: models improve faster than institutions change, so they should decide now how to respond as capabilities increase. We offer tentative positions, release the agents' unedited output as of 25 September 2026, and invite readers to repeat the experiment in their own fields.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sergey Gusev, David E. Bernal Neira. 2026-09-28. AI Agent Swarms as Researchers: Progress, Challenges, and Open Questions. https://arxiv.org/abs/2609.35719

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Meme Template Identification in the Wild: Comparing Methods for Semi-Open-Set Recognition

Image-with-text memes are a dominant form of online communication, and much of their spread happens through meme templates which are recurring visual formats that users adapt with new text or imagery. Most prior work scores memes individually for engagement or harmful content, an approach that is structurally blind to templates. Templates might amplify these patterns and enable coordinated harassment, which becomes visible only when memes are grouped by shared template. We formalize meme template identification as a semi-open-set recognition problem, requiring methods to both classify known templates and reject template-free memes and non-meme content. We introduce an evaluation framework spanning a controlled setting (1,704 ImgFlip templates) and a heterogeneous real-world social media sample, comparing supervised CNN- and distance-based methods, unsupervised density-based clustering, a novel fused SigLIP2+DINOv2 representation, and a retrieval-augmented LLM pipeline. In real-world social media sources dedicated to meme sharing, we find that only 21% of images use a known template, while 58% are template-free and 21% are not memes at all; identification performance also drops sharply from the controlled setting (best MCC 0.974 to 0.684). The loss comes mainly from deciding whether an image is a template-based meme at all, rather than identifying the template. Our dataset of 1.1M soft-labeled social media images, together with code for reproducibility, is available upon request.

cs.CY↗

(Mis-)Informed Consent: Predatory Apps and the Exploitation of Populations with Limited Literacy

Among populations with limited literacy in emerging digital markets, the adoption of mobile phones, combined with comprehension barriers and poor cybersecurity hygiene, has created hidden privacy risks. This paper examines how informed consent is often abused by predatory financial applications, leading to financial scams that disproportionately affect users with low literacy. We focus on predatory loan, gambling, and trading apps, analyzing a dataset of 50 Google Play Store apps to measure how many omit or obfuscate critical privacy disclosures. We also evaluate comprehension gaps among users with low literacy via a targeted user study and assess whether Large Language Model (LLM)-generated summaries, translations, and visual cues can improve consent clarity. Our findings show that 85% of study participants did not understand basic app permissions, underscoring the urgent need for stronger regulatory oversight and scalable LLM-driven privacy-literacy tools.

cs.CY↗

Validated Behavioral Hypotheses as a Lens for Evaluating Participant Simulation

We propose using validated behavioral hypotheses as a lens for evaluating LLM agents as simulated human participants. This approach makes behavioral agreement in participant simulation measurable and decomposable, revealing which human effects agents reproduce, where they diverge, and how agent design changes that agreement. To operationalize this idea, we build HumanStudy-Bench, an open benchmarking platform that reconstructs experimental protocols from published human studies and administers these protocols to agents serving as silicon participants. The benchmark compares population-level effects derived from agent responses with the corresponding published human findings using two metrics: the Probability Alignment Score (PAS) for inferential agreement and the Effect Consistency Score (ECS) for effect-magnitude agreement. We apply HumanStudy-Bench across 12 human studies, evaluating 10 models under four agent designs, with 6,588 simulated participants per agent configuration. By making published human studies and their validated behavioral hypotheses reusable for evaluating simulated participants, we hope to make the behavioral assumptions underlying LLM-based social simulations and their sensitivity to agent design explicit and empirically testable.

cs.CY↗