arXiv ScienceSearch

arXiv subjects

Boris Bolliet

Publications and source records attributed to Boris Bolliet.

At least 19 recordsLinked to original sources

AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions

Agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span literature synthesis, code generation, data analysis, hypothesis proposal, and model criticism. We argue that this transition is qualitative rather than incremental, and that suitably designed multi-agent systems may evolve from passive computational tools into ``AI scientists'' that can expand the hypothesis-generating and verification capacity of science. Such systems must be developed and deployed within a scientific ecosystem fit for purpose: institutions must be redesigned for verification, accountability, interpretability, and dual-use safety. We sketch how multi-agent architectures, illustrated by the prototype framework \textit{Denario}, accelerate the discovery cycle and traverse model spaces beyond human reach; examine what this implies for authorship, peer review, and the enduring role of human scientists; and close with recommendations for governing AI as an epistemic actor rather than a mere instrument.

cs.AI

Impact of non-Gaussian likelihood on cosmological constraints from the thermal Sunyaev--Zel'dovich power spectrum: a simulation-based inference analysis

The thermal Sunyaev--Zel'dovich (tSZ) power spectrum is a sensitive probe of cosmology and cluster astrophysics, but its statistics are non-Gaussian because the signal receives a significant contribution from rare, massive, low-redshift galaxy clusters. As a result, a Gaussian likelihood fails to describe the statistics of its power spectrum on large scales. We use simulation-based inference (SBI) to test the accuracy of the standard Gaussian power-spectrum likelihood for a \textit{Planck}-like tSZ analysis. Using halo-based simulations of full-sky Compton-$y$ maps, we train neural posterior and likelihood estimators and compare the resulting constraints with those from a Gaussian likelihood assumption. Using only multipoles $\ell < 1000$, we find that the Gaussian likelihood assumption gives unbiased cosmological constraints, while the SBI-based inference shows a mild broadening of the posterior distributions for the amplitudes of residual foregrounds. This suggests that the Gaussian likelihood assumption is sufficiently accurate for cosmological inference for a \textit{Planck}-like tSZ analysis, while SBI provides a useful validation tool to model non-Gaussian likelihoods beyond analytic approximations.

astro-ph.CO

Competing with AI Scientists: Agent-Driven Approach to Astrophysics Research

We present an agent-driven approach to the construction of parameter inference pipelines for scientific data analysis. Our method leverages a multi-agent system, Cmbagent (the analysis system of the AI scientist Denario), in which specialized agents collaborate to generate research ideas, write and execute code, evaluate results, and iteratively refine the overall pipeline. As a case study, we apply this approach to the FAIR Universe Weak Lensing Uncertainty Challenge, a competition under time constraints focused on robust cosmological parameter inference with realistic observational uncertainties. While the fully autonomous exploration initially did not reach expert-level performance, the integration of human intervention enabled our agent-driven workflow to achieve a first-place result in the challenge. This demonstrates that semi-autonomous agentic systems can compete with, and in some cases surpass, expert solutions. We describe our workflow in detail, including both the autonomous and semi-autonomous exploration by Cmbagent. Our final inference pipeline utilizes parameter-efficient convolutional neural networks, likelihood calibration over a known parameter grid, and multiple regularization techniques. Our results suggest that agent-driven research workflows can provide a scalable framework to rapidly explore and construct pipelines for inference problems.

cs.AI

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

astro-ph.IM

Self-consistent secondary cosmic microwave background anisotropies and extragalactic foregrounds in the FLAMINGO simulations

Secondary anisotropies in the cosmic microwave background (CMB) contain information that can be used to test both cosmological models and models of galaxy formation. Starting from lightcone-based HEALPix maps and catalogues, we present a new set of mock CMB maps constructed in a self-consistent manner from the FLAMINGO suite of cosmological hydrodynamical simulations, including CMB lensing, thermal and kinetic Sunyaev-Zeldovich effects, cosmic infrared background, radio point source and anisotropic screening maps. We show that these simulations reproduce a wide range of observational constraints. We also compare our simulations with previous predictions based on dark matter-only simulations which generally model the secondary anisotropies independently from one another, concluding that our hydrodynamical simulation mocks perform at least as well as previous mocks in matching the observations whilst retaining self-consistency in the predictions of the different components. Using the model variations in FLAMINGO, we further explore how the signals depend on cosmology and feedback modelling, and we predict cross-correlations between some of the signals that differ significantly from those in previous mocks. The mock CMB maps should provide a valuable resource for exploring correlations between different secondary anisotropies and other large-scale structure tracers, and can be applied to forecasts for upcoming surveys.

astro-ph.CO

Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities

We show that multi-agent systems guided by vision-language models (VLMs) improve end-to-end autonomous scientific discovery. By treating plots as verifiable checkpoints, a VLM-as-a-judge evaluates figures against dynamically generated domain-specific rubrics, enabling agents to correct their own errors and steer exploratory data analysis in real-time. Case studies in cosmology and astrochemistry demonstrate recovery from faulty reasoning paths and adaptation to new datasets without human intervention. On a 10-task benchmark for data-driven discovery, VLM-augmented systems achieve pass at 1 scores of 0.7-0.8, compared to 0.2-0.3 for code-only and 0.4-0.5 for code-and-text baselines, while also providing auditable reasoning traces that improve interpretability. Code available here: https://github.com/CMBAgents/cmbagent

cs.CL

The Denario project: Deep knowledge AI agents for scientific discovery

We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific analysis using Cmbagent as a deep-research backend. In this work, we describe in detail Denario and its modules, and illustrate its capabilities by presenting multiple AI-generated papers generated by it in many different scientific disciplines such as astrophysics, biology, biophysics, biomedical informatics, chemistry, material science, mathematical physics, medicine, neuroscience and planetary science. Denario also excels at combining ideas from different disciplines, and we illustrate this by showing a paper that applies methods from quantum physics and machine learning to astrophysical data. We report the evaluations performed on these papers by domain experts, who provided both numerical scores and review-like feedback. We then highlight the strengths, weaknesses, and limitations of the current system. Finally, we discuss the ethical implications of AI-driven research and reflect on how such technology relates to the philosophy of science. We publicly release the code at https://github.com/AstroPilot-AI/Denario. A Denario demo can also be run directly on the web at https://huggingface.co/spaces/astropilot-ai/Denario, and the full app will be deployed on the cloud.

cs.AI

CMB component-separated power spectrum estimation by Spectral Internal Linear Combination (SpILC)

Component separation methods mitigate the cross-contamination between different extragalactic and galactic contributions to cosmic microwave background (CMB) data. This is often done by linearly combining CMB maps from different frequency channels using internal linear combination (ILC) methods. We demonstrate that deriving power spectrum estimators directly by linearly combining auto- and cross-spectra instead of maps allows us to obtain a different constrained-optimization problem that allows fewer (deprojection) constraint equations than combining at map level using the constrained ILC method. Through simulations, we show that our Spectral internal linear combination (SpILC) produces CMB power spectrum estimators with more than 7 times smaller errorbars than constrained ILC (with thermal Sunyaev-Zel'dovich and cosmic infrared background deprojections) at $\ell\gtrsim 4000$ for Simons Observatory-like observations. Spectral ILC outperforms constrained ILC methods when some modeled components are spatially uncorrelated, e.g. the primary CMB is uncorrelated with foregrounds, and the difference in performance is most significant at noise-dominated scales. More generally, our work shows that component-separated maps with foreground deprojections do not necessarily produce minimum-variance two-or-higher-point estimators.

astro-ph.CO

CLAPP: The CLASS LLM Agent for Pair Programming

We introduce CLAPP (CLASS LLM Agent for Pair Programming), an interactive AI assistant designed to support researchers working with the Einstein-Boltzmann solver CLASS. CLAPP leverages large language models (LLMs) and domain-specific retrieval to provide conversational coding support for CLASS-answering questions, generating code, debugging errors, and producing plots. Its architecture combines multi-agent LLM orchestration, semantic search across CLASS documentation, and a live Python execution environment. Deployed as a user-friendly web application, CLAPP lowers the entry barrier for scientists unfamiliar with AI tools and enables more productive human-AI collaboration in computational and numerical cosmology. The app is available at https://classclapp.streamlit.app

astro-ph.IM

The Atacama Cosmology Telescope: High-redshift measurement of structure growth from the cross-correlation of Quaia quasars and CMB lensing from ACT DR6 and $\textit{Planck}$ PR4

We measure the amplitude of matter fluctuations over a wide range of redshifts by combining CMB lensing observations from ACT DR6 and $\textit{Planck}$ PR4 with the overdensity of quasars from Quaia, a $\textit{Gaia}$ and $\textit{unWISE}$ quasar catalog. Our analysis includes the CMB lensing power spectrum from ACT DR6, the auto-correlation of two Quaia quasar samples centered at $z \simeq 1.0$ and $z \simeq 2.1$, and their cross-correlations with CMB lensing from both ACT DR6 and $\textit{Planck}$ PR4. By performing a series of contamination and systematic null tests, we find no evidence for contamination in the lensing maps, contrary to what was suggested in previous Quaia cross-correlation analyses using $\textit{Planck}$ PR4 CMB lensing data. From the joint analysis of the quasar auto- and cross-correlations with CMB lensing, and including BOSS BAO data to break the degeneracy between $\Omega_m$ and $\sigma_8$, we obtain $\sigma_8 = 0.802^{+0.045}_{-0.057}$, consistent with $\Lambda$CDM predictions from $\textit{Planck}$ primary CMB measurements. Combining the CMB lensing auto-spectrum with the cross-correlation measurement improves the constraint on $\sigma_8$ by $12\%$ relative to the lensing auto-spectrum alone, yielding $\sigma_8 = 0.804 \pm 0.013$. This dataset combination also enables a reconstruction of structure growth across redshifts. We infer a $12\%$ constraint on the amplitude of matter fluctuations at $z > 3$, with a measurement at the median redshift of the signal of $\sigma_8(\tilde{z}=5.1) = 0.146^{+0.021}_{-0.014}$, consistent with $\textit{Planck}$ at the $1.4\sigma$ level. These results provide one of the highest redshift constraints on the growth of structure to date.

astro-ph.CO

CLASS_SZ II: Notes and Examples of Fast and Accurate Calculations of Halo Model, Large Scale Structure and Cosmic Microwave Background Observables

These notes are very much work-in-progress and simply intended to showcase, in various degrees of details (and rigour), some of the cosmology calculations that class_sz can do. We describe the class_sz code in C, Python and Jax. Based on the Boltzmann code class, it can compute a wide range of observables relevant to current and forthcoming CMB and Large Scale Structure surveys. This includes galaxy shear and clustering, CMB lensing, thermal and kinetic Sunyaev and Zeldovich observables, Cosmic Infrared Background, cross-correlations and three-point statistics. Calculations can be done either within the halo model or the linear bias model. For standard $\Lambda$CDM cosmology and extensions, class_sz uses high-accuracy cosmopower emulators of the CMB and matter power spectrum to accelerate calculations. With this, along with efficient numerical integration routines, most class_sz output can be obtained in less than 500 ms (CMB $C_\ell$'s or matter $P(k)$ take $\mathcal{O}(1\mathrm{ms})$), allowing for fast or ultra-fast parameter inference analyses. Parts of the calculations are "jaxified", so the software can be integrated into differentiable pipelines.

astro-ph.CO

Evaluating Retrieval-Augmented Generation Agents for Autonomous Scientific Discovery in Astrophysics

We evaluate 9 Retrieval Augmented Generation (RAG) agent configurations on 105 Cosmology Question-Answer (QA) pairs that we built specifically for this purpose.The RAG configurations are manually evaluated by a human expert, that is, a total of 945 generated answers were assessed. We find that currently the best RAG agent configuration is with OpenAI embedding and generative model, yielding 91.4\% accuracy. Using our human evaluation results we calibrate LLM-as-a-Judge (LLMaaJ) system which can be used as a robust proxy for human evaluation. These results allow us to systematically select the best RAG agent configuration for multi-agent system for autonomous scientific discovery in astrophysics (e.g., cmbagent presented in a companion paper) and provide us with an LLMaaJ system that can be scaled to thousands of cosmology QA pairs. We make our QA dataset, human evaluation results, RAG pipelines, and LLMaaJ system publicly available for further use by the astrophysics community.

astro-ph.IM

Open Source Planning & Control System with Language Agents for Autonomous Scientific Discovery

We present a multi-agent system for automation of scientific research tasks, cmbagent (https://github.com/CMBAgents/cmbagent). The system is formed by about 30 Large Language Model (LLM) agents and implements a Planning & Control strategy to orchestrate the agentic workflow, with no human-in-the-loop at any point. Each agent specializes in a different task (performing retrieval on scientific papers and codebases, writing code, interpreting results, critiquing the output of other agents) and the system is able to execute code locally. We successfully apply cmbagent to carry out a PhD level cosmology task (the measurement of cosmological parameters using supernova data) and evaluate its performance on two benchmark sets, finding superior performance over state-of-the-art LLMs. The source code is available on GitHub, demonstration videos are also available, and the system is deployed on HuggingFace and will be available on the cloud.

cs.AI

The Atacama Cosmology Telescope: DR6 Power Spectrum Foreground Model and Validation

We discuss the model of astrophysical emission at millimeter wavelengths used to characterize foregrounds in the multi-frequency power spectra of the Atacama Cosmology Telescope (ACT) Data Release 6 (DR6), expanding on Louis et al. (2025). We detail several tests to validate the capability of the DR6 parametric foreground model to describe current observations and complex simulations, and show that cosmological parameter constraints are robust against model extensions and variations. We demonstrate consistency of the model with pre-DR6 ACT data and observations from Planck and the South Pole Telescope. We evaluate the implications of using different foreground templates and extending the model with new components and/or free parameters. In all scenarios, the DR6 $\Lambda$CDM and $\Lambda$CDM+$N_{\rm eff}$ cosmological parameters shift by less than $0.5\sigma$ relative to the baseline constraints. Some foreground parameters shift more; we estimate their systematic uncertainties associated with modeling choices. From our constraint on the kinematic Sunyaev-Zel'dovich power, we obtain a conservative limit on the duration of reionization of $\Delta z_{\rm rei} < 4.4$, assuming a reionization midpoint consistent with optical depth measurements and a minimal low-redshift contribution, with varying assumptions for this component leading to tighter limits. Finally, we analyze realistic non-Gaussian, correlated microwave sky simulations containing Galactic and extragalactic foreground fields, built independently of the DR6 parametric foreground model. Processing these simulations through the DR6 power spectrum and likelihood pipeline, we recover the input cosmological parameters of the underlying cosmic microwave background field, a new demonstration for small-scale CMB analysis. These tests validate the robustness of the ACT DR6 foreground model and cosmological parameter constraints.

astro-ph.CO

Unified and consistent structure growth measurements from joint ACT, SPT and \textit{Planck} CMB lensing

We present the tightest cosmic microwave background (CMB) lensing constraints to date on the growth of structure by combining CMB lensing measurements from the Atacama Cosmology Telescope (ACT), the South Pole Telescope (SPT) and \textit{Planck}. Each of these surveys individually provides lensing measurements with similarly high statistical power, achieving signal-to-noise ratios of approximately 40. The combined lensing bandpowers represent the most precise CMB lensing power spectrum measurement to date with a signal-to-noise ratio of 61 and an amplitude of $A_\mathrm{lens}^\mathrm{recon} = 1.025 \pm 0.017$ with respect to the theory prediction from the best-fit CMB \textit{Planck}-ACT cosmology. The bandpowers from all three lensing datasets, analyzed jointly, yield a $1.6\%$ measurement of the parameter combination $S_8^\mathrm{CMBL} \equiv \sigma_8\,(\Omega_m/0.3)^{0.25} = 0.825^{+0.015}_{-0.013}$. Including Dark Energy Spectroscopic Instrument (DESI) Baryon Acoustic Oscillation (BAO) data improves the constraint on the amplitude of matter fluctuations to $\sigma_8 = 0.829 \pm 0.009$ (a $1.1\%$ determination). When combining with uncalibrated supernovae from \texttt{Pantheon+}, we present a $4\%$ sound-horizon-independent estimate of $H_0=66.4\pm2.5\,\mathrm{km\,s^{-1}\,Mpc^{-1}} $. The joint lensing constraints on structure growth and present-day Hubble rate are fully consistent with a $\Lambda$CDM model fit to the primary CMB data from \textit{Planck} and ACT. While the precise upper limit is sensitive to the choice of data and underlying model assumptions, when varying the neutrino mass sum within the $\Lambda\mathrm{CDM}$ cosmological model, the combination of primary CMB, BAO and CMB lensing drives the probable upper limit for the mass sum towards lower values, comparable to the minimum mass prior required by neutrino oscillation experiments.

astro-ph.CO

Extracting cosmological information from the abundance of galaxy clusters with simulation-based inference

The abundance of galaxy clusters as a function of mass and redshift is a well-established and powerful cosmological probe. Cosmological analyses based on galaxy cluster number counts have traditionally relied on explicitly computed likelihoods, which are often challenging to develop with the required accuracy and expensive to evaluate. In this work, we implement an alternative approach based on simulation-based inference (SBI) methods that relies solely on synthetic galaxy cluster catalogues generated under a given model. These catalogues are much easier to produce than it is to develop and validate a likelihood. We validate this approach in the context of the galaxy cluster survey of the upcoming Simons Observatory for a setup in which we can also evaluate an exact explicit likelihood. We find that our SBI-based approach yields cosmological parameter posterior means that are within $0.2\,\sigma$ of those obtained with the explicit likelihood and with biases smaller than $0.1\,\sigma$. We also introduce and validate a procedure to assess the goodness of fit using only synthetic catalogues similar to those used for training. This demonstrates, for the first time, that a galaxy cluster number count cosmological analysis can be performed fully without resorting to a likelihood at any stage. Finally, we apply our SBI-based approach to the real Planck MMF3 cosmology sample, obtaining cosmological parameter constraints that are within $0.1\,\sigma$ of their likelihood-based counterparts. This constitutes the first SBI-based number count cosmological analysis of a real galaxy cluster catalogue.

astro-ph.CO

The Atacama Cosmology Telescope: DR6 Maps

We present Atacama Cosmology Telescope (ACT) Data Release 6 (DR6) maps of the Cosmic Microwave Background temperature and polarization anisotropy at arcminute resolution over three frequency bands centered on 98, 150 and 220 GHz. The maps are based on data collected with the AdvancedACT camera over the period 2017--2022 and cover 19,000 square degrees with a median combined depth of 10 uK arcmin. We describe the instrument, mapmaking and map properties and illustrate them with a number of figures and tables. The ACT DR6 maps and derived products are available on LAMBDA at https://lambda.gsfc.nasa.gov/product/act/actadv_prod_table.html. We also provide an interactive web atlas at https://phy-act1.princeton.edu/public/snaess/actpol/dr6/atlas and HiPS data sets in Aladin (e.g. https://alasky.cds.unistra.fr/ACT/DR4DR6/color_CMB).

astro-ph.CO

The Atacama Cosmology Telescope: DR6 Power Spectra, Likelihoods and $\Lambda$CDM Parameters

We present power spectra of the cosmic microwave background (CMB) anisotropy in temperature and polarization, measured from the Data Release 6 maps made from Atacama Cosmology Telescope (ACT) data. These cover 19,000 deg$^2$ of sky in bands centered at 98, 150 and 220 GHz, with white noise levels three times lower than Planck in polarization. We find that the ACT angular power spectra estimated over 10,000 deg$^2$, and measured to arcminute scales in TT, TE and EE, are well fit by the sum of CMB and foregrounds, where the CMB spectra are described by the $\Lambda$CDM model. Combining ACT with larger-scale Planck data, the joint P-ACT dataset provides tight limits on the ingredients, expansion rate, and initial conditions of the universe. We find similar constraining power, and consistent results, from either the Planck power spectra or from ACT combined with WMAP data, as well as from either temperature or polarization in the joint P-ACT dataset. When combined with CMB lensing from ACT and Planck, and baryon acoustic oscillation data from DESI DR1, we measure a baryon density of $\Omega_b h^2=0.0226\pm0.0001$, a cold dark matter density of $\Omega_c h^2=0.118\pm0.001$, a Hubble constant of $H_0=68.22\pm0.36$ km/s/Mpc, a spectral index of $n_s=0.974\pm0.003$, and an amplitude of density fluctuations of $\sigma_8=0.813\pm0.005$. Including the DESI DR2 data tightens the Hubble constant to $H_0=68.43\pm0.27$ km/s/Mpc; $\Lambda$CDM parameters agree between the P-ACT and DESI DR2 data at the $1.6\sigma$ level. We find no evidence for excess lensing in the power spectrum, and no departure from spatial flatness. The contribution from Sunyaev-Zel'dovich (SZ) anisotropy is detected at high significance; we find evidence for a tilt with suppressed small-scale power compared to our baseline SZ template spectrum, consistent with hydrodynamical simulations with feedback.

astro-ph.CO