arXiv ScienceSearch

arXiv subjects

Francisco Villaescusa-Navarro

Publications and source records attributed to Francisco Villaescusa-Navarro.

At least 19 recordsLinked to original sources

Inductive Biases in Field-Level Cosmological Inference from Galaxy Catalogs

We perform field-level likelihood-free inference of the matter density parameter $\Omega_m$ from simulated galaxy catalogs using machine learning models with differing inductive biases. Using hydrodynamic simulations from CAMELS, we examine how observable choice and architecture govern cosmological information extraction. We consider galaxy positions and line-of-sight peculiar velocities, separately and jointly, and compare permutation-invariant Deep Sets, implemented with either multilayer perceptrons (MLPs) or Kolmogorov-Arnold Networks (KANs), to graph neural networks (GNNs), which explicitly encode spatial relations. We test in-distribution and out-of-distribution (OOD) performance across simulations with different subgrid galaxy-formation prescriptions. Deep Sets infer $\Omega_m$ from velocities alone with mean relative errors of approximately $18\%$ in-distribution and $\sim25\%$ OOD, with KANs and MLPs achieving comparable performance. In contrast, the same set-based approach does not yield useful $\sigma_8$ predictions in either in-distribution or cross-suite tests. Adding positions does not improve Deep Sets, while GNNs infer $\Omega_m$ with mean relative errors of about $10\%$ in-distribution and $10$--$17\%$ OOD. These results indicate that peculiar velocities provide the dominant source of $\Omega_m$ information for set-based models in this setting, while spatial information is most effectively used by architectures that explicitly encode galaxy-galaxy relations. Because the velocity inputs are exact simulated peculiar velocities, applications to survey data will require validation under realistic velocity-measurement noise, selection effects, and survey geometry.

astro-ph.CO

Deep Learning for Astrophysics: An Open Textbook from the NASA Cosmic Origins AI/ML Science and Technology Interest Group

Recent community assessments identify education as a principal barrier to adopting modern machine learning in astronomy. We present Deep Learning for Astrophysics, a freely available textbook at https://deeplearning4astro.com, curated from the NASA Cosmic Origins Artificial Intelligence and Machine Learning Science and Technology Interest Group (AI/ML STIG) lecture series. The book collects 23 chapters by 17 lecturers across six parts, moving from computational foundations and deep-learning architectures through generative modeling, simulation-based inference, reinforcement learning, and large-language-model agents to the practice of AI-laden science. Many include executable notebooks using astronomical data.

astro-ph.IM

AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions

Agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span literature synthesis, code generation, data analysis, hypothesis proposal, and model criticism. We argue that this transition is qualitative rather than incremental, and that suitably designed multi-agent systems may evolve from passive computational tools into ``AI scientists'' that can expand the hypothesis-generating and verification capacity of science. Such systems must be developed and deployed within a scientific ecosystem fit for purpose: institutions must be redesigned for verification, accountability, interpretability, and dual-use safety. We sketch how multi-agent architectures, illustrated by the prototype framework \textit{Denario}, accelerate the discovery cycle and traverse model spaces beyond human reach; examine what this implies for authorship, peer review, and the enduring role of human scientists; and close with recommendations for governing AI as an epistemic actor rather than a mere instrument.

cs.AI

Learning the Universe with the 2nd Generation of CAMELS: Varying 35 parameters of the IllustrisTNG model in (50Mpc/h)^3 boxes

We present a new set of 1,192 cosmological simulations as part of the CAMELS project, in which a space of 35 cosmological, astrophysical, and numerical parameters is explored around the fiducial IllustrisTNG model. The volume of each of these simulations is (50Mpc/h)^3, eight times larger than that of previous CAMELS simulations. This provides lower sample variance as well as access to more massive halos and more diverse environments. We focus this work on exploring the advantages these differences provide for parameter inference powered by neural networks. We generate training sets based on the matter power spectra, projected maps of the volumes, graphs representing galaxy spatial distributions, and thermodynamical properties of massive halos. We employ multilayer perceptrons, convolutional neural networks, graph neural networks, and Gaussian processes, respectively, to extract information on the simulation parameters from these inputs while comparing systematically to analogous results from our previous generation of (25Mpc/h)^3 simulations. We generally find that the new, larger volumes produce tighter marginal constraints on the parameters, to degrees that vary between the different inputs. The improvements, however, scale more weakly than with the square root of the increase in the amount of data (i.e., physical volume). We interpret this as originating either from information loss due to mode coupling or from complex degeneracies in parameter space. We also discuss the effects on statistics of the intergalactic medium temperature from four new parameters that are varied in these simulations, which control the amplitude and timing of the ionizing background radiation. We publicly release the simulation outputs and ancillary data at https://camels.readthedocs.io.

astro-ph.CO

Reconciling the Fundamental Plane of Early-Type Galaxies with hydrodynamical simulations: The case of IllustrisTNG100-1

The Fundamental Plane (FP) of Early-Type Galaxies (ETGs) encapsulates a tight correlation among their structural and dynamical properties and provides an important benchmark for galaxy formation models. However, cosmological hydrodynamical simulations have historically struggled to reproduce the observed FP tilt, with discrepancies often attributed to to flawed feedback physics or insufficient resolution. Using the IllustrisTNG100-1 simulation, we show that adopting observationally motivated measurements, including S\'ersic-derived photometric parameters and dynamically inferred velocity dispersions designed to minimise softening-length effects, substantially reduces the discrepancy between simulated and observed FPs. We further explore the impact of non-universal, mass-dependent Initial Mass Function (IMF) variations through forward modelling of their effects on galaxy structural and dynamical quantities. In particular, bottom-heavy IMF variations produce FP coefficients fully consistent with observational constraints for both direct and orthogonal fits. Our results suggest that a significant fraction of the long-standing FP tension arises from how galaxy observables are extracted and interpreted in simulations, although residual discrepancies may still reflect limitations in the underlying baryonic physics. These findings highlight the importance of observational realism and IMF variations for interpreting galaxy scaling relations and for improving the predictive power of hydrodynamical simulations of ETG formation.

astro-ph.GA

Efficiently emulating distribution functions in gigaparsec volumes for varying cosmological parameters

We present a new method for emulating the halo mass function (HMF) and other distribution functions in large effective volumes, down to low halo masses, whilst simultaneously modifying large ranges of parameters, for a fraction of the cost of traditional periodic cosmological simulations. We demonstrate the method by selecting small regions, $V \sim (50 \,h^{-1}{\rm Mpc})^3$, with a range of overdensities from the Quijote suite, consisting of tens of thousands of $(1 \,h^{-1}{\rm Gpc})^3$ $N$-body simulation volumes run with varying $\Lambda$CDM parameters. We train a differentiable emulator, conditioned on the overdensity of the region and these global parameters, to reproduce the halo mass function in these regions. We then successfully recover the global distribution of halo masses of the entire box by integrating over the overdensity distribution. Our approach uses just $\sim\,$0.026% of the original simulation volume, and suggests that suites of targeted `zoom' simulations, extracted from low resolution parent volumes, can be used to emulate large volume simulations at a fraction of the computational cost, whilst simultaneously pushing the dynamic range to much lower masses than can be achieved in periodic simulations. We discuss emulation of other key dark matter and baryonic distribution functions, as well as higher order statistics, with implications for the interpretation of upcoming wide field surveys on observatories such as Euclid, Roman and Rubin.

astro-ph.CO

Field-Level Inference of Primordial Non-Gaussianity with the Quijote Simulation Suite

Local primordial non-Gaussianity, parameterised as $f_{\rm NL}^{\rm local}$, will be stringently constrained using state-of-the-art methods applied to next-generation galaxy redshift survey data. In this paper, in preparation for the upcoming data sets, we demonstrate for the first time the joint field-level inference of $f_{\rm NL}^{\rm local}$, nuisance parameters, and the initial conditions in realistic halo catalogues, ones which are generated through full dark-matter-only $N$-body simulations. The field-level inference algorithm optimally constrains $f_{\rm NL}^{\rm local}$ through a Bayesian forward-modelling approach at the field level, which outperforms traditional methods by leveraging the full statistical power of the data at the scales considered. In addition, we assess its performance under various design choices in the forward model, including tests of the structure formation model and resolution. We demonstrate the robustness of our approach by applying it to a subset of the \textit{Quijote} simulation suite, performing the inference at scales down to $k_{\rm max} \approx 0.1 h \rm{Mpc}^{-1}$. Compared with a power spectrum and bispectrum estimator, we find a $\sim1.3$ improvement in $\sigma(f_{\rm NL}^{\rm local})$ when applying \borg{}, while marginalising over the initial conditions and bias parameters. From the small-scale information sensitivity tests, we show that the constraints on $f_{\rm NL}^{\rm local}$ improve as we increase the resolution of the inference. These findings underscore the transformative potential of field-level inference to leverage the information available in ongoing surveys such as \textit{Euclid}, providing accurate insights into the physics of cosmic inflation and the number of fields driving it.

astro-ph.CO

Cosmology with one galaxy: An analytic formula relating $\Omega_{\rm m}$ with galaxy properties

Standard cosmological analyses typically treat galaxy formation and cosmological parameter inference as decoupled problems, relying on population-level statistics such as clustering, lensing, or halo abundances. However, classical studies of baryon fractions in massive galaxy clusters have long suggested that gravitationally bound systems may retain cosmological information through their baryonic content. Building on this insight, we present the first analytic and physically interpretable cosmological tracer that links the matter density parameter, $\Omega_m$, directly to intrinsic galaxy-scale observables, demonstrating that cosmological information can be extracted from individual galaxies. Using symbolic regression applied to state-of-the-art hydrodynamical simulations from the CAMELS project, we identify a compact functional form that robustly recovers $\Omega_m$ across multiple simulation suites (IllustrisTNG, ASTRID, SIMBA, and Swift-EAGLE), requiring only modest recalibration of a small number of coefficients. The resulting expression admits a transparent physical interpretation in terms of baryonic retention and enrichment efficiency regulated by gravitational potential depth, providing a clear explanation for why $\Omega_m$ is locally encoded in galaxy properties. Our work establishes a direct, interpretable bridge between small-scale galaxy physics and large-scale cosmology, opening a complementary pathway to cosmological inference that bypasses traditional clustering-based statistics and enables new synergies between galaxy formation theory and precision cosmology.

astro-ph.CO

Cosmological back-reaction of baryons on dark matter in the CAMELS simulations

Baryonic processes such as radiative cooling and feedback from massive stars and active galactic nuclei (AGN) directly redistribute baryons in the Universe but also indirectly redistribute dark matter due to changes in the gravitational potential. In this work, we investigate this "back-reaction" of baryons on dark matter using thousands of cosmological hydrodynamic simulations from the Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) project, including parameter variations in the SIMBA, IllustrisTNG, ASTRID, and Swift-EAGLE galaxy formation models. Matching haloes to corresponding N-body (dark matter-only) simulations, we find that virial masses decrease owing to the ejection of baryons by feedback. Relative to N-body simulations, halo profiles show an increased dark matter density in the center (due to radiative cooling) and a decrease in density farther out (due to feedback), with both effects being strongest in SIMBA (> 450% increase at r < 0.01 Rvir). The clustering of dark matter strongly responds to changes in baryonic physics, with dark matter power spectra in some simulations from each model showing as much as 20% suppression or increase in power at k ~ 10 h/Mpc relative to N-body simulations. We find that the dark matter back-reaction depends intrinsically on cosmology (Omega_m and sigma_8) at fixed baryonic physics, and varies strongly with the details of the feedback implementation. These results emphasize the need for marginalizing over uncertainties in baryonic physics to extract cosmological information from weak lensing surveys as well as their potential to constrain feedback models in galaxy evolution.

astro-ph.CO

CAMELS Environments: The Impact of Local Neighbours on Galaxy Evolution across the SIMBA, IllustrisTNG, ASTRID, and Swift-EAGLE Simulations

Internal feedback from massive stars and active galactic nuclei (AGN) play a key role in galaxy evolution, but external environmental effects can also strongly influence galaxies. We investigate the impact of environment on galaxy evolution, and its dependence on baryonic physics implementation, using Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) spanning a wide range of stellar and AGN feedback implementations in the SIMBA, IllustrisTNG, ASTRID, and Swift-EAGLE galaxy formation models. We show that satellite galaxies are significantly affected by the environment in all simulation models, with their gas fraction and star formation rate (SFR) suppressed in overdense regions compared to similar mass satellites in underdense environments at $z=0$. Central galaxies are less sensitive to environment but tend to show lower gas fraction and SFR in overdense regions at low stellar mass, transitioning to higher gas fraction and SFR for massive galaxies in higher-density environments. Halo baryon fraction ($f_{\rm B}$) and circumgalactic medium mass fraction ($f_{\rm CGM}$) at $z=0$ show clear environmental effects. In SIMBA, low-mass haloes in overdense regions have systematically lower $f_{\rm B}$ and $f_{\rm CGM}$ at fixed halo mass, while Swift-EAGLE haloes in overdense regions have systematically higher $f_{\rm B}$ and $f_{\rm CGM}$ across the full halo mass range, and IllustrisTNG and ASTRID show opposite trends at the low and high mass ends. Environmental effects can flip at higher redshift, with SFR and $f_{\rm B}$ increasing with local density in low-mass haloes before quenching at an increasing overdensity threshold. Our results demonstrate that the impact of environment on galaxy evolution depends significantly on galaxy formation model, and higher-density environments can either suppress or enhance star formation depending on galaxy mass and cosmic epoch.

astro-ph.GA

Galaxy Phase-Space and Field-Level Cosmology: The Strength of Semi-Analytic Models

Semi-analytic models are a widely used approach to simulate galaxy properties within a cosmological framework, relying on simplified yet physically motivated prescriptions. They have also proven to be an efficient alternative for generating accurate galaxy catalogs, offering a faster and less computationally expensive option compared to full hydrodynamical simulations. In this paper, we demonstrate that using only galaxy $3$D positions and radial velocities, we can train a graph neural network coupled to a moment neural network to obtain a robust machine learning based model capable of estimating the matter density parameters, $\Omega_{\rm m}$, with a precision of approximately 10%. The network is trained on ($25 h^{-1}$Mpc)$^3$ volumes of galaxy catalogs from L-Galaxies and can successfully extrapolate its predictions to other semi-analytic models (GAEA, SC-SAM, and Shark) and, more remarkably, to hydrodynamical simulations (Astrid, SIMBA, IllustrisTNG, and SWIFT-EAGLE). Our results show that the network is robust to variations in astrophysical and subgrid physics, cosmological and astrophysical parameters, and the different halo-profile treatments used across simulations. This suggests that the physical relationships encoded in the phase-space of semi-analytic models are largely independent of their specific physical prescriptions, reinforcing their potential as tools for the generation of realistic mock catalogs for cosmological parameter inference.

astro-ph.CO

The DREAMS Project: Disentangling the Impact of Halo-to-Halo Variance and Baryonic Feedback on Milky Way Dark Matter Density Profiles

In this work, we utilize a new suite of Milky Way-mass halos from the DREAMS Project, simulated with Cold Dark Matter (CDM), to quantify the influence of baryon feedback and intrinsic halo-to-halo variance on dark matter density profiles. Our suite of 1024 halos varies over supernova and black hole feedback parameters from the IllustrisTNG model, as well as variations in two cosmological parameters. We find that, for the DREAMS parameter variations, Milky Way-mass dark matter density profiles in the IllustrisTNG model are largely insensitive to astrophysics and cosmology variations, with the dominant source of scatter instead arising from halo-to-halo variance. However, most of the (comparatively minor) feedback-driven variations come from the changes to supernova prescriptions. By comparing to dark matter-only simulations, we find that the strongest supernova wind energies are so effective at preventing galaxy formation that the halos are nearly entirely collisionless dark matter. Finally, regardless of physics variation, all the DREAMS halos are roughly consistent with a halo contracting adiabatically from the presence of baryons, unlike models that have bursty stellar feedback. This work represents a step toward assessing the uncertainty in Milky Way dark matter profiles, with direct implications for dark matter searches where systematic uncertainty in the density profile remains a major challenge.

astro-ph.GA

The DREAMS Project: Disentangling the Impact of Halo-to-Halo Variance and Baryonic Feedback on Milky Way Satellite Galaxies

We analyze the properties of satellite galaxies around 1,024 Milky Way-mass hosts from the DREAMS Project, simulated within a $\Lambda$CDM cosmology. Utilizing the TNG galaxy-formation model, the DREAMS simulations incorporate both baryonic physics and cosmological uncertainties for a large sample of galaxies with diverse environments and formation histories. We investigate the relative impact of the physical uncertainty from the galaxy-formation model on predicted satellite properties using four metrics: the satellite stellar mass function, radial distribution, inner slope of dark matter density profile, and stellar half-light radius. We compare these predictions to observations from the SAGA Survey and the DREAMS N-body simulations and find that uncertainties from baryonic physics modeling are subdominant to the scatter arising from halo-to-halo variance. Where baryonic modeling does affect satellites, the supernova wind energy has the largest effect on the satellite properties that we investigate. Specifically, increased supernova wind energy suppresses the stellar mass of satellites and results in more extended stellar half-light radii. The adopted wind speed has only a minor impact, and other astrophysical and cosmological parameters show no measurable effect. Our findings highlight the robustness of satellite properties against uncertainties in baryonic physics modeling.

astro-ph.GA

The DREAMS Project: A New Suite of 1,024 Simulations to Contextualize the Milky Way and Assess Physics Uncertainties

We introduce a new suite of 1,024 cosmological and hydrodynamical zoom-in simulations of Milky Way-mass halos, run with Cold Dark Matter, as part of the DREAMS Project. Each simulation in the suite has a unique set of initial conditions and combination of cosmological and astrophysical parameters. The suite is designed to quantify theoretical uncertainties from halo-to-halo variance, as well as stellar and black hole feedback. We develop a novel weighting scheme that prioritizes regions of the input parameter space, yielding galaxies consistent with the observed present-day stellar mass--halo mass relation. The resulting galaxy population exhibits a wide diversity in structural properties that encompasses those of the actual Milky Way, providing a powerful statistical sample for galactic archaeology. To demonstrate the suite's scientific utility, we investigate the connection between a galaxy's merger history, focusing on Gaia-Sausage-Enceladus~(GSE) analogs, and its present-day properties. We find that galaxies with a GSE analog have lower star formation rates, more compact disks, and more spherical stellar halos. Crucially, significant halo-to-halo scatter remains, demonstrating that matching more than the most significant events in the Milky Way's past is necessary to recover its present-day properties. Our results highlight the necessity for large statistical samples to disentangle the stochastic nature of galaxy formation and robustly model the Milky Way's unique history.

astro-ph.GA

Learning Cosmology from Nearest Neighbour Statistics

Extracting cosmological parameters from galaxy/halo catalogues with sub-percent level accuracy is an important aspect of modern cosmology, especially in view of ongoing and upcoming surveys such as Euclid, DESI, and LSST. While traditional two-point statistics have been known to be suboptimal for this task, recently proposed k-Nearest Neighbour (kNN) based summary statistics have demonstrated tighter constraining power. Building on the kNN statistics, we introduce a new field-level representation of discrete halo catalogues - NN distance maps. We employ this technique on the halo catalogues obtained from Quijote N-body simulation suites. By combining these maps with kNN-based summary statistics, we train a hybrid neural network to infer cosmological parameters, showing that the resulting constraints achieve state-of-the-art, if not the best, accuracy. In addition, our hybrid framework is 5-10 times more computationally efficient than some of the existing point-cloud-based ML methods.

astro-ph.CO

Linking Warm Dark Matter to Merger Tree Histories via Deep Learning Networks

Dark matter (DM) halos form hierarchically in the Universe through a series of merger events. Cosmological simulations can represent this series of mergers as a graph-like ``tree'' structure. Previous work has shown these merger trees are sensitive to cosmology simulation parameters, but as DM structures, the outstanding question of their sensitivity to DM models remains unanswered. In this work, we investigate the feasibility of deep learning methods trained on merger trees to infer Warm Dark Matter (WDM) particles masses from the DREAMS simulation suite. We organize the merger trees from 1,024 zoom-in simulations into graphs with nodes representing the merger history of galaxies and edges denoting hereditary links. We vary the complexity of the node features included in the graphs ranging from a single node feature up through an array of several galactic properties (e.g., halo mass, star formation rate, etc.). We train a Graph Neural Network (GNN) to predict the WDM mass using the graph representation of the merger tree as input. We find that the GNN can predict the mass of the WDM particle ($R^2$ from 0.07 to 0.95), with success depending on the graph complexity and node features. We extend the same methods to supernovae and active galactic nuclei feedback parameters $A_\text{SN1}$, $A_\text{SN2}$, and $A_\text{AGN}$, successfully inferring the supernovae parameters. The GNN can even infer the WDM mass from merger tree histories without any node features, indicating that the structure of merger trees alone inherits information about the cosmological parameters of the simulations from which they form.

astro-ph.GA

The Denario project: Deep knowledge AI agents for scientific discovery

We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific analysis using Cmbagent as a deep-research backend. In this work, we describe in detail Denario and its modules, and illustrate its capabilities by presenting multiple AI-generated papers generated by it in many different scientific disciplines such as astrophysics, biology, biophysics, biomedical informatics, chemistry, material science, mathematical physics, medicine, neuroscience and planetary science. Denario also excels at combining ideas from different disciplines, and we illustrate this by showing a paper that applies methods from quantum physics and machine learning to astrophysical data. We report the evaluations performed on these papers by domain experts, who provided both numerical scores and review-like feedback. We then highlight the strengths, weaknesses, and limitations of the current system. Finally, we discuss the ethical implications of AI-driven research and reflect on how such technology relates to the philosophy of science. We publicly release the code at https://github.com/AstroPilot-AI/Denario. A Denario demo can also be run directly on the web at https://huggingface.co/spaces/astropilot-ai/Denario, and the full app will be deployed on the cloud.

cs.AI

MG-NECOLA: Fast Neural Emulators for Modified Gravity Cosmologies

Observations of the large-scale structure (LSS) provide a powerful test of gravity on cosmological scales, but high-resolution N-body simulations of modified gravity (MG) are prohibitively expensive. We present MG-NECOLA, a convolutional neural network that enhances fast MG-PICOLA simulations to near-N-body fidelity at a fraction of the cost. MG-NECOLA reproduces QUIJOTE-MG N-body results in the power spectrum and bispectrum with better than 1% accuracy down to non-linear scales ($k \simeq 1~h~\mathrm{Mpc}^{-1}$), while reducing computational time by several orders of magnitude. Importantly, although trained only on $f(R)$ models with massless neutrinos, the network generalizes robustly to scenarios with massive neutrinos, preserving accuracy to within 5% at non-linear scales. This combination of precision and robustness establishes MG-NECOLA as a practical emulator for producing large ensembles of high-fidelity simulations, enabling efficient exploration of modified gravity and beyond-$\Lambda$CDM cosmologies in upcoming surveys.

astro-ph.CO