arXiv ScienceSearch

arXiv subjects

Davide Piras

Publications and source records attributed to Davide Piras.

At least 19 recordsLinked to original sources

Nonlinear velocity power spectrum: modeling the cosmological dependence on the Hubble constant and cold dark matter density

In this paper we present a semi-analytical model for the velocity power spectrum in $\La$CDM cosmology for wave numbers $k<1/$Mpc. We mainly concentrate on the dominant divergence part but also present some results on the vorticity contribution. We divide cosmological parameters into evolution and shape parameters and model the dependence of the evolution parameter $h$ and of the shape parameter $\om_{\rm cdm}$ with an accuracy better than 2.5\%. A surprising finding of our study is that the velocity power spectrum becomes independent of $\om_{\rm cdm}$ on nonlinear scales. A python implementation of the model is publicly available.

astro-ph.CO

Machine Learning and the SKA for Cosmic Dawn and the Epoch of Reionization

When operational, the SKA will generate unprecedented amounts of data and provide exquisite sensitivity for 21 cm tomography of Cosmic Dawn (CD) and the Epoch of Reionization (EoR). With this comes opportunities for new data-driven algorithms that unlock new methods for instrument modelling, data analysis, theoretical simulation, and inference for understanding the high-redshift universe. In this chapter, we provide an overview of some machine learning algorithms that have been proposed for CD and EoR science with the SKA

astro-ph.IM

Field-level weak lensing cosmology with $60$ simulations using multifidelity simulation-based inference

We perform a realistic KiDS-Legacy mock analysis with field-level neural compression and simulation-based inference using just 60 $N$-body simulations. The weak lensing shear field encodes substantially more cosmological information than standard two-point summary statistics such as the power spectrum. Field-level inference can fully exploit this information, but physical realism at the field-level requires very high-fidelity simulations. This poses a major challenge for simulation-based inference (SBI): accurate empirical density modelling and deep-learning-based neural compression require tens of thousands of training samples, but achieving physical realism at the field level makes each simulation extremely costly. We demonstrate that multifidelity SBI can alleviate this tension by substantially reducing the number of high-fidelity simulations needed for accurate cosmological inference. We pre-train neural inference models on realistic KiDS-Legacy-like shear mocks using fast log-normal \texttt{GLASS} simulations and fine-tune them on a small set of high-fidelity $N$-body simulations. We show that $60$ high-fidelity simulations are sufficient to obtain informative and well-calibrated cosmological posteriors, enabling at least an order-of-magnitude reduction in simulation cost for accurate field-level inference in a realistic setting.

astro-ph.CO

Tailoring the birefringence of femtosecond-laser-written multi-scan waveguides in glass

Femtosecond-laser direct waveguide writing is progressively emerging as an alternative to conventional techniques to develop complex photonic devices, for applications ranging from classical and quantum information processing, to sensing and metrology. Laser written waveguides typically offer low modal birefringence, thus preserving coherence of polarization-encoded information. Integrated waveplates have been reported, as waveguides with tilted birefringence axis, but with limited flexibility in terms of achievable rotation angle, birefringence magnitude or control in the modal shape. Here we investigate the multi-scan approach to realize low-loss optical waveguides in fused silica substrate with controlled modal birefringence. We show that by tuning the horizontal and vertical shifts between subsequent scans we can independently change both the magnitude and the axis inclination of the birefringence, while keeping efficient mode coupling with standard fibers.

physics.optics

Implicit inference of the reionization history with higher-order statistics of the 21-cm signal

The Epoch of Reionization (EoR), when the first luminous sources ionised the intergalactic medium, represents a new frontier in cosmology. The Square Kilometre Array Observatory (SKAO) will offer unprecedented insights into this era through observations of the redshifted 21-cm signal, enabling constraints on the Universe's reionization history. We investigate the information content of the average neutral hydrogen fraction ($\bar{x}_{\rm HI}$) in several Gaussian (spherical and cylindrical power spectra) and non-Gaussian (Betti numbers and bispectrum) summary statistics of the 21-cm signal. Mock 21-cm observations are generated using the AA* configuration of SKAO's low-frequency telescope, incorporating noise levels for 100 and 1000 hours. We employ a state-of-the-art implicit inference framework to learn posterior distributions of $\bar{x}_{\rm HI}$ in redshift bins centred at $z=8.0,7.2$ and $6.5$, for each statistic and noise scenario, validating the posteriors through calibration tests. Using the figure of merit to assess constraining power, we find that Betti numbers alone are on average more informative than the power spectra, while the bispectrum provides limited constraints. However, combining higher-order statistics with the cylindrical power spectrum improves the mean figure of merit by $\sim$0.25 dex ($\sim33\%$ reduction in $\sigma(\bar{x}_{\rm HI})$). The relative contribution of each statistic varies with the stage of reionization. With SKAO observations approaching, our results show that combining power spectra with higher-order statistics can significantly increase the information retrieved from the EoR, maximising the scientific return of future 21-cm observations.

astro-ph.CO

Exploring the Early Universe with Deep Learning

Hydrogen is the most abundant element in our Universe. The first generation of stars and galaxies produced photons that ionized hydrogen gas, driving a cosmological event known as the Epoch of Reionization (EoR). The upcoming Square Kilometre Array Observatory (SKAO) will map the distribution of neutral hydrogen during this era, aiding in the study of the properties of these first-generation objects. Extracting astrophysical information will be challenging, as SKAO will produce a tremendous amount of data where the hydrogen signal will be contaminated with undesired foreground contamination and instrumental systematics. To address this, we develop the latest deep learning techniques to extract information from the 2D power spectra of the hydrogen signal expected from SKAO. We apply a series of neural network models to these measurements and quantify their ability to predict the history of cosmic hydrogen reionization, which is connected to the increasing number and efficiency of early photon sources. We show that the study of the early Universe benefits from modern deep learning technology. In particular, we demonstrate that dedicated machine learning algorithms can achieve more than a $0.95$ $R^2$ score on average in recovering the reionization history. This enables accurate and precise cosmological and astrophysical inference of structure formation in the early Universe.

astro-ph.CO

High-Dimensional Bayesian Model Comparison in Cosmology with GPU-accelerated Nested Sampling and Neural Emulators

We demonstrate a GPU-accelerated nested sampling framework for efficient high-dimensional Bayesian inference in cosmology. Using JAX-based neural emulators and likelihoods for cosmic microwave background and cosmic shear analyses, our approach provides parameter constraints and direct calculation of Bayesian evidence. In the 39-dimensional $\Lambda$CDM vs $w_0w_a$ shear analysis, we produce Bayes factors and a robust error bar in just 2 days on a single A100 GPU, without loss of accuracy. Where CPU-based nested sampling can now be outpaced by methods relying on MCMC sampling and decoupled evidence estimation, we demonstrate that with GPU acceleration nested sampling offers the necessary speed-up to put it on equal computational footing with these methods, especially where reliable model comparison is paramount. We also explore interpolation in the matter power spectrum for cosmic shear analysis, finding a further factor of 4 speed-up with consistent posterior contours and Bayes factor. We put forward both nested and gradient-based sampling as useful tools for the modern cosmologist, where cutting-edge inference pipelines can yield orders of magnitude improvements in computation time.

astro-ph.CO

Savage-Dickey density ratio estimation with normalizing flows for Bayesian model comparison

A core motivation of science is to evaluate which scientific model best explains observed data. Bayesian model comparison provides a principled statistical approach to comparing scientific models and has found widespread application within cosmology and astrophysics. Calculating the Bayesian evidence is computationally challenging, especially as we continue to explore increasingly more complex models. The Savage-Dickey density ratio (SDDR) provides a method to calculate the Bayes factor (evidence ratio) between two nested models using only posterior samples from the super model. The SDDR requires the calculation of a normalised marginal distribution over the extra parameters of the super model, which has typically been performed using classical density estimators, such as histograms. Classical density estimators, however, can struggle to scale to high-dimensional settings. We introduce a neural SDDR approach using normalizing flows that can scale to settings where the super model contains a large number of extra parameters. We demonstrate the effectiveness of this neural SDDR methodology applied to both toy and realistic cosmological examples. For a field-level inference setting, we show that Bayes factors computed for a Bayesian hierarchical model (BHM) and simulation-based inference (SBI) approach are consistent, providing further validation that SBI extracts as much cosmological information from the field as the BHM approach. The SDDR estimator with normalizing flows is implemented in the open-source harmonic Python package.

astro-ph.CO

Transfer learning for multifidelity simulation-based inference in cosmology

Simulation-based inference (SBI) enables cosmological parameter estimation when closed-form likelihoods or models are unavailable. However, SBI relies on machine learning for neural compression and density estimation. This requires large training datasets which are prohibitively expensive for high-quality simulations. We overcome this limitation with multifidelity transfer learning, combining less expensive, lower-fidelity simulations with a limited number of high-fidelity simulations. We demonstrate our methodology on dark matter density maps from two separate simulation suites in the hydrodynamical CAMELS Multifield Dataset. Pre-training on dark-matter-only $N$-body simulations reduces the required number of high-fidelity hydrodynamical simulations by a factor between $8$ and $15$, depending on the model complexity, posterior dimensionality, and performance metrics used. By leveraging cheaper simulations, our approach enables performant and accurate inference on high-fidelity models while substantially reducing computational costs.

astro-ph.CO

Anchors no more: Using peculiar velocities to constrain $H_0$ and the primordial Universe without calibrators

We develop a novel approach to constrain the Hubble parameter $H_0$ and the primordial power spectrum amplitude $A_\mathrm{s}$ using type Ia supernovae (SNIa) data. By considering SNIa as tracers of the peculiar velocity field, we can model their distance and their covariance as a function of cosmological parameters without the need of calibrators like Cepheids; this yields a new independent probe of the large-scale structure based on SNIa data without distance anchors. Crucially, we implement a differentiable pipeline in JAX, including efficient emulators and affine sampling, reducing inference time from years to hours on a single GPU. We first validate our method on mock datasets, demonstrating that we can constrain $H_0$ and $\log 10^{10}A_\mathrm{s}$ within $10\%$ and $15\%$, respectively, using $\mathcal{O}(10^3)$ SNIa. We then test our pipeline with SNIa from an $N$-body simulation, obtaining $6\%$-level unbiased constraints on $H_0$ with a moderate noise level. We finally apply our method to Pantheon+ data, constraining $H_0$ at the $15\%$ level without Cepheids when fixing $A_\mathrm{s}$ to its $\it{Planck}$ value. On the other hand, we obtain $20\%$-level constraints on $\log 10^{10}A_\mathrm{s}$ in agreement with $\it{Planck}$ when including Cepheids in the analysis. In light of upcoming observations of low redshift SNIa from the Zwicky Transient Facility and the Vera Rubin Legacy Survey of Space and Time, surveys for which our method will develop its full potential, we make our code publicly available.

astro-ph.CO

Constraining the primordial power spectrum using a differentiable likelihood

The simplest inflationary models predict the primordial power spectrum (PPS) of curvature perturbations to be nearly scale-invariant. However, various other models of inflation predict deviations from this behaviour, motivating a data-driven approach to reconstruct the PPS and constrain its shape. In this work, we present a novel method that employs a fully differentiable pipeline to reconstruct the PPS using Gaussian Processes and uses neural network emulators for fast and differentiable theoretical predictions. By leveraging gradient-based sampling techniques, such as Hamiltonian Monte Carlo, our approach efficiently samples the high-dimensional parameter space of cosmological parameters and the free-form PPS, enabling joint constraints on both. Applying this framework to Planck 2018 Cosmic Microwave Background (CMB) temperature anisotropy data we find our reconstructed PPS to be consistent with near scale-invariance on small scales, while exhibiting large uncertainties at large scales, driven mostly by cosmic variance. Our results show an overestimation of the PPS amplitude compared to $\Lambda$CDM predictions from the Planck 2018 analysis, which we attribute to our choice of a wider prior on the optical depth $\tau$ based on Planck 2015 measurements. Adopting a prior consistent with Planck 2018 measurements brings our results into full agreement with previous work. To ensure robustness of our results, we validate our differentiable pipeline against a non-differentiable framework, and also demonstrate that our results are insensitive to the choice of Gaussian process hyperparameters. These promising results and the flexibility of our pipeline make it ideally suited for application to additional data sets such as CMB polarisation as well as Large-Scale Structure probes, thus moving towards multi-probe primordial power spectrum reconstruction.

astro-ph.CO

$\Lambda$CDM and early dark energy in latent space: a data-driven parametrization of the CMB temperature power spectrum

Finding the best parametrization for cosmological models in the absence of first-principle theories is an open question. We propose a data-driven parametrization of cosmological models given by the disentangled 'latent' representation of a variational autoencoder (VAE) trained to compress cosmic microwave background (CMB) temperature power spectra. We consider a broad range of $\Lambda$CDM and beyond-$\Lambda$CDM cosmologies with an additional early dark energy (EDE) component. We show that these spectra can be compressed into 5 ($\Lambda$CDM) or 8 (EDE) independent latent parameters, as expected when using temperature power spectra alone, and which reconstruct spectra at an accuracy well within the Planck errors. These latent parameters have a physical interpretation in terms of well-known features of the CMB temperature spectrum: these include the position, height and even-odd modulation of the acoustic peaks, as well as the gravitational lensing effect. The VAE also discovers one latent parameter which entirely isolates the EDE effects from those related to $\Lambda$CDM parameters, thus revealing a previously unknown degree of freedom in the CMB temperature power spectrum. We further showcase how to place constraints on the latent parameters using Planck data as typically done for cosmological parameters, obtaining latent values consistent with previous $\Lambda$CDM and EDE cosmological constraints. Our work demonstrates the potential of a data-driven reformulation of current beyond-$\Lambda$CDM phenomenological models into the independent degrees of freedom to which the data observables are sensitive.

astro-ph.CO

Testing interacting dark energy with Stage IV cosmic shear surveys through differentiable neural emulators

We employ a novel framework for accelerated cosmological inference, based on neural emulators and gradient-based sampling methods, to forecast constraints on dark energy models from Stage IV cosmic shear surveys. We focus on dark scattering (DS), an interacting dark energy model with pure momentum exchange in the dark sector, and train COSMOPOWER emulators to accurately and efficiently model the DS non-linear matter power spectrum produced by the halo model reaction framework, including the effects of baryon feedback and massive neutrinos. We embed the emulators within a fully-differentiable pipeline for gradient-based cosmological inference for which the batch likelihood call is up to $O(10^5)$ times faster than with traditional approaches, producing parameter constraints from simulated Stage IV cosmic shear data running on a single graphics processing unit (GPU). We also perform model comparison on the output chains from the inference process, employing the learnt harmonic mean estimator implemented in the software HARMONIC. We investigate degeneracies between dark energy and systematics parameters and assess the impact of scale cuts on the final constraints. Assuming a DS model for the mock data vector, we find that a Stage IV survey cosmic shear analysis can constrain the DS amplitude parameter $A_{\mathrm{ds}}$ with an uncertainty roughly an order of magnitude smaller than current constraints from Stage III surveys, even after marginalising over baryonic feedback, intrinsic alignments and redshift distribution uncertainties. These results show great promise for constraining DS with Stage IV data; furthermore, our methodology can be straightforwardly extended to a wide range of dark energy and modified gravity models.

astro-ph.CO

Psi-GAN: A power-spectrum-informed generative adversarial network for the emulation of large-scale structure maps across cosmologies and redshifts

Simulations of the dark matter distribution throughout the Universe are essential in order to analyse data from cosmological surveys. $N$-body simulations are computationally expensive, and many cheaper alternatives (such as lognormal random fields) fail to reproduce accurate statistics of the smaller, non-linear scales. In this work, we present \textsc{Psi-GAN} (\textbf{P}ower-\textbf{s}pectrum-\textbf{i}nformed \textbf{G}enerative \textbf{A}dversarial \textbf{N}etwork), a machine learning model which takes a two-dimensional lognormal dark matter density field and transforms it into a more realistic field. We construct \textsc{Psi-GAN} so that it is continuously conditional, and can therefore generate realistic realisations of the dark matter density field across a range of cosmologies and redshifts in $z \in [0, 3]$. We train \textsc{Psi-GAN} as a generative adversarial network on $2\,000$ simulation boxes from the Quijote simulation suite. We use a novel critic architecture that utilises the power spectrum as the basis for discrimination between real and generated samples. \textsc{Psi-GAN} shows agreement with $N$-body simulations over a range of redshifts and cosmologies, consistently outperforming the lognormal approximation on all tests of non-linear structure, such as being able to reproduce both the power spectrum up to wavenumbers of $1~h~\mathrm{Mpc}^{-1}$, and the bispectra of target $N$-body simulations to within ${\sim}5$ per cent. Our improved ability to model non-linear structure should allow more robust constraints on cosmological parameters when used in techniques such as simulation-based inference.

astro-ph.CO

Self-Supervised Learning on MeerKAT Wide-Field Continuum Images

Self-supervised learning (SSL) applied to natural images has demonstrated a remarkable ability to learn meaningful, low-dimension representations without labels, resulting in models that are adaptable to many different tasks. Until now, applications of SSL to astronomical images have been limited to Galaxy Zoo datasets, which require a significant amount of pre-processing to prepare sparse images centered on a single galaxy. With wide-field survey instruments at the forefront of the Square Kilometer Array (SKA) era, this approach to gathering training data is impractical. We demonstrate that continuum images from surveys like the MeerKAT Galactic Cluster Legacy Survey (MGCLS) can be successfully used with SSL, without extracting single-galaxy cutouts. Using the SSL framework DINO, we experiment with various preprocessing steps, augmentations, and architectures to determine the optimal approach for this data. We train both ResNet50 and Vision Transformer (ViT) backbones. Our models match state-of-the-art results (trained on Radio Galaxy Zoo) for FRI/FRII morphology classification. Furthermore, they predict the number of compact sources via linear regression with much higher accuracy. However, fine-tuning results in similar performance between our models, the state-of-the-art, and open-source models on multi-class morphology classification. Using source-rich crops from wide-field images to train multi-purpose models is an easily scalable approach that significantly reduces data preparation time. For the tasks evaluated in this work, twenty thousand crops is sufficient training data for models that produce results similar to state-of-the-art. In the future, complex tasks like source detection and characterization, together with domain-specific tasks, ought to demonstrate the true advantages of training models with radio astronomy data over natural-image foundation models.

astro-ph.IM

Deep learning insights into non-universality in the halo mass function

The abundance of dark matter haloes is a key cosmological probe in forthcoming galaxy surveys. The theoretical understanding of the halo mass function (HMF) is limited by our incomplete knowledge of the origin of non-universality and its cosmological parameter dependence. We present a deep learning model which compresses the linear matter power spectrum into three independent factors which are necessary and sufficient to describe the $z=0$ HMF from the state-of-the-art AEMULUS emulator to sub-per cent accuracy in a $w$CDM$+N_\mathrm{eff}$ parameter space. Additional information about growth history does not improve the accuracy of HMF predictions if the matter power spectrum is already provided as input, because required aspects of the former can be inferred from the latter. The three factors carry information about the universal and non-universal aspects of the HMF, which we interrogate via the information-theoretic measure of mutual information. We find that non-universality is captured by recent growth history after matter-dark-energy equality and $N_\mathrm{eff}$ for $M\sim 10^{13} \, \mathrm{M_\odot}\, h^{-1}$ haloes, and by $\Omega_{\rm m}$ for $M\sim 10^{15} \, \mathrm{M_\odot}\, h^{-1}$. The compact representation learnt by our model can inform the design of emulator training sets to achieve high emulator accuracy with fewer simulations.

astro-ph.CO

The future of cosmological likelihood-based inference: accelerated high-dimensional parameter estimation and model comparison

We advocate for a new paradigm of cosmological likelihood-based inference, leveraging recent developments in machine learning and its underlying technology, to accelerate Bayesian inference in high-dimensional settings. Specifically, we combine (i) emulation, where a machine learning model is trained to mimic cosmological observables, e.g. CosmoPower-JAX; (ii) differentiable and probabilistic programming, e.g. JAX and NumPyro, respectively; (iii) scalable Markov chain Monte Carlo (MCMC) sampling techniques that exploit gradients, e.g. Hamiltonian Monte Carlo; and (iv) decoupled and scalable Bayesian model selection techniques that compute the Bayesian evidence purely from posterior samples, e.g. the learned harmonic mean implemented in harmonic. This paradigm allows us to carry out a complete Bayesian analysis, including both parameter estimation and model selection, in a fraction of the time of traditional approaches. First, we demonstrate the application of this paradigm on a simulated cosmic shear analysis for a Stage IV survey in 37- and 39-dimensional parameter spaces, comparing $\Lambda$CDM and a dynamical dark energy model ($w_0w_a$CDM). We recover posterior contours and evidence estimates that are in excellent agreement with those computed by the traditional nested sampling approach while reducing the computational cost from 8 months on 48 CPU cores to 2 days on 12 GPUs. Second, we consider a joint analysis between three simulated next-generation surveys, each performing a 3x2pt analysis, resulting in 157- and 159-dimensional parameter spaces. Standard nested sampling techniques are simply unlikely to be feasible in this high-dimensional setting, requiring a projected 12 years of compute time on 48 CPU cores; on the other hand, the proposed approach only requires 8 days of compute time on 24 GPUs. All packages used in our analyses are publicly available.

astro-ph.CO

Learned harmonic mean estimation of the Bayesian evidence with normalizing flows

We present the learned harmonic mean estimator with normalizing flows - a robust, scalable and flexible estimator of the Bayesian evidence for model comparison. Since the estimator is agnostic to sampling strategy and simply requires posterior samples, it can be applied to compute the evidence using any Markov chain Monte Carlo (MCMC) sampling technique, including saved down MCMC chains, or any variational inference approach. The learned harmonic mean estimator was recently introduced, where machine learning techniques were developed to learn a suitable internal importance sampling target distribution to solve the issue of exploding variance of the original harmonic mean estimator. In this article we present the use of normalizing flows as the internal machine learning technique within the learned harmonic mean estimator. Normalizing flows can be elegantly coupled with the learned harmonic mean to provide an approach that is more robust, flexible and scalable than the machine learning models considered previously. We perform a series of numerical experiments, applying our method to benchmark problems and to a cosmological example in up to 21 dimensions. We find the learned harmonic mean estimator is in agreement with ground truth values and nested sampling estimates. The open-source harmonic Python package implementing the learned harmonic mean, now with normalizing flows included, is publicly available.

astro-ph.IM