arXiv ScienceSearch

arXiv subjects

Matteo Biagetti

Publications and source records attributed to Matteo Biagetti.

At least 19 recordsLinked to original sources

TopoFisher: Learning Topological Summary Statistics by Maximizing Fisher Information

Persistence diagrams provide stable, interpretable summaries of geometric and topological structure and are useful for simulation-based inference when low-order statistics miss key information. Yet persistence-based pipelines require hand-chosen filtrations, vectorizations, and compressors, typically without an objective tied to parameter uncertainty. We introduce \textbf{TopoFisher}, a differentiable persistent-homology pipeline that learns topological summaries by maximizing local Gaussian Fisher information. Using simulations near a fiducial parameter, TopoFisher optimizes trainable filtrations, diagram vectorizations, and compressors without posterior samples or supervised regression targets, while retaining stable topological inductive bias. We also give sufficient regularity conditions for the log-determinant Fisher loss to be locally Lipschitz in trainable parameters. Controlled experiments on noisy spirals and Gaussian random fields, where total Fisher information is known, show that TopoFisher recovers much of the available information and outperforms fixed topological vectorizations. Our main results are on weak gravitational lensing, a high-dimensional non-Gaussian cosmological field-inference problem. Learned topological summaries reach higher Fisher information than state-of-the-art cosmological summaries and approach an unconstrained Information Maximising Neural Network baseline with up to $\sim80\times$ fewer parameters. The learned filtrations also generalize better: under simulator shift from lognormal to LPT-based maps it retains most Fisher information, while the neural baseline drops, and in neural posterior estimation they give tighter constraints than the neural baseline, and of state-of-the-art cosmological summaries. These results support Fisher-based topological optimization as a robust, parameter-efficient front end for simulation-based inference.

stat.ML

Zigzag Persistence of Neural Responses to Time-Varying Stimuli

We use topological data analysis to study neural population activity in the Sensorium 2023 dataset, which records responses from thousands of mouse visual cortex neurons to diverse video stimuli. For each video, we build frame-by-frame cubical complexes from neuronal activity and apply zigzag persistent homology to capture how topological structure evolves over time. These dynamics are summarized with persistence landscapes, providing a compact vectorized representation of temporal features. We focus on one-dimensional topological features-loops in the data-that reflect coordinated, cyclical patterns of neural co-activation. To test their informativeness, we compare repeated trials of different videos by clustering their resulting topological neural representations. Our results show that these topological descriptors reliably distinguish neural responses to distinct stimuli. This work highlights a connection between evolving neuronal activity and interpretable topological signatures, advancing the use of topological data analysis for uncovering neural coding in complex dynamical systems.

math.AT

Understanding and inverse design of implicit bias in stochastic learning: a geometric perspective

A key challenge in machine learning is to explain how learning dynamics select among the many solutions that achieve identical loss values in overparameterized models - a phenomenon known as implicit bias. Controlling this bias provides a direct mechanism on learned representations, which are central to interpretability, robustness, and reasoning in modern AI systems. Yet, despite its importance, existing explanations remain largely ad hoc and lack a unifying mechanism. We develop a theoretical and constructive framework in which implicit bias emerges as a geometric correction induced by the interplay between gradient noise and continuous symmetries of the loss. We compute the induced bias across a range of architectures, predicting new behaviors and explaining known ones. The approach also enables inverse design: by engineering predictor - preserving parameterizations, it is possible to shape the bias, with sparsity and spectral sparsity emerging as canonical instances. Numerical experiments support the theory and validate the inverse - design framework in controlled settings.

cs.LG

The Intrinsic Dimension of Prompts in Internal Representations of Large Language Models

We study the geometry of token representations at the prompt level in large language models through the lens of intrinsic dimension. Viewing transformers as mean-field particle systems, we estimate the intrinsic dimension of the empirical measure at each layer and demonstrate that it correlates with next-token uncertainty. Across models and intrinsic dimension estimators, we find that intrinsic dimension peaks in early to middle layers and increases under syntactic and semantic disruption (by shuffling tokens), and that it is strongly correlated with average surprisal, with a simple analysis linking logits geometry to entropy via softmax. As a case study in practical interpretability and safety, we train a linear probe on the per-layer intrinsic dimension profile to distinguish malicious from benign prompts before generation. This probe achieves accuracy of 90 to 95\% in different datasets, outperforming widely used guardrails such as Llama Guard and Shield Gemma. We further compare against linear probes built from layerwise entropy derived via the Tuned Lens and find that the intrinsic dimension-based probe is competitive and complementary, offering a compact, interpretable signal distributed across layers. Our findings suggest that prompt-level geometry provides actionable signals for monitoring and controlling LLM behavior, and offers a bridge between mechanistic insights and practical safety tools.

cs.CL

Persistent Topological Features in Large Language Models

Understanding the decision-making processes of large language models is critical given their widespread applications. To achieve this, we aim to connect a formal mathematical framework - zigzag persistence from topological data analysis - with practical and easily applicable algorithms. Zigzag persistence is particularly effective for characterizing data as it dynamically transforms across model layers. Within this framework, we introduce topological descriptors that measure how topological features, $p$-dimensional holes, persist and evolve throughout the layers. Unlike methods that assess each layer individually and then aggregate the results, our approach directly tracks the full evolutionary path of these features. This offers a statistical perspective on how prompts are rearranged and their relative positions changed in the representation space, providing insights into the system's operation as an integrated whole. To demonstrate the expressivity and applicability of our framework, we highlight how sensitive these descriptors are to different models and a variety of datasets. As a showcase application to a downstream task, we use zigzag persistence to establish a criterion for layer pruning, achieving results comparable to state-of-the-art methods while preserving the system-level perspective.

cs.CL

Cosmology with Persistent Homology: a Fisher Forecast

Persistent homology naturally addresses the multi-scale topological characteristics of the large-scale structure as a distribution of clusters, loops, and voids. We apply this tool to the dark matter halo catalogs from the Quijote simulations, and build a summary statistic for comparison with the joint power spectrum and bispectrum statistic regarding their information content on cosmological parameters and primordial non-Gaussianity. Through a Fisher analysis, we find that constraints from persistent homology are tighter for 8 out of the 10 parameters by margins of 13-50%. The complementarity of the two statistics breaks parameter degeneracies, allowing for a further gain in constraining power when combined. We run a series of consistency checks to consolidate our results, and conclude that our findings motivate incorporating persistent homology into inference pipelines for cosmological survey data.

astro-ph.CO

On the range of validity of perturbative models for galaxy clustering and its uncertainty

We explore the reach of analytical models at one-loop in Perturbation Theory (PT) to accurately describe measurements of the galaxy power spectrum from numerical simulations in redshift space. We consider the validity range in terms of three different diagnostics: 1) the goodness of fit; 2) a figure-of-bias quantifying the error in recovering the fiducial value of a cosmological parameter; 3) an internal consistency check of the theoretical model quantifying the running of the model parameters with the scale cut. We consider different sets of measurements corresponding to an increasing cumulative simulation volume in redshift space. For each volume we define a median value and the associated scatter for the largest wavenumber where the model is valid (the $k$-reach of the model). We find, as a rather general result, that the median value of the reach decreases with the simulation volume, as expected since the smaller statistical errors provide a more stringent test for the model. This is true for all the three definitions considered, with the one given in terms of the figure-of-bias providing the most stringent scale cut. More interestingly, we find as well that the error associated with the $k$-reach value is quite large, with a significant probability of being as low as 0.1$\, h \, {\rm Mpc}^{-1}$ (or, more generally, up to 40% smaller than the median) for all the simulation volumes considered. We explore as well the additional information on the growth rate parameter encoded in the power spectrum hexadecapole, compared to the analysis of monopole and quadrupole, as a function of simulation volume. While our analysis is, in many ways, rather simplified, we find that the gain in the determination of the growth rate is quite small in absolute value and well within the statistical error on the corresponding figure of merit.

astro-ph.CO

A Model for the Squeezed Bispectrum in the Non-Linear Regime

We present a model for the squeezed dark matter bispectrum, where the short modes are deep in the non-linear regime. We exploit the consistency relations for large-scale structures combined with a response function approach to write the squeezed bispectrum in terms of a few unknown functions of the short modes. We provide an ansatz for a fitting function for these response functions, checking that the resulting model is reliable when compared to the one-loop squeezed bispectrum. We then test the model against measured bispectra from numerical simulations for short modes ranging between $k \sim 0.1 \, h/$Mpc, and $k \sim 0.7 \, h/$Mpc at redshift $z=0$. To evaluate the goodness of the fit of our model we implement a non-Gaussian covariance and find agreement within $1$-$\sigma$ standard deviation of the simulated data.

astro-ph.CO

Primordial non-Gaussianity and non-Gaussian Covariance

In the pursuit of primordial non-Gaussianities, we hope to access smaller scales across larger comoving volumes. At low redshift, the search for primordial non-Gaussianities is hindered by gravitational collapse, to which we often associate a scale $k_{\rm NL}$. Beyond these scales, it will be hard to reconstruct the modes sensitive to the primordial distribution. When forecasting future constraints on the amplitude of primordial non-Gaussianity, $f_{\rm NL}$, off-diagonal components are usually neglected in the covariance because these are small compared to the diagonal. We show that the induced non-Gaussian off-diagonal components in the covariance degrade forecast constraints on primordial non-Gaussianity, even when all modes are well within what is usually considered the linear regime. As a testing ground, we examine the effects of these off-diagonal components on the constraining power of the matter bispectrum on $f_{\rm NL}$ as a function of $k_{\rm max}$ and redshift, confirming our results against N-body simulations out to redshift $z=10$. We then consider these effects on the hydrogen bispectrum as observed from a PUMA-like 21-cm intensity mapping survey at redshifts $2<z<6$ and show that not including off-diagonal covariance over-predicts the constraining power on $f_{\rm NL}$ by up to a factor of $5$. For future surveys targeting even higher redshifts, such as Cosmic Dawn and the Dark Ages, which are considered ultimate surveys for primordial non-Gaussianity, we predict that non-Gaussian covariance would severely limit prospects to constrain $f_{\rm NL}$ from the bispectrum.

astro-ph.CO

Fitting covariance matrix models to simulations

Data analysis in cosmology requires reliable covariance matrices. Covariance matrices derived from numerical simulations often require a very large number of realizations to be accurate. When a theoretical model for the covariance matrix exists, the parameters of the model can often be fit with many fewer simulations. We write a likelihood-based method for performing such a fit. We demonstrate how a model covariance matrix can be tested by examining the appropriate $\chi^2$ distributions from simulations. We show that if model covariance has amplitude freedom, the expectation value of second moment of $\chi^2$ distribution with a wrong covariance matrix will always be larger than one using the true covariance matrix. By combining these steps together, we provide a way of producing reliable covariances without ever requiring running a large number of simulations. We demonstrate our method on two examples. First, we measure the two-point correlation function of halos from a large set of $10000$ mock halo catalogs. We build a model covariance with $2$ free parameters, which we fit using our procedure. The resulting best-fit model covariance obtained from just $100$ simulation realizations proves to be as reliable as the numerical covariance matrix built from the full $10000$ set. We also test our method on a setup where the covariance matrix is large by measuring the halo bispectrum for thousands of triangles for the same set of mocks. We build a block diagonal model covariance with $2$ free parameters as an improvement over the diagonal Gaussian covariance. Our model covariance passes the $\chi^2$ test only partially in this case, signaling that the model is insufficient even using free parameters, but significantly improves over the Gaussian one.

astro-ph.CO

Inflation: Theory and Observations

Cosmic inflation provides a window to the highest energy densities accessible in nature, far beyond those achievable in any realistic terrestrial experiment. Theoretical insights into the inflationary era and its observational probes may therefore shed unique light on the physical laws underlying our universe. This white paper describes our current theoretical understanding of the inflationary era, with a focus on the statistical properties of primordial fluctuations. In particular, we survey observational targets for three important signatures of inflation: primordial gravitational waves, primordial non-Gaussianity and primordial features. With the requisite advancements in analysis techniques, the tremendous increase in the raw sensitivities of upcoming and planned surveys will translate to leaps in our understanding of the inflationary paradigm and could open new frontiers for cosmology and particle physics. The combination of future theoretical and observational developments therefore offer the potential for a dramatic discovery about the nature of cosmic acceleration in the very early universe and physics on the smallest scales.

astro-ph.CO

Fisher Forecasts for Primordial non-Gaussianity from Persistent Homology

We study the information content of summary statistics built from the multi-scale topology of large-scale structures on primordial non-Gaussianity of the local and equilateral type. We use halo catalogs generated from numerical N-body simulations of the Universe on large scales as a proxy for observed galaxies. Besides calculating the Fisher matrix for halos in real space, we also check more realistic scenarios in redshift space. Without needing to take a distant observer approximation, we place the observer on a corner of the box. We also add redshift errors mimicking spectroscopic and photometric samples. We perform several tests to assess the reliability of our Fisher matrix, including the Gaussianity of our summary statistics and convergence. We find that the marginalized 1-$\sigma$ uncertainties in redshift space are $\Delta f_{\rm NL}^{\rm loc} \sim 16$ and $\Delta f_{\rm NL}^{\rm equi} \sim 41 $ on a survey volume of $1$ $($Gpc$/h)^3$. These constraints are weakly affected by redshift errors. We close by speculating as to how this approach can be made robust against small-scale uncertainties by exploiting (non)locality.

astro-ph.CO

Bispectrum-window convolution via Hankel transform

We present a method to perform the exact convolution of the model prediction for bispectrum multipoles in redshift space with the survey window function. We extend a widely applied method for the power spectrum convolution to the bispectrum, taking advantage of a 2D-FFTlog algorithm. As a preliminary test of its accuracy, we consider the toy model of a spherical window function in real space. This setup provides an analytical evaluation of the 3-point function of the window, and therefore it allows to isolate and quantify possible systematic errors of the method. We find that our implementation of the convolution in terms of a mixing matrix shows differences at the percent level in comparison to the measurements from a very large set of mock halo catalogs. It is also able to recover unbiased constraints on halo bias parameters in a likelihood analysis of a set of numerical simulations with a total volume of $100\, h^{-3} \, {\rm Gpc}^3$. For the level of accuracy required by these tests, the multiplication with the mixing matrix is performed in the time of one second or less.

astro-ph.CO

The Covariance of Squeezed Bispectrum Configurations

We measure the halo bispectrum covariance in a large set of N-body simulations and compare it with theoretical expectations. We find a large correlation among (even mildly) squeezed halo bispectrum configurations. A similarly large correlation can be found between squeezed triangles and the long-wavelength halo power spectrum. This shows that the diagonal Gaussian contribution fails to describe, even approximately, the full covariance in these cases. We compare our numerical estimate with a model that includes, in addition to the Gaussian one, only the non-Gaussian terms that are large for squeezed configurations. We find that accounting for these large terms in the modeling greatly improves the agreement of the full covariance with simulations. We apply these results to a simple Fisher matrix forecast, and find that constraints on primordial non-Gaussianity are degraded by a factor of $\sim 2$ when a non-Gaussian covariance is assumed instead of the diagonal, Gaussian approximation.

astro-ph.CO

The reach of next-to-leading-order perturbation theory for the matter bispectrum

We provide a comparison between the matter bispectrum derived with different flavours of perturbation theory at next-to-leading order and measurements from an unprecedentedly large suite of $N$-body simulations. We use the $\chi^2$ goodness-of-fit test to determine the range of accuracy of the models as a function of the volume covered by subsets of the simulations. We find that models based on the effective-field-theory (EFT) approach have the largest reach, standard perturbation theory has the shortest, and `classical' resummed schemes lie in between. The gain from EFT, however, is less than in previous studies. We show that the estimated range of accuracy of the EFT predictions is heavily influenced by the procedure adopted to fit the amplitude of the counterterms. For the volumes probed by galaxy redshift surveys, our results indicate that it is advantageous to set three counterterms of the EFT bispectrum to zero and measure the fourth from the power spectrum. We also find that large fluctuations in the estimated reach occur between different realisations. We conclude that it is difficult to unequivocally define a range of accuracy for the models containing free parameters. Finally, we approximately account for systematic effects introduced by the $N$-body technique either in terms of a scale- and shape-dependent bias or by boosting the statistical error bars of the measurements (as routinely done in the literature). We find that the latter approach artificially inflates the reach of EFT models due to the presence of tunable parameters.

astro-ph.CO

The Formation Probability of Primordial Black Holes

We calculate the exact formation probability of primordial black holes generated during the collapse at horizon re-entry of large fluctuations produced during inflation, such as those ascribed to a period of ultra-slow-roll. We show that it interpolates between a Gaussian at small values of the average density contrast and a Cauchy probability distribution at large values. The corresponding abundance of primordial black holes may be larger than the Gaussian one by several orders of magnitude. The mass function is also shifted towards larger masses.

astro-ph.CO

Topological Echoes of Primordial Physics in the Universe at Large Scales

We present a pipeline for characterizing and constraining initial conditions in cosmology via persistent homology. The cosmological observable of interest is the cosmic web of large scale structure, and the initial conditions in question are non-Gaussianities (NG) of primordial density perturbations. We compute persistence diagrams and derived statistics for simulations of dark matter halos with Gaussian and non-Gaussian initial conditions. For computational reasons and to make contact with experimental observations, our pipeline computes persistence in sub-boxes of full simulations and simulations are subsampled to uniform halo number. We use simulations with large NG ($f_{\rm NL}^{\rm loc}=250$) as templates for identifying data with mild NG ($f_{\rm NL}^{\rm loc}=10$), and running the pipeline on several cubic volumes of size $40~(\textrm{Gpc/h})^{3}$, we detect $f_{\rm NL}^{\rm loc}=10$ at $97.5\%$ confidence on $\sim 85\%$ of the volumes for our best single statistic. Throughout we benefit from the interpretability of topological features as input for statistical inference, which allows us to make contact with previous first-principles calculations and make new predictions.

astro-ph.CO

Primordial Non-Gaussianity from Biased Tracers: Likelihood Analysis of Real-Space Power Spectrum and Bispectrum

Upcoming galaxy redshift surveys promise to significantly improve current limits on primordial non-Gaussianity (PNG) through measurements of 2- and 3-point correlation functions in Fourier space. However, realizing the full potential of this dataset is contingent upon having both accurate theoretical models and optimized analysis methods. Focusing on the local model of PNG, parameterized by $f_{\rm NL}$, we perform a Monte-Carlo Markov Chain analysis to confront perturbation theory predictions of the halo power spectrum and bispectrum in real space against a suite of N-body simulations. We model the halo bispectrum at tree-level, including all contributions linear and quadratic in $f_{\rm NL}$, and the halo power spectrum at 1-loop, including tree-level terms up to quadratic order in $f_{\rm NL}$ and all loops induced by local PNG linear in $f_{\rm NL}$. Keeping the cosmological parameters fixed, we examine the effect of informative priors on the linear non-Gaussian bias parameter on the statistical inference of $f_{\rm NL}$. A conservative analysisof the combined power spectrum and bispectrum, in which only loose priors are imposed and all parameters are marginalized over, can improve the constraint on $f_{\rm NL}$ by more than a factor of 5 relative to the power spectrum-only measurement. Imposing a strong prior on $b_\phi$, or assuming bias relations for both $b_\phi$ and $b_{\phi\delta}$ (motivated by a universal mass function assumption), improves the constraints further by a factor of few. In this case, however, we find a significant systematic shift in the inferred value of $f_{\rm NL}$ if the same range of wavenumber is used. Likewise, a Poisson noise assumption can lead to significant systematics, and it is thus essential to leave all the stochastic amplitudes free.

astro-ph.CO