arXiv ScienceSearch

arXiv subjects

Lucia A. Perez

Publications and source records attributed to Lucia A. Perez.

At least 19 recordsLinked to original sources

Learning the Universe with cosmological rescaling of merger trees and semi-analytic galaxy formation models

Learning cosmology from galaxy surveys requires large suites of simulations spanning the cosmological and astrophysical parameter space, yet hydrodynamical simulations of galaxy formation remain prohibitively expensive. Semi-analytic models offer an inexpensive, physically grounded alternative, but still require halo merger trees from $N$-body simulations, and densely sampling cosmological parameters in sufficient volume remains expensive. We address this by extending cosmological rescaling to operate directly on merger trees and applying it in the $\Omega_{\rm m}$-$\sigma_8$ plane, running the Santa Cruz semi-analytic model for galaxy formation on the rescaled trees to produce galaxy populations across new cosmological and astrophysical parameters at negligible additional cost. A novel halo-profile-based correction, controlled by a single free parameter, suppresses systematic bias in rescaled halo masses to below the per cent level. We apply the method to parameter estimation of $\Omega_{\rm m}$ and $\sigma_8$ given either the stellar mass function or the two-point correlation function, finding that as few as 64, and potentially fewer, base $N$-body simulations, rescaled to $\sim1000$ training samples, match the accuracy of 750 dedicated $N$-body simulations; rescaling to 3200 realisations improves the prediction of $\Omega_{\rm m}$ by $\sim25\%$. Rescaling all merger trees from a single CAMELS-SAM $N$-body simulation costs $\sim0.1$ CPUh, compared to several thousand CPUh to run the simulation itself. We demonstrate a practical route to obtaining predictions of galaxy summary statistics across cosmological and astrophysical parameters, even with a relatively small number of base $N$-body simulations.

astro-ph.CO

Informative Priors on Primordial Non-Gaussianity Bias $b_{\phi}$ From Galaxy Formation

Constraining primordial non-Gaussianity via its scale-dependent imprint on galaxy clustering requires knowledge of the bias parameter $b_{\phi}$, which is exactly degenerate with $f^{\rm{loc}}_{\rm{NL}}$ at leading order. To break this degeneracy, current analyses adopt the relation $\left(b_{\phi} = 2\delta_c\left(b_1 - 1\right)\right)$ based on the assumption of a universal mass function. This relation is known to break down for physically motivated galaxy selections, introducing systematic errors in the inferred $f^{\rm{loc}}_{\rm{NL}}$ that scale directly with the assumed $b_{\phi}$ prior. We present a framework to construct physically motivated, observation-conditioned priors on $b_{\phi}$ by marginalizing over galaxy formation uncertainties. We use the CAMELS-SAM simulation suite, augmented by separate Universe simulations, to measure galaxy formation observables, like the stellar mass function (SMF) and the stellar-to-halo mass relationship (SHMR), and $b_{\phi}$ across a range of galaxy formation parameters. From these measurements, we construct a distribution of $b_{\phi}$ conditioned on observations, and we select our galaxy sample to resemble the DESI Emission Line Galaxy (ELG) sample. Conditioning on the SMF or SHMR decreases $\sigma_{b_{\phi}}$ from $0.69$ to $0.08$ and $0.02$ respectively -- reductions of $88\%$ and $97\%$ -- with consistent results when conditioning on the observed data directly. Despite substantial shifts in the galaxy formation posteriors driven by known SC-SAM discrepancies at high halo masses, the resulting $b_{\phi}$ distributions remain mutually consistent across all observables. The SMF and SHMR are found to carry sufficient constraining power to reduce the galaxy formation uncertainty in $b_{\phi}$ relevant for $f^{\rm{loc}}_{\rm{NL}}$ inference with next-generation spectroscopic surveys

astro-ph.CO

Introducing sapphire: Towards Hybrid Physics-Informed, Data-Driven Modeling of Galaxy Formation

Semi-analytic models (SAMs) have been treating galaxy populations as dynamical systems for $\gtrsim50$ years, but their evolution equations remain poorly constrained. We introduce sapphire, a modular, automatically differentiable, GPU-accelerated SAM written in JAX. For the first time, we compute exact Jacobian and Hessian matrices of a galaxy formation SAM, using the Pandya et al. (2023) nonlinear differential equation system as an example. These allow efficient, interpretable local and global sensitivity analyses, which reveal that supernova energy loading is the key astrophysical parameter. We use gradient descent and Hamiltonian Monte Carlo (HMC) to perform comprehensive mock parameter recovery tests. These indicate that the $z=0$ stellar-to-halo-mass relation alone does not contain enough information to infer many astrophysical parameters. Using observations of star-forming galaxies from the MaNGA survey and the Behroozi et al. (2019) empirical model as one baseline, we derive multiple posteriors assuming different combinations of data, including $z=0$ interstellar medium gas fractions and metallicities. The inferred physical parameters suggest that galaxies self-regulate their star formation primarily through preventative rather than ejective feedback, though this remains uncertain due to the lack of satellite galaxies, black holes and multi-phase galactic atmosphere physics. Both Fisher and HMC forecasts demonstrate the potential of sapphire to enable precision inference for galaxy formation and cosmology in a hybrid physics-informed, data-driven way, but more work is needed to expand its library of models and methods. We make sapphire publicly available at https://github.com/virajpandya/sapphire.

astro-ph.GA

The Impact of Galaxy Formation on Galaxy Biasing, and Implications for Primordial non-Gaussianity Constraints

The parameter $f_{\textrm{NL}}$ measures the local non-Gaussianity in the primordial energy fluctuations of the Universe, with any deviation from $f_{\textrm{NL}}=0$ providing key constraints on inflationary models. Galaxy clustering is sensitive to $f_{\textrm{NL}}$ at large scale modes and the next generation of galaxy surveys will approach a statistical error of $\sigma_{f_{\textrm{NL}}}\sim1$. However, the systematic errors on these constraints are dominated by the degeneracy of $f_{\textrm{NL}}$ with the galaxy bias parameters $b_1$ (galaxy overdensities caused by mass perturbations) and $b_{\phi}$ (galaxy overdensities caused by primordial potential perturbations). It has been shown that the assumed scaling of $b_{\phi}(z)=2\delta_c (b_1(z)-1)$ is not accurate for realistically simulated galaxies, and depends both on the galaxy selection and the way that galaxies are modeled. To address this, we leverage the CAMELS-SAM pipeline to explore how varying parameters of galaxy formation affects $b_{\phi}$ and $b_1$ for various galaxy selections. We run separate-universe N-body simulations of $L=205 h^{-1}$ cMpc and $N=1280^3$ to measure $b_{\phi}$, and run 55 unique instances of the Santa Cruz semi-analytic model with varying parameters of stellar and AGN feedback. We find the behavior and evolution of a SC-SAM model's stellar-, SFR- and sSFR- to halo mass relationships track well with how $b_1$ and $b_{\phi}(b_1)$ change across redshift and selection for the SC-SAM. We find our variations of the SC-SAM encapsulate the $b_{\phi}$ behavior previously measured in IllustrisTNG, the Munich SAM, and Galacticus.Finally, we identify sSFR selections as particularly robust to varied galaxy modeling.

astro-ph.CO

Galaxy Phase-Space and Field-Level Cosmology: The Strength of Semi-Analytic Models

Semi-analytic models are a widely used approach to simulate galaxy properties within a cosmological framework, relying on simplified yet physically motivated prescriptions. They have also proven to be an efficient alternative for generating accurate galaxy catalogs, offering a faster and less computationally expensive option compared to full hydrodynamical simulations. In this paper, we demonstrate that using only galaxy $3$D positions and radial velocities, we can train a graph neural network coupled to a moment neural network to obtain a robust machine learning based model capable of estimating the matter density parameters, $\Omega_{\rm m}$, with a precision of approximately 10%. The network is trained on ($25 h^{-1}$Mpc)$^3$ volumes of galaxy catalogs from L-Galaxies and can successfully extrapolate its predictions to other semi-analytic models (GAEA, SC-SAM, and Shark) and, more remarkably, to hydrodynamical simulations (Astrid, SIMBA, IllustrisTNG, and SWIFT-EAGLE). Our results show that the network is robust to variations in astrophysical and subgrid physics, cosmological and astrophysical parameters, and the different halo-profile treatments used across simulations. This suggests that the physical relationships encoded in the phase-space of semi-analytic models are largely independent of their specific physical prescriptions, reinforcing their potential as tools for the generation of realistic mock catalogs for cosmological parameter inference.

astro-ph.CO

How does feedback affect the star formation histories of galaxies?

Star formation in galaxies is regulated by the interplay of a range of processes that shape the multiphase gas in the interstellar and circumgalactic media. Using the CAMELS suite of cosmological simulations, we study the effects of varying feedback and cosmology on the average star formation histories (SFHs) of galaxies at $z\sim0$ across the IllustrisTNG, SIMBA and ASTRID galaxy formation models. We find that galaxy SFHs in all three models are sensitive to changes in stellar feedback, which affects the efficiency of baryon cycling and the rates at which central black holes grow, while effects of varying AGN feedback depend on model-dependent implementations of black hole seeding, accretion and feedback. We also find strong interaction terms that couple stellar and AGN feedback, usually by regulating the amount of gas available for the central black hole to accrete. Using a double power-law to describe the average SFHs, we derive a general set of equations relating the shape of the SFHs to physical quantities like baryon fraction and black hole mass across all three models. We find that a single set of equations (albeit with different coefficients) can describe the SFHs across all three CAMELS models, with cosmology dominating the SFH at early times, followed by halo accretion, and feedback and baryon cycling at late times. Galaxy SFHs provide a novel, complementary probe to constrain cosmology and feedback, and can connect the observational constraints from current and upcoming galaxy surveys with the physical mechanisms responsible for regulating galaxy growth and quenching.

astro-ph.GA

CosmoBench: A Multiscale, Multiview, Multitask Cosmology Benchmark for Geometric Deep Learning

Cosmological simulations provide a wealth of data in the form of point clouds and directed trees. A crucial goal is to extract insights from this data that shed light on the nature and composition of the Universe. In this paper we introduce CosmoBench, a benchmark dataset curated from state-of-the-art cosmological simulations whose runs required more than 41 million core-hours and generated over two petabytes of data. CosmoBench is the largest dataset of its kind: it contains 34 thousand point clouds from simulations of dark matter halos and galaxies at three different length scales, as well as 25 thousand directed trees that record the formation history of halos on two different time scales. The data in CosmoBench can be used for multiple tasks -- to predict cosmological parameters from point clouds and merger trees, to predict the velocities of individual halos and galaxies from their collective positions, and to reconstruct merger trees on finer time scales from those on coarser time scales. We provide several baselines on these tasks, some based on established approaches from cosmological modeling and others rooted in machine learning. For the latter, we study different approaches -- from simple linear models that are minimally constrained by symmetries to much larger and more computationally-demanding models in deep learning, such as graph neural networks. We find that least-squares fits with a handful of invariant features sometimes outperform deep architectures with many more parameters and far longer training time. Still there remains tremendous potential to improve these baselines by combining machine learning and cosmology to fully exploit the data. CosmoBench sets the stage for bridging cosmology and geometric deep learning at scale. We invite the community to push the frontier of scientific discovery by engaging with this dataset, available at https://cosmobench.streamlit.app

cs.LG

Learning the Universe: physically-motivated priors for dust attenuation curves

Understanding the impact of dust on the spectral energy distributions (SEDs) of galaxies is crucial for inferring their physical properties and for studying the nature of interstellar dust. We analyze dust attenuation curves for $\sim 6400$ galaxies ($M_{\star} \sim 10^9 - 10^{11.5}\,M_{\odot}$) at $z=0.07$ in the IllustrisTNG50 and TNG100 simulations. Using radiative transfer post-processing, we generate synthetic attenuation curves and fit them with a parametric model that captures known extinction and attenuation laws (e.g., Calzetti, MW, SMC, LMC) and more exotic forms. We present the distributions of the best-fitting parameters: UV slope ($c_1$), optical-to-NIR slope ($c_2$), FUV slope ($c_3$), 2175 Angstrom bump strength ($c_4$), and normalization ($A_{\rm V}$). Key correlations emerge between $A_{\rm V}$ and the star formation rate surface density $\Sigma_{\rm SFR}$, as well as the UV slope $c_1$. The UV and FUV slopes ($c_1, c_3$) and the bump strength and visual attenuation ($c_4, A_{\rm V}$) exhibit robust internal correlations. Using these insights from simulations, we provide a set of scaling relations that predict a galaxy's median (averaged over line of sight) dust attenuation curve based solely on its $\Sigma_{\rm SFR}$ and/or $A_{\rm V}$. These predictions agree well with observed attenuation curves from the GALEX-SDSS-WISE Legacy Catalog despite minor differences in bump strength. This study delivers the most comprehensive library of synthetic attenuation curves for local galaxies, providing a foundation for physically motivated priors in SED fitting and galaxy inference studies, such as those performed as part of the Learning the Universe Collaboration.

astro-ph.GA

LtU-ILI: An All-in-One Framework for Implicit Inference in Astrophysics and Cosmology

This paper presents the Learning the Universe Implicit Likelihood Inference (LtU-ILI) pipeline, a codebase for rapid, user-friendly, and cutting-edge machine learning (ML) inference in astrophysics and cosmology. The pipeline includes software for implementing various neural architectures, training schemata, priors, and density estimators in a manner easily adaptable to any research workflow. It includes comprehensive validation metrics to assess posterior estimate coverage, enhancing the reliability of inferred results. Additionally, the pipeline is easily parallelizable and is designed for efficient exploration of modeling hyperparameters. To demonstrate its capabilities, we present real applications across a range of astrophysics and cosmology problems, such as: estimating galaxy cluster masses from X-ray photometry; inferring cosmology from matter power spectra and halo point clouds; characterizing progenitors in gravitational wave signals; capturing physical dust parameters from galaxy colors and luminosities; and establishing properties of semi-analytic models of galaxy formation. We also include exhaustive benchmarking and comparisons of all implemented methods as well as discussions about the challenges and pitfalls of ML inference in astronomical sciences. All code and examples are made publicly available at https://github.com/maho3/ltu-ili.

astro-ph.IM

Field-level simulation-based inference with galaxy catalogs: the impact of systematic effects

It has been recently shown that a powerful way to constrain cosmological parameters from galaxy redshift surveys is to train graph neural networks to perform field-level likelihood-free inference without imposing cuts on scale. In particular, de Santi et al. (2023) developed models that could accurately infer the value of $\Omega_{\rm m}$ from catalogs that only contain the positions and radial velocities of galaxies that are robust to uncertainties in astrophysics and subgrid models. However, observations are affected by many effects, including 1) masking, 2) uncertainties in peculiar velocities and radial distances, and 3) different galaxy selections. Moreover, observations only allow us to measure redshift, intertwining galaxies' radial positions and velocities. In this paper we train and test our models on galaxy catalogs, created from thousands of state-of-the-art hydrodynamic simulations run with different codes from the CAMELS project, that incorporate these observational effects. We find that, although the presence of these effects degrades the precision and accuracy of the models, and increases the fraction of catalogs where the model breaks down, the fraction of galaxy catalogs where the model performs well is over 90 %, demonstrating the potential of these models to constrain cosmological parameters even when applied to real data.

astro-ph.CO

Constraints on the Epoch of Reionization with Roman Space Telescope and the Void Probability Function of Lyman-Alpha Emitters

We use large simulations of Lyman-Alpha Emitters with different fractions of ionized intergalactic medium to quantify the clustering of Ly$\alpha$ emitters as measured by the Void Probability function (VPF), and how it evolves under different ionization scenarios. We quantify how well we might be able to distinguish between these scenarios with a deep spectroscopic survey using the future Nancy Grace Roman Space Telescope. Since Roman will be able to carry out blind spectroscopic surveys of Ly$\alpha$ emitters continuously between $7 3-4\sigma$ at several redshifts between $7 5-8\sigma$) across the epoch of reionization, and would yield a detailed history of the reionization of the IGM and its effect on Lyman-$\alpha$ Emitter clustering.

astro-ph.CO

Probing Patchy Reionization with the Void Probability Function of Lyman-$\alpha$ Emitters

We probe what constraints for the global ionized hydrogen fraction the Void Probability Function (VPF) clustering can give for the Lyman-Alpha Galaxies in the Epoch of Reionization (LAGER) narrowband survey as a function of area. Neutral hydrogen acts like a fog for Lyman-alpha emission, and measuring the drop in the luminosity function of Lyman-$\alpha$ emitters (LAEs) has been used to constrain the ionization fraction in narrowband surveys. However, the clustering of LAEs is independent from the luminosity function's inherent evolution, and can offer additional constraints for reionization under different models. The VPF measures how likely a given circle is to be empty. It is a volume-averaged clustering statistic that traces the behavior of higher order correlations, and its simplicity offers helpful frameworks for planning surveys. Using the \citet{Jensen2014} simulations of LAEs within various amount of ionized intergalactic medium, we predict the behavior of the VPF in one (301x150.5x30 Mpc$^3$), four (5.44$\times 10^6$ Mpc$^3$), or eight (1.1$\times 10^7$ Mpc$^3$) fields of LAGER imaging. We examine the VPF at 5 and 13 arcminutes, corresponding to the minimum scale implied by the LAE density and the separation of the 2D VPF from random, and the maximum scale from the 8-field 15.5 deg$^2$ LAGER area. We find that even a single DECam field of LAGER (2-3 deg$^2$) could discriminate between mostly neutral vs. ionized. Additionally, we find four fields allows the distinction between 30, 50, and 95 percent ionized; and that eight fields could even distinguish between 30, 50, 73, and 95 percent ionized.

astro-ph.CO

Constraining cosmology with machine learning and galaxy clustering: the CAMELS-SAM suite

As the next generation of large galaxy surveys come online, it is becoming increasingly important to develop and understand the machine learning tools that analyze big astronomical data. Neural networks are powerful and capable of probing deep patterns in data, but must be trained carefully on large and representative data sets. We developed and generated a new `hump' of the Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) project: CAMELS-SAM, encompassing one thousand dark-matter only simulations of (100 $h^{-1}$ cMpc)$^3$ with different cosmological parameters ($\Omega_m$ and $\sigma_8$) and run through the Santa Cruz semi-analytic model for galaxy formation over a broad range of astrophysical parameters. As a proof-of-concept for the power of this vast suite of simulated galaxies in a large volume and broad parameter space, we probe the power of simple clustering summary statistics to marginalize over astrophysics and constrain cosmology using neural networks. We use the two-point correlation function, count-in-cells, and the Void Probability Function, and probe non-linear and linear scales across $0.68<$ R $<27\ h^{-1}$ cMpc. Our cosmological constraints cluster around 3-8$\%$ error on $\Omega_{\text{M}}$ and $\sigma_8$, and we explore the effect of various galaxy selections, galaxy sampling, and choice of clustering statistics on these constraints. We additionally explore how these clustering statistics constrain and inform key stellar and galactic feedback parameters in the Santa Cruz SAM. CAMELS-SAM has been publicly released alongside the rest of CAMELS, and offers great potential to many applications of machine learning in astrophysics: https://camels-sam.readthedocs.io.

astro-ph.GA

The CAMELS project: public data release

The Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) project was developed to combine cosmology with astrophysics through thousands of cosmological hydrodynamic simulations and machine learning. CAMELS contains 4,233 cosmological simulations, 2,049 N-body and 2,184 state-of-the-art hydrodynamic simulations that sample a vast volume in parameter space. In this paper we present the CAMELS public data release, describing the characteristics of the CAMELS simulations and a variety of data products generated from them, including halo, subhalo, galaxy, and void catalogues, power spectra, bispectra, Lyman-$\alpha$ spectra, probability distribution functions, halo radial profiles, and X-rays photon lists. We also release over one thousand catalogues that contain billions of galaxies from CAMELS-SAM: a large collection of N-body simulations that have been combined with the Santa Cruz Semi-Analytic Model. We release all the data, comprising more than 350 terabytes and containing 143,922 snapshots, millions of halos, galaxies and summary statistics. We provide further technical details on how to access, download, read, and process the data at \url{https://camels.readthedocs.io}.

astro-ph.CO

New spectroscopic confirmations of Lyman-$\alpha$ emitters at z $\sim$ 7 from the LAGER survey

We report spectroscopic confirmations of 15 Lyman-alpha galaxies at $z\sim7$, implying a spectroscopic confirmation rate of $\sim$80% on candidates selected from LAGER (Lyman-Alpha Galaxies in the Epoch of Reionization), which is the largest (24 deg$^2$) survey aimed at finding Lyman-alpha emitters (LAEs) at $z\sim7$ using deep narrow-band imaging from DECam at CTIO. LAEs at high-redshifts are sensitive probes of cosmic reionization and narrow-band imaging is a robust and effective method for selecting a large number of LAEs. In this work, we present results from the spectroscopic follow-up of LAE candidates in two LAGER fields, COSMOS and WIDE-12, using observations from Keck/LRIS. We report the successful detection of Ly$\alpha$ emission in 15 candidates (11 in COSMOS and 4 in WIDE-12 fields). Three of these in COSMOS have matching confirmations from a previous LAGER spectroscopic follow-up and are part of the overdense region, LAGER-$z7$OD1. Additionally, two candidates that were not detected in the LRIS observations have prior spectroscopic confirmations from Magellan. Including these, we obtain a spectroscopic confirmation success rate of $\sim$$80$% for LAGER LAE candidates. Apart from Ly$\alpha$, we do not detect any other UV nebular lines in our LRIS spectra; however, we estimate a 2$\sigma$ upper limit for the ratio of NV/Ly$\alpha$, $f_{NV}/f_{Ly\alpha} \lesssim 0.27$, which implies that ionizing emission from these sources is mostly dominated by star formation. Including confirmations from this work, a total of 33 LAE sources from LAGER are now spectroscopically confirmed. LAGER has more than doubled the sample of spectroscopically confirmed LAE sources at $z\sim7$.

astro-ph.GA

The CAMELS Multifield Dataset: Learning the Universe's Fundamental Parameters with Artificial Intelligence

We present the Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) Multifield Dataset, CMD, a collection of hundreds of thousands of 2D maps and 3D grids containing many different properties of cosmic gas, dark matter, and stars from 2,000 distinct simulated universes at several cosmic times. The 2D maps and 3D grids represent cosmic regions that span $\sim$100 million light years and have been generated from thousands of state-of-the-art hydrodynamic and gravity-only N-body simulations from the CAMELS project. Designed to train machine learning models, CMD is the largest dataset of its kind containing more than 70 Terabytes of data. In this paper we describe CMD in detail and outline a few of its applications. We focus our attention on one such task, parameter inference, formulating the problems we face as a challenge to the community. We release all data and provide further technical details at https://camels-multifield-dataset.readthedocs.io.

cs.LG

LAGER Ly$\alpha$ Luminosity Function at $z\sim7$, Implications for Reionization

We present a new measurement of the Ly$\alpha$ luminosity function at redshift $z=6.9$, finding moderate evolution from $z=5.7$ that is consistent with a fully or largely ionized $z\sim7$ intergalactic medium. Our result is based on four fields of the LAGER (Lyman Alpha Galaxies in the Epoch of Reionization) project. Our survey volume of $6.1\times10^{6}$ Mpc$^{3}$ is double that of the next largest $z\sim 7$ survey. We combine two new LAGER fields (WIDE12 and GAMA15A) with two previously reported LAGER fields (COSMOS and CDFS). In the new fields, we identify $N=95$ new $z=6.9$ Ly$\alpha$ emitters (LAEs); characterize our survey's completeness and reliability; and compute Ly$\alpha$ luminosity functions. The best-fit Schechter luminosity function parameters for all four LAGER fields are in good general agreement. Two fields (COSMOS and WIDE12) show evidence for a bright-end excess above the Schechter function fit. We find that the Ly$\alpha$ luminosity density declines at the same rate as the UV continuum LF from $z=5.7$ to $z=6.9$. This is consistent with an intergalactic medium that was fully ionized as early as redshift $z\sim 7$, or with a volume-averaged neutral hydrogen fraction of $x_{HI} < 0.33$ at $1\sigma$.

astro-ph.GA

Correlations between H$\alpha$ Equivalent Width and Galaxy Properties at $z = 0.47$: Physical or Selection-driven?

The H$\alpha$ equivalent width (EW) is an observational proxy for specific star formation rate (sSFR) and a tracer of episodic star-formation activity. Previous assessments show that EW strongly anti-correlates with stellar mass as $M^{-0.25}$ similar to the sSFR -- stellar mass relation. However, such a correlation may be driven/formed by selection effects. In this study, we investigate how H$\alpha$ EWs correlate with galaxy properties and how selection biases could alter such correlations using a $z = 0.47$ narrowband-selected sample of 1572 H$\alpha$ emitters from the Ly$\alpha$ Galaxies in the Epoch of Reionization (LAGER) survey. The sample covers 3 deg$^2$ of COSMOS and $1.1\times10^5$ cMpc$^3$. We assume an intrinsic EW distribution to form mock samples of H$\alpha$ emitters (HAEs) and propagate the selection criteria to match observations, giving us control on how selection biases can affect the underlying results. We find EW intrinsically correlates with stellar mass as $W_0 \propto M^{-0.16\pm0.03}$ and decreases by a factor of $\sim 3$ from $10^{7}$ to $10^{10}$ M$_\odot$. We find low-mass HAEs to be $\sim 320$ times more likely to have rest-frame EW$ > 200$\AA compared to high-mass HAEs. Combining the intrinsic EW -- stellar mass correlation with an observed SMF correctly reproduces the observed H$\alpha$ LF, while not correcting for selection effects underestimates the number of bright HAEs. This suggests that the intrinsic EW -- stellar mass correlation is physically significant and reproduces three statistical distributions of galaxy populations (LF, SMF, EW distribution). At lower masses, we find there are more high-EW outliers compared to high masses, even after taking into account selection effects. Our results suggest that high sSFR outliers indicative of bursty SF activity are intrinsically more prevalent in low-mass HAEs and not a byproduct of selection effects.

astro-ph.GA