arXiv ScienceSearch

arXiv subjects

Simran Arora

Publications and source records attributed to Simran Arora.

At least 19 recordsLinked to original sources

Dark Photon - ALP Freeze-in: 511 keV and H$\alpha$ Constraints

We investigate a freeze-in scenario of two component dark matter consisting of an axion-like particle (ALP) and a dark photon. The dark sector connects to the Standard Model through a dimension-five ALP-dark photon interaction, while a small kinetic mixing governs dark photon decays. Solving the coupled Boltzmann equations, we determine the parameter space consistent with the observed relic abundance. We find that, for a nearly degenerate dark sector, dark photon decay into an electron-positron pair through the kinetic mixing explains the Galactic 511 keV line while the dimension five operator can source the necessary production of dark photon and ALP particles to satisfy the relic density. We further confront the model with recent H$\alpha$ observations of dwarf galaxies, together with constraints from the cosmic microwave background, diffuse gamma rays, direct detection and collider searches. We identify viable regions of parameter space yielding $\Omega_{\rm DM}h^2\simeq0.12$, with an effective scale $\Lambda\sim10^{10}$-$10^{12}$ GeV, and dark photon lifetimes of order $10^{26}$-$10^{29}$ s, while remaining consistent with the observed 511 keV photon flux, $H\alpha$ constraints from Leo T and all other astrophysical constraints.

hep-ph

Does DESI prefer Damped Oscillating Dark Energy over Cosmological constant?

We investigate a dark-energy equation of state governed by a damped harmonic oscillator equation, admitting underdamped, critically damped, and overdamped solutions. Confronting the model with Planck CMB distance priors, DESI BAO, BBN, cosmic chronometers, and three Type~Ia supernova compilations, we find that the data select an underdamped solution yielding $H_0 = 70.9 \pm 1.1$ km/s/Mpc with DES-Dovekie and $H_0 = 72.0^{+1.4}_{-2.1}$ km/s/Mpc with Union3, without any local $H_0$ prior. These higher values of $H_0$ arise along the $\Omega_{\rm m}$--$H_0$ degeneracy direction while the sound horizon remains nearly unchanged at $r_{\rm d} \simeq 145$~Mpc, indicating that the enhancement of the late-time expansion rate is a geometrical effect that does not address the early-time calibration of $r_{\rm d}$. In contrast, the Pantheon+ compilation selects a near-critically damped solution with a prior-limited positive $w_0$ and $H_0 = 66.23 \pm 0.85$ km/s/Mpc, highlighting the sensitivity of the model to the low-redshift distance information encoded in the different supernova compilations. The Bayesian evidence relative to $\Lambda$CDM is inconclusive for the DES-Dovekie and Union3 combinations, whereas Pantheon+ shows a strong preference for the damped-oscillator model, driven by the departure from $w=-1$ at $z\lesssim0.1$.

astro-ph.CO

Data-Driven Discovery of a Simple Phantom-Crossing Dark Energy Parametrization

We develop a data-driven reconstruction programme for the dark-energy equation of state within VCDM, a minimally modified gravity framework in which both background and linear perturbations can be consistently evolved across the phantom divide. Using CMB, BAO, and type-Ia supernova data, we first perform a Bayesian spline reconstruction of $w(a)$, finding a preference for smooth, monotonic phantom-crossing trajectories. Bayesian evidence disfavors increasingly complex spline models, indicating that current observations exhibit a statistical preference for low-complexity dark-energy dynamics. Motivated by this result, we apply Exhaustive Symbolic Regression, an interpretable machine-learning technique that systematically searches over analytic expressions of fixed complexity, identifying the remarkably simple one-parameter form $w(a)={w_0}/{\sqrt a}$, which reproduces the reconstructed behaviour and fits the data at a level comparable to standard two-parameter parametrizations such as CPL. The model naturally crosses the phantom divide for $w_0<0$, suppresses early dark energy, and predicts a transient accelerating and phantom phase without a future big-rip singularity. As a one-parameter model, it is highly predictive, being a genuinely dynamical deformation of the cosmological constant rather than containing it as a limit. Bayesian model comparison yields mild-to-moderate support for this parametrization relative to standard two-parameter alternatives, and stronger evidence relative to $\Lambda$CDM. Our results suggest that current observations favour surprisingly simple dark-energy dynamics and illustrate how Bayesian reconstruction and symbolic regression can be combined into a principled model-discovery framework for cosmology.

astro-ph.CO

Constraining Spatial Curvature with Priors from Swampland Conjectures

We study a string-motivated theoretical prior on the quintessential dark energy model with exponential potential, \( V(\phi) = V_0 e^{-\lambda \phi} \), allowing for non-zero spatial curvature. First, we formulate the corresponding dynamical system and investigate its cosmological evolution numerically, illustrating the phase-space behaviour and the influence of curvature on the background dynamics. In open universes (\( \Omega_k > 0 \)), it has been suggested that a curvature-related fixed point may support accelerated expansion even for relatively steep potentials compatible with swampland considerations. Next, we explicitly impose swampland-motivated priors on the slope parameter $\lambda$, restricting it to values consistent with the de Sitter conjecture that excludes the (curved) $\Lambda$CDM limit. Furthermore, we restrict our considerations to the range of field excursion that is consistent with the swampland distance conjecture. Our primary interest is the possibility that such theoretically-motivated priors may shift values of cosmological parameters inferred by observational data, compared with the standard analysis based on theory-agnostic priors such as a sufficiently wide flat prior. We examine this possibility using a combination of Planck CMB data, DESI BAO measurements, and recent Type Ia supernova samples, performing a Bayesian inference of the model parameters. Our analysis indicates that the swampland-motivated prior mildly shifts the values of $\Omega_k$.

astro-ph.CO

Towards Testable Type-III Leptogenesis in Non-Standard Early Universe Scenarios

Leptogenesis is an elegant way to explain the baryon asymmetry of the Universe in connection to the neutrino mass and mixing. Although leptogenesis from the decay of a heavy Majorana neutrino has been the minimal set up, it is also motivating to look for leptogenesis from the decay of triplet fermion as it can have detectable signatures in the experiments. However, due to strong gauge annihilations and constraints from neutrino sector, the triplet fermions have to be as heavy as $10^{10}$ GeV or more to generate the observed baryon asymmetry. While this prediction is based on the standard radiation dominated history of the early Universe, it is also possible to have a non-standard expansion history of the Universe prior to the big-bang nucleosynthesis. In this work we study triplet leptogenesis in two non-standard cosmological scenarios, where the Universe expands faster than radiation and a scalar tensor theory of gravity. We show that it is possible to have successful leptogenesis with a few TeV triplet fermion for fast expanding Universe and a few hundered TeV for a scalar tensor gravity theory.

hep-ph

The eV-Scale Sterile Neutrino and Neutrinoless Double Beta Decay

In short-baseline experiments such as LSND and MiniBooNE, an excess of electron neutrinos has been observed, originating from a muon neutrino beam. To address this anomaly, in the line of many works, we investigate various neutrino mixing schemes involving eV-scale sterile neutrinos alongside three active neutrinos. Using updated experimental and global fit data, we studied neutrinoless double beta decay for three different schemes such as 3+1, 1 + 3, and 2 + 2, which involve one sterile neutrino and three active neutrinos. We have done analysis of these schemes for normal hierarchy (NH) as well as for inverted hierarchy (IH) frameworks, and constrained the sterile neutrino mass in light of current and future neutrinoless double beta decay experiments. The 3+1 scheme is found to be the most viable and at the level of $3\sigma$ the mass of sterile neutrino with respect to the lightest neutrino mass ($m_{\text{lightest}}$) is restricted to $4.75~eV$ for the NH and $4.72~eV$ for the IH. Additionally, the limits on the sum of four neutrino masses are determined to be $4.81~eV$ for the normal hierarchy and $4.78~eV$ for the inverted hierarchy. The updated analysis of all these schemes would help us in understanding physics governing neutrinoless double beta decay and limit on the mass of sterile neutrinos.

hep-ph

Observational constraints on viscous free-$\gamma$ fluid in $f(Q)$ gravity

We study the late-time cosmological dynamics of a spatially flat FLRW universe in the framework of $f(Q)$ gravity, where $Q$ denotes the nonmetricity scalar. The matter sector is modeled as a bulk viscous fluid with a free equation-of-state parameter $\gamma$, allowing for a generalized description of cosmic matter beyond the standard dust approximation. We derive the background evolution equations and analyze the resulting expansion history. The model parameters are constrained using a combination of observational datasets, including cosmic chronometers (CC), baryon acoustic oscillations from DESI DR2, and Type~Ia supernovae (GRBs and Union3). Using the best-fit parameters, we further employ the statefinder and $\mathrm{Om}(z)$ diagnostics to distinguish the viscous $f(Q)$ scenario from the standard $\Lambda$CDM model. In addition, we examine the evolution of the deceleration parameter, which exhibits a transition from an early decelerated phase to the current accelerated expansion, and analyze the effective equation of state behavior. Our results show that bulk viscosity within $f(Q)$ gravity provides a viable and observationally consistent description of late-time cosmic acceleration.

gr-qc

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System

Large language models(LLMs) are now used to power complex multi-turn agentic workflows. Existing systems run agentic inference by loosely assembling isolated components: an LLM inference engine (e.g., vLLM) and a tool orchestrator (e.g., Kubernetes). Although agentic workflows involve multiple LLM and tool requests, these systems schedule and allocate resources separately on a per-request basis, without end-to-end knowledge of the workflow. This leads to sub-optimal management of KV cache and tool execution environments. To address the challenges, we propose ThunderAgent, a fast, simple, and program-aware agentic inference system. We first abstract agentic workflows as LLM Programs, enabling a unified view of heterogeneous resources, including KV caches, system states, and external tool assets such as disk memory and network ports. Built upon this abstraction, ThunderAgent introduces a program-aware scheduler and a tool resource manager designed to maximize KV cache hit rates, mitigate memory imbalances, and enable asynchronous environment preparation. Evaluations across coding, routing, and scientific discovery agents demonstrate that ThunderAgent achieves 1.5-3.6x throughput improvements in serving, 1.8-3.9x in RL rollout, and up to 4.2x disk memory savings compared to state-of-the-art inference systems. To facilitate reproducibility and support future development, we open-source the system implementations of the whole ThunderAgent at: https://github.com/Agentic-Kinetics/ThunderAgent.

cs.OS

ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels

Inter-GPU communication has become a major bottleneck for modern AI workloads as models scale and improvements in hardware compute throughput outpace improvements in interconnect bandwidth. Existing systems mitigate this through compute-communication overlap but often fail to meet theoretical peak performance across heterogeneous workloads and new accelerators. Instead of operator-specific techniques, we ask whether a small set of simple, reusable principles can systematically guide the design of optimal multi-GPU kernels. We present ParallelKittens (PK), a minimal CUDA framework that drastically simplifies the development of overlapped multi-GPU kernels. PK extends the ThunderKittens framework and embodies the principles of multi-GPU kernel design through eight core primitives and a unified programming template, derived from a comprehensive analysis of the factors that govern multi-GPU performance$\unicode{x2014}$data-transfer mechanisms, resource scheduling, and design overheads. We validate PK on both Hopper and Blackwell architectures. With fewer than 50 lines of device code, PK achieves up to $2.33 \times$ speedup for data- and tensor-parallel workloads, $4.08 \times$ for sequence-parallel workloads, and $1.22 \times$ for expert-parallel workloads.

cs.DC

HipKittens: Fast and Furious AMD Kernels

AMD GPUs offer state-of-the-art compute and memory bandwidth; however, peak performance AMD kernels are written in raw assembly. To address the difficulty of mapping AI algorithms to hardware, recent work proposes C++ embedded and PyTorch-inspired domain-specific languages like ThunderKittens (TK) to simplify high performance AI kernel development on NVIDIA hardware. We explore the extent to which such primitives -- for explicit tile-based programming with optimized memory accesses and fine-grained asynchronous execution across workers -- are NVIDIA-specific or general. We provide the first detailed study of the programming primitives that lead to performant AMD AI kernels, and we encapsulate these insights in the HipKittens (HK) programming framework. We find that tile-based abstractions used in prior DSLs generalize to AMD GPUs, however we need to rethink the algorithms that instantiate these abstractions for AMD. We validate the HK primitives across CDNA3 and CDNA4 AMD platforms. In evaluations, HK kernels compete with AMD's hand-optimized assembly kernels for GEMMs and attention, and consistently outperform compiler baselines. Moreover, assembly is difficult to scale to the breadth of AI workloads; reflecting this, in some settings HK outperforms all available kernel baselines by $1.2-2.4\times$ (e.g., $d=64$ attention, GQA backwards, memory-bound kernels). These findings help pave the way for a single, tile-based software layer for high-performance AI kernels that translates across GPU vendors. HipKittens is released at: https://github.com/HazyResearch/HipKittens.

cs.LG

Bayesian and Machine-Learning Analyses of Nonminimal $f(Q)$ Gravity and $H_0$ Tension

In this study, the cosmological implications of nonminimally coupled $f(Q)$ gravity are examined within the metric-affine formalism, in which the nonmetricity scalar $Q$ couples directly to the matter Lagrangian. Within the symmetric teleparallel framework, a representative $f(Q)$ model is constructed, and the corresponding background cosmological equations are derived. The analysis aims to test whether this geometric formulation yields more consistent realizations of nonminimal matter-geometry couplings. A comprehensive statistical MCMC analysis is performed using cosmic chronometers, DESI BAO DR2, and Type~Ia supernovae from the Pantheon+, DESY5, and Union3 samples and CMB. To complement the statistical study, we employ machine learning methods, such as linear regression, support vector regression (SVR), and random forest algorithms, to evaluate the predictive performance and robustness of the data. The results indicate that a partial alleviation of the $H_0$ tension can be achieved for a broad range of parameter choices. Nonetheless, $f(Q)$ gravity emerges as a promising and flexible framework for late-time cosmology, motivating further exploration of extended models consistent with all observations.

gr-qc

Leptogenesis from Dark Matter Coannihilation

We propose a minimal extension of the type-I seesaw model to realise leptogenesis from the co-annihilation of dark sector particles. The type-I seesaw model is extended with a singlet fermion and two singlet scalars charged under a $Z_{2}$ symmetry. The $Z_{2}$-odd singlet scalar is the dark matter candidate. Here the usual type-I seesaw mechanism generates neutrino mass, and a net lepton asymmetry is generated from the co-annihilation of the dark matter and the $Z_2$-odd singlet fermion. The $Z_{2}$-even singlet scalar is important in dark matter phenomenology. Successful leptogenesis is possible at TeV-scale, unlike the vanilla case. This minimal extension provides an elegant explanation of successful leptogenesis with direct connection to the dark matter abundance in the Universe.

hep-ph

Interacting bosonic dark energy and fermionic dark matter in Einstein scalar Gauss-Bonnet gravity

We explore a cosmological framework in which a Gauss-Bonnet (GB) coupled scalar field, acting as dark energy, interacts with a fermionic dark matter field through a coupling obtained from the point of view of particle physics. This setup is inspired by string/M-theory, and two representative scalar field potentials are investigated: exponential and power-law. A distinctive feature of the GB-coupled models is their potential to alter the propagation speed of gravitational waves (GWs), a property with significant implications in light of recent multi-messenger astrophysical observations. To account for this, we analyze models under two scenarios: one where the GW speed differs from that of light and the other where they are equal, but all consistent with current observational constraints. The dynamical evolution of the system is investigated by reformulating the field equations into an autonomous dynamical system, enabling a detailed analysis of the Universe's long-term behavior, including the radiation-, matter- and dark energy-dominated epochs. We constrain the model parameters using a broad set of recent observational data, including mock high-redshift measurements from the Roman Space Telescope. Our findings indicate that both potentials yield cosmologies that are in excellent agreement with current data, closely tracking the expansion history predicted by the standard \(\Lambda\)CDM model, while still allowing room for subtle deviations that could be tested by future observations.

astro-ph.CO

Dynamical dark energy parameterizations in VCDM

In the context of a theory of minimally modified gravity called VCDM, one can realize any cosmological behavior at the level of the homogeneous and isotropic background without introducing fatal instabilities for perturbations. Therefore, VCDM provides a theoretically-consistent and observationally-testable framework of dynamical dark energy parameterizations with or without phantom behaviors. In this paper, we propose the VCDM realizations of various phenomenological parameterizations present in the literature: the Chevallier-Polarski-Linder (CPL), Barboza-Alcaniz (BA), Jassal-Bagla-Padmanabhan (JBP), Exponential (EXP), and Logarithmic (LOG) models. Using the VCDM equations for cosmological perturbations, we test them against the recent cosmological datasets, Planck 2018 and DESI BAO DR2, and then discuss their implications.

gr-qc

Interacting Scalar Fields as Dark Energy and Dark Matter in Einstein scalar Gauss Bonnet Gravity

A Gauss-Bonnet (GB) coupled scalar field $\phi$, responsible for the late-time cosmic acceleration and interacting with a coherent scalar field $\psi$ through an interaction potential $W(\phi,\psi)$, is considered from the point of view of particle physics for two different models. The non-minimal coupling between the GB curvature term and the field $\phi$ leads to a time-dependent speed of gravitational waves (GWs), which is fixed to unity in order to be consistent with current GW observations, rendering the GB coupling function model-independent. We investigate the dynamical stability of the system by formulating it as an autonomous system, and provide a detailed discussion on the choice of initial conditions required to obtain stable background evolution of the models. We constrain the model parameters using various sets of observational data, including both early- and late-time probes. We incorporate the improved Dark Energy Survey (DES) 5-year Type Ia supernova sample (DES-SN5YR), referred to as DES-Dovekie, which exhibits substantially lower tension with the Pantheon+ supernova sample. We find that both models are physically viable and closely follow the $\Lambda$CDM trend for the Pantheon+ and DES samples. However, upon including the Roman mock data, a significant departure is observed at higher redshifts, yielding statistically strong preference over the flat $\Lambda$CDM model.

gr-qc

Cartridges: Lightweight and general-purpose long context representations via self-study

Large language models are often used to answer queries grounded in large text corpora (e.g. codebases, legal documents, or chat histories) by placing the entire corpus in the context window and leveraging in-context learning (ICL). Although current models support contexts of 100K-1M tokens, this setup is costly to serve because the memory consumption of the KV cache scales with input length. We explore an alternative: training a smaller KV cache offline on each corpus. At inference time, we load this trained KV cache, which we call a Cartridge, and decode a response. Critically, the cost of training a Cartridge can be amortized across all the queries referencing the same corpus. However, we find that the naive approach of training the Cartridge with next-token prediction on the corpus is not competitive with ICL. Instead, we propose self-study, a training recipe in which we generate synthetic conversations about the corpus and train the Cartridge with a context-distillation objective. We find that Cartridges trained with self-study replicate the functionality of ICL, while being significantly cheaper to serve. On challenging long-context benchmarks, Cartridges trained with self-study match ICL performance while using 38.6x less memory and enabling 26.4x higher throughput. Self-study also extends the model's effective context length (e.g. from 128k to 484k tokens on MTOB) and surprisingly, leads to Cartridges that can be composed at inference time without retraining.

cs.CL

Probing the Dynamics of Gaussian Dark Energy Equation of State Using DESI BAO

We present an updated reconstruction of the DE equation of state (EoS), $w(a)$, employing the newly released DESI DR2 Baryon Acoustic Oscillation data. This analysis constrains the cosmological scenarios influenced by different models through the joint examination of a range of recently available cosmological probes, specifically the Pantheon+ sample and the DESY5 sample of Type Ia Supernovae, baryon acoustic oscillations, Hubble parameter measurements derived from cosmic chronometers, and cosmic microwave background distance priors based on the Planck 2018 data. Furthermore, we provide a concise perspective on the dynamical evolution of all models (CPL, PADE, GEDE, GDE, BellDE) and their interrelations. A Bayesian inference procedure is adopted to estimate the models parameters that yield the best fit to the data. The EoS remains within the phantom regime at higher redshifts, while favoring the quintessence regime in the current epoch. In this context, we propose a new Gaussian-like form of EoS, termed BellDE, which avoids phantom behavior (\(w \geq -1\)) at higher redshifts while remaining precisely calibrated at lower redshifts. Interestingly, BellDE exhibits a transient phantom nature (\(w < -1\)) around the transition redshift \(z \sim 0.5\), subsequently evolving into a quintessential regime (\(w > -1\)). In particular, the BellDE model provides competitive statistical preference while offering greater flexibility in the redshift regime $z \sim 0.5-1$, where DE is observationally significant.

astro-ph.CO

Towards Learning High-Precision Least Squares Algorithms with Sequence Models

This paper investigates whether sequence models can learn to perform numerical algorithms, e.g. gradient descent, on the fundamental problem of least squares. Our goal is to inherit two properties of standard algorithms from numerical analysis: (1) machine precision, i.e. we want to obtain solutions that are accurate to near floating point error, and (2) numerical generality, i.e. we want them to apply broadly across problem instances. We find that prior approaches using Transformers fail to meet these criteria, and identify limitations present in existing architectures and training procedures. First, we show that softmax Transformers struggle to perform high-precision multiplications, which prevents them from precisely learning numerical algorithms. Second, we identify an alternate class of architectures, comprised entirely of polynomials, that can efficiently represent high-precision gradient descent iterates. Finally, we investigate precision bottlenecks during training and address them via a high-precision training recipe that reduces stochastic gradient noise. Our recipe enables us to train two polynomial architectures, gated convolutions and linear attention, to perform gradient descent iterates on least squares problems. For the first time, we demonstrate the ability to train to near machine precision. Applied iteratively, our models obtain 100,000x lower MSE than standard Transformers trained end-to-end and they incur a 10,000x smaller generalization gap on out-of-distribution problems. We make progress towards end-to-end learning of numerical algorithms for least squares.

cs.LG