arXiv ScienceSearch

arXiv subjects

Jing Luo

Publications and source records attributed to Jing Luo.

At least 19 recordsLinked to original sources

Generalized Non-linear Bayesian Pulsar Timing with Enterprise

In this study, we use the Bayesian methods in the Enterprise package to examine the fully general parameterization of pulsar timing models in tandem with noise. We investigate four pulsars, PSR J1600$-$3053, PSR J2043+1711, PSR J0740+6620, and PSR J1640+2224, through the lens of Bayesian timing. These four are selected as they are well-studied, but exhibit interesting characteristics under the lens of Bayesian timing. Our new pulsar mass constraints (medians and 68\% confidence intervals) for our fully general non-linear Bayesian timing models are $m_{\mathrm{p}}=1.6(1)~\mathrm{M}_{\odot}$ for PSR J2043+1711 and $m_{\mathrm{p}}=2.3^{+0.9}_{-0.7}~\mathrm{M}_{\odot}$ for PSR J1600$-$3053 both using the NANOGrav 12.5-yr data release, and $m_{\mathrm{p}}=2.06(6)~\mathrm{M}_{\odot}$ for PSR J0740+6620 using the data from Fonseca, et al., 2021. We investigate the effects on placing physical priors on timing model parameters, including restricting the upper limit on the pulsar mass for PSR J1640+2224, which has a mass often estimated to be greater than $3~\mathrm{M}_{\odot}$. We find \ark{that restricting the allowed sampling space of the pulsar mass for PSR J1640+2224 to} $m_{\mathrm{p}}<3~\mathrm{M}_{\odot}$ results in a pulsar mass of $m_{\mathrm{p}}=2.2(5)~\mathrm{M}_{\odot}$ for PSR J1640+2224 using the NANOGrav 12.5-yr data release. For the first time, we find evidence for intrinsic red noise in PSR J2043+1711. We show how fully general Bayesian timing can better model the interplay of the intrinsic noise and the timing parameters.

astro-ph.HE

Pulsar Timing Array Sensitivity to Anisotropy: Empirical Sensitivity Curves, Scaling Relations, and the Multi-Resolution Pixel Basis

We quantify pulsar timing array (PTA) sensitivity to anisotropy in the gravitational wave background using the cross-correlation based Fisher information matrix in the pixel and spherical harmonic bases. We use a set of simulations to empirically determine scaling relations of a PTA's sensitivity to anisotropy with the number of pulsars $N_\mathrm{psr}$ in the array, the error $\delta t$ on the times of arrival, the frequency $f_\mathrm{GW}$ of the gravitational waves, and the angular scale $\Delta\Omega$ of the anisotropy. The sensitivity scales approximately as $N_\mathrm{psr}^{0.8}$, $\delta t^{-0.08}$, and $\Delta\Omega^{1.6}-\Delta\Omega^{2.1}$ (depending on the ranges of $\ell$ and $m$ under consideration). In addition, we use realistic simulations to project the NANOGrav PTA sensitivity to a 30-year baseline and quantify the growth in sensitivity at several timeslices. Except at the lowest frequencies, we find negligible effect on sensitivity through increasing the observation duration only. Finally, we introduce a multi-resolution pixel basis motivated by the large dependence of the sensitivity on sky location, and demonstrate the operation of the basis through a set of injections and recoveries.

astro-ph.IM

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length song SQA requires modeling how singing quality varies across different audio segments and how these local variations influence the overall evaluation of vocal performance. Moreover, the scarcity of segment-level annotations makes effective supervision challenging, as directly assigning a single overall score label to every segment tends to treat different segment qualities as equivalent. To address these challenges, we propose SongSQA, a two-stage framework for full-length song SQA. In the first stage, a Segment Score Predictor is trained with pseudo labels generated by a pre-trained teacher model, enabling segment-level singing quality prediction without requiring manual segment annotations. In the second stage, a Song Quality Aggregator integrates segment features and predicted segment scores into unified segment embeddings, and employs a learnable song embedding together with self-attention to capture the connection between segment-level vocal performance and overall song quality. In this way, SongSQA dynamically aggregates critical quality cues across the song to produce a holistic quality prediction, while also generating a temporal segment-level quality curve. Experimental results demonstrate the effectiveness of SongSQA for full-length song SQA, achieving up to a 13.95% relative improvement in KTAU over the strongest baseline, while consistently improving other evaluation metrics across all datasets.

cs.SD

Evaluating the Fourier Approximation in Pulsar Timing Array Analysis

Pulsar timing arrays search for stochastic processes such as gravitational waves by comparing pulse time of arrival data for millisecond pulsars to expectations from a background with a given power spectral density (PSD). To make the analysis computationally tractable, the Bayesian likelihood is usually computed using an approximation in which the signal is taken to be a sum of Fourier modes appropriate to the total time of observation, even though the true signal is not periodic. We study the difference between likelihoods computed with this Fourier approximation method for power law spectra and those computed exactly (or using more-closely spaced frequencies as a proxy for the exact result) in the NANOGrav 15-year dataset. We find that the true marginal likelihoods for power-law PSDs are on average about half as large as the likelihoods computed using the Fourier approximation. This could lead to an error of a factor of two in model comparison. However, in the important comparison of uncorrelated vs. Hellings-Downs correlated models, a very similar correction appears in both, so the model comparison is essentially unaffected. We also compare parameter estimation results for power law PSDs, finding little difference between the methods. We briefly discuss spectra with sharper features, for which the approximation could be much worse.

gr-qc

The NANOGrav 15 yr Data Set: Impacts of Customized Chromatic Noise Models on Gravitational Wave Analyses

We report updated nHz gravitational wave (GW) significance, characterization, and interpretations using the customized chromatic-noise models (CNMs) developed in Larsen, Baier et al. (2026). for the NANOGrav 15-year data set. We find increased evidence for the Hellings-Downs (HD) correlation signature of the stochastic gravitational wave background (GWB), with a Bayes factor of $1571\pm14$ for HD-correlations over a common uncorrelated red-noise process using a power-law model with $14$ Fourier modes. We find this $\sim8\times$ increase in Bayes factor from Agazie et al. (2023a) is a result of improved noise mitigation. Assuming an analytic null distribution for the frequentist interpulsar correlation statistic, this corresponds to a slightly more significant measurement from $3.16\sigma$ to $3.32\sigma$ against the no-correlation scenario. Spectral inference with CNMs brings the power-law GWB amplitude down to $A_{\rm GWB} = 2.1^{+0.6}_{-0.5}\times10^{-15}$ at fixed $\gamma_{\rm GWB} = 13/3$. In a varied-$\gamma$ analysis, the spectral index increases to $\gamma_{\rm GWB}=3.5^{+0.7}_{-0.6}$. We report updates on an all-sky continuous gravitational wave (CW) search as well as select targeted searches and calculate a $3.2\times$ larger detection volume for the NANOGrav detector. With CNMs, we find reduced evidence for a non-Einsteinian, scalar-transverse mode of gravity. Finally, we reinterpret the GWB first with the assumption of an astrophysical background sourced by SMBHBs and then assuming the more exotic origins of cosmic inflation, a first-order cosmological phase transition, and stable cosmic strings. Under both the SMBHB hypothesis and the cosmological hypotheses, we see only marginal shifts in model parameter posteriors which are consistent with the slightly quieter and steeper power-law GWB spectrum.

astro-ph.CO

The NANOGrav 15 yr Data Set: Customized Chromatic Noise Models

Pulsar timing arrays conduct low-frequency gravitational wave searches, which require comprehensive accounting of various noise sources to achieve robust results. Interstellar propagation effects (e.g., dispersion and scattering) are especially complex noise sources, introducing chromatic delays that can reduce sensitivity to gravitational waves and bias their inference if left unmodeled. These delays also strongly depend on the line of sight properties to each individual pulsar. To address this, we present customized chromatic noise models for 67 pulsars in the NANOGrav 15 yr dataset. These models are selected from an expanded suite of Gaussian processes to simultaneously characterize multiple types of chromatic delays and are tailored to each pulsar's dataset. Alongside probing the interstellar medium, we use these models to infer the solar wind electron density over the course of $\sim 1.5$ solar cycles. We also find evidence for non-dispersive chromatic delays in 21 out of 67 NANOGrav pulsars. After applying our chromatic models, we observe significant impacts on the inference of achromatic noise in 19 out of 67 pulsars, finding in several cases that a previously significant achromatic noise process can be partially or entirely described as chromatic. These results demonstrate that refined noise modeling is essential to enhance the sensitivity and accuracy of low-frequency gravitational wave searches with pulsar timing arrays.

astro-ph.HE

Can Coding Agents Reproduce Findings in Computational Materials Science?

Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject matter experts, we curate a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims. We then evaluate multiple representative coding agent settings across several foundation models. Our results show that current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 53%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and that they fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.

cs.CL

Analysis of the $D_0^*(2300)$ resonance from lattice QCD under chiral symmetry

We reanalyze the lattice spectra for $I=1/2$ $D\pi$ scattering in the $A_1^+$ irreducible representation from [Phys. Rev. D 111, 014503 (2025)] to investigate the impact of chiral and SU(3) flavor symmetries in $S$-wave $D\pi$ scattering and the $D_0^*(2300)$ resonance. By fitting the phase shifts obtained via L\"uscher's formula with both traditional and chirally modified effective-range expansion and $K$-matrix parameterizations, we find that the chiral factor shifts the extracted pole mass closer to the threshold (especially for resonances) and substantially reduces the resonance width. These findings are confirmed by unitarized chiral perturbation theory through a direct fit to the lattice spectra with both the single-channel and the $D\pi$-$D\eta$-$D_s\bar{K}$ coupled-channel schemes. Once the coupled channels are incorporated, the two-pole structure of the $D_0^*(2300)$ emerges. The trajectories of the two poles are investigated by varying the pion mass.

hep-ph

From Papers to Property Tables: A Priority-Based LLM Workflow for Materials Data Extraction

Scientific data are widely dispersed across research articles and are often reported inconsistently across text, tables, and figures, making manual data extraction and aggregation slow and error-prone. We present a prompt-driven, hierarchical workflow that uses a large language model (LLM) to automatically extract and reconstruct structured, shot-level shock-physics experimental records by integrating information distributed across text, tables, figures, and physics-based derivations from full-text published research articles, using alloy spall strength as a representative case study. The pipeline targeted 37 experimentally relevant fields per shot and applied a three-level priority strategy: (T1) direct extraction from text/tables, (T2) physics-based derivation using verified governing relations, and (T3) digitization from figures when necessary. Extracted values were normalized to canonical units, tagged by priority for traceability, and validated with physics-based consistency and plausibility checks. Evaluated on a benchmark of 30 published research articles comprising 11,967 evaluated data points, the workflow achieved high overall accuracy, with priority-wise accuracies of 94.93% (T1), 92.04% (T2), and 83.49% (T3), and an overall weighted accuracy of 94.69%. Cross-model testing further indicated strong agreement for text/table and equation-derived fields, with lower agreement for figure-based extraction. Implementation through an API interface demonstrated the scalability of the approach, achieving consistent extraction performance and, in a subset of test cases, matching or exceeding chat-based accuracy. This workflow demonstrates a practical approach for converting unstructured technical literature into traceable, analysis-ready datasets without task-specific fine-tuning, enabling scalable database construction in materials science.

cs.AI

The NANOGrav 15 yr and 20 yr Datasets: Timing Events and Pulse Shape Changes

The average pulse shape of a pulsar is typically stable over decadal timescales, enabling estimation of pulse times of arrival to better than a small fraction of the pulse width using matched filtering techniques. However, in North American Nanohertz Observatory for Gravitational Waves (NANOGrav) observations of PSR J1713+0747, three discrete timing events that depart from the prevailing timing model have been seen in the last 20 yr. All three correspond to morphological changes in pulse shape. Using principal component analysis, we analyze the pulse profiles of nine NANOGrav pulsars, including seven with profiles from the 15 yr dataset and two with additional profiles from the forthcoming 20 yr dataset. We recover the three known pulse shape change events in PSR J1713+0747 and another previously known event in PSR J1643$-$1224. We implement a ranking metric for candidate events and address four highly ranked candidates in this nine-pulsar sample. We also recover known slow pulse shape variations in PSR J1643$-$1224, PSR J1903+0327, and PSR B1937+21 and report an unexpected recurrence after ~10 yr of one such variation in PSR B1937+21.

astro-ph.HE

PAC-DP: Personalized Adaptive Clipping for Differentially Private Federated Learning

Differential privacy (DP) is crucial for safeguarding sensitive client information in federated learning (FL), yet traditional DP-FL methods rely predominantly on fixed gradient clipping thresholds. Such static clipping neglects significant client heterogeneity and varying privacy sensitivities, which may lead to an unfavorable privacy-utility trade-off. In this paper, we propose PAC-DP, a Personalized Adaptive Clipping framework for federated learning under record-level local differential privacy. PAC-DP introduces a Simulation-CurveFitting approach leveraging a server-hosted public proxy dataset to learn an effective mapping between personalized privacy budgets epsilon and gradient clipping thresholds C, which is then deployed online with a lightweight round-wise schedule. This design enables budget-conditioned threshold selection while avoiding data-dependent tuning during training. We provide theoretical analyses establishing convergence guarantees under the per-example clipping and Gaussian perturbation mechanism and a reproducible privacy accounting procedure. Extensive evaluations on multiple FL benchmarks show that PAC-DP surpasses conventional fixed-threshold approaches under matched privacy budgets, improving accuracy by up to 26% and accelerating convergence by up to 45.5% in our evaluated settings.

cs.CR

Gravitational Wave Measurement of the $M_\mathrm{BH}$-$M_\mathrm{bulge}$ Intrinsic Scatter at High Redshift

The observed GWB spectrum is higher in amplitude than model predictions by a factor of 2-3. Using a semi-analytic model, we evaluate the effect of a high-scatter supermassive black hole (SMBH) scaling relation ($M_\mathrm{BH}$-$M_\mathrm{bulge}$) on models of the nanohertz gravitational wave background (GWB). By implementing an intrinsic scatter of the $M_\mathrm{BH}$-$M_\mathrm{bulge}$ relation, which is larger at higher redshift, but matches local observations, we find that the amplitude of GWB models increases to be consistent with the low-frequency end of the GWB spectrum. This amplitude increase is not uniform across frequencies, a strongly evolving scatter preferentially increases the number density of the most massive SMBHs which, in the GWB spectrum, minimizes the strength of the low-frequency turnover. Our models with positively evolving intrinsic scatter can reproduce the electromagnetically observed overmassive SMBHs at $4 < z < 6$ without changing the $M_\mathrm{BH}$-$M_\mathrm{bulge}$ normalization though we find that including moderate normalization evolution marginally improves fits to the GWB data. We conclude that the $M_\mathrm{BH}$-$M_\mathrm{bulge}$ relation which best describes the available GWB and electromagnetic data sets has intrinsic scatter that evolves as $\varepsilon(z) = \varepsilon_0 + (0.56 \pm 0.4) \log_{10}(1 + z)$ and normalization that evolves as $\alpha(z) = \alpha_0 (1 + z)^{0.84 \pm 0.35}$. The results of this work imply that the $M_\mathrm{BH}$-$M_\mathrm{bulge}$ relation we see today is not universal throughout cosmic time and that a diversity of seeding models and growth mechanisms may be at play in the early stages of SMBH-galaxy evolution.

astro-ph.HE

Learning Ordinal Probabilistic Reward from Preferences

Reward models are crucial for aligning large language models (LLMs) with human values and intentions. Existing approaches follow either Generative (GRMs) or Discriminative (DRMs) paradigms, yet both suffer from limitations: GRMs typically demand costly point-wise supervision, while DRMs produce uncalibrated relative scores that lack probabilistic interpretation. To address these challenges, we introduce a novel reward modeling paradigm: Probabilistic Reward Model (PRM). Instead of modeling reward as a deterministic scalar, our approach treats it as a random variable, learning a full probability distribution for the quality of each response. To make this paradigm practical, we present its closed-form, discrete realization: the Ordinal Probabilistic Reward Model (OPRM), which discretizes the quality score into a finite set of ordinal ratings. Building on OPRM, we propose a data-efficient training strategy called Region Flooding Tuning (RgFT). It enables rewards to better reflect absolute text quality by incorporating quality-level annotations, which guide the model to concentrate the probability mass within corresponding rating sub-regions. Experiments on various reward model benchmarks show that our method improves accuracy by $\textbf{2.9%}\sim\textbf{7.4%}$ compared to prior reward models, demonstrating strong performance and data efficiency. Analysis of the score distribution provides evidence that our method captures not only relative rankings but also absolute quality.

cs.CL

Design of a 60.8 K superconducting hydride LiMgZr2H12 at ambient pressure via Lithium doping

High-pressure hydrogen-rich compounds have long been regarded as promising room-temperature superconductor candidates; however, their practical applications are limited by their reliance on extreme compression. This study explores hydrogen-rich superconductors that may be stable at ambient pressures. Inspired by recent investigations of the MgZrH2n family, the LiMgZr2H12 structure with a Pmmm symmetry was constructed, and its thermodynamic, mechanical, and dynamical stability were evaluated using first-principles calculations. Electron-phonon coupling (EPC) analysis suggests that LiMgZr2H12 reaches a superconducting critical temperature (Tc) of 60.8 K at ambient pressure. Compared with MgZrH6, Li doping significantly increases the contribution of hydrogen atoms to the electron density of states near the Fermi level (EF) and enhances the EPC constant of the LiMgZr2H12 structure. LiMgZr2H12 exhibits a superconducting figure of merit of 1.56, which is significantly greater than that of MgZrH6, demonstrating its outstanding potential for practical applications. This work guides ambient-pressure design of high-Tc hydrides.

cond-mat.supr-con

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling

Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics underexplored. Moreover, conventional approaches often predict a precise Mean Opinion Score (MOS) value directly, which struggles to capture the nuances of human perception in song aesthetics evaluation. This paper proposes a song-oriented aesthetics evaluation framework, featuring two novel modules: 1) Multi-Stem Attention Fusion (MSAF) builds bidirectional cross-attention between mixture-vocal and mixture-accompaniment pairs, fusing them to capture complex musical features; 2) Hierarchical Granularity-Aware Interval Aggregation (HiGIA) learns multi-granularity score probability distributions, aggregates them into a score interval, and applies a regression within the interval to produce the final score. We evaluated on two datasets of full-length songs: SongEval dataset (AI-generated) and an internal aesthetics dataset (human-created), and compared with two state-of-the-art (SOTA) models. Results show that the proposed method achieves stronger performance for multi-dimensional song aesthetics evaluation. The inference code and checkpoint are publicly available at https://github.com/yisan33/song-aesthetics-evaluation.

cs.SD

The NANOGrav 15 yr Data Set: Piecewise Power-Law Reconstruction of the Gravitational-Wave Background

The NANOGrav 15-year (NG15) data set provides evidence for a gravitational-wave background (GWB) signal at nanohertz frequencies, which is expected to originate either from a cosmic population of inspiraling supermassive black-hole binaries or new particle physics in the early Universe. A firm identification of the source of the NG15 signal requires an accurate reconstruction of its frequency spectrum. In this paper, we provide such a spectral characterization of the NG15 signal based on a piecewise power-law (PPL) ansatz that strikes a balance between existing alternatives in the literature. Our PPL reconstruction is more flexible than the standard constant-power-law model, which describes the GWB spectrum in terms of only two parameters: an amplitude A and a spectral index gamma. Concurrently, it better approximates physically realistic GWB spectra -- especially those of cosmological origin -- than the free spectral model, since the latter allows for arbitrary variations in the GWB amplitude from one frequency bin to the next. Our PPL reconstruction of the NG15 signal relies on individual PPL models with a fixed number of internal nodes (i.e., constant power law, broken power law, doubly broken power law, etc.) that are ultimately combined in a Bayesian model average. The data products resulting from our analysis provide the basis for fast refits of spectral GWB models.

astro-ph.HE

MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning

Traditional workflow-based agents exhibit limited intelligence when addressing real-world problems requiring tool invocation. Tool-integrated reasoning (TIR) agents capable of autonomous reasoning and tool invocation are rapidly emerging as a powerful approach for complex decision-making tasks involving multi-step interactions with external environments. In this work, we introduce MindWatcher, a TIR agent integrating interleaved thinking and multimodal chain-of-thought (CoT) reasoning. MindWatcher can autonomously decide whether and how to invoke diverse tools and coordinate their use, without relying on human prompts or workflows. The interleaved thinking paradigm enables the model to switch between thinking and tool calling at any intermediate stage, while its multimodal CoT capability allows manipulation of images during reasoning to yield more precise search results. We implement automated data auditing and evaluation pipelines, complemented by manually curated high-quality datasets for training, and we construct a benchmark, called MindWatcher-Evaluate Bench (MWE-Bench), to evaluate its performance. MindWatcher is equipped with a comprehensive suite of auxiliary reasoning tools, enabling it to address broad-domain multimodal problems. A large-scale, high-quality local image retrieval database, covering eight categories including cars, animals, and plants, endows model with robust object recognition despite its small size. Finally, we design a more efficient training infrastructure for MindWatcher, enhancing training speed and hardware utilization. Experiments not only demonstrate that MindWatcher matches or exceeds the performance of larger or more recent models through superior tool invocation, but also uncover critical insights for agent training, such as the genetic inheritance phenomenon in agentic RL.

cs.AI

The NANOGrav 12.5-year Data Set: Chromatic Noise Characterization & Mitigation with Time-Domain Kernels

Pulsar timing arrays (PTAs) have recently entered the detection era, quickly moving beyond the goal of simply improving sensitivity at the lowest frequencies for the sake of observing the stochastic gravitational wave background (GWB), and focusing on its accurate spectral characterization. While all PTA collaborations around the world use Fourier-domain Gaussian processes to model the GWB and intrinsic long time-correlated (red) noise, techniques to model the time-correlated radio frequency-dependent (chromatic) processes have varied from collaboration to collaboration. Here we test a new class of models for PTA data, Gaussian processes based on time-domain kernels that model the statistics of the chromatic processes starting from the covariance matrix. As we will show, these models can be effectively equivalent to Fourier-domain models in mitigating chromatic noise. This work presents a method for Bayesian model selection across the various choices of kernel as well as deterministic chromatic models for non-stationary chromatic events and the solar wind. As PTAs turn towards high frequency (>1/yr) sensitivity, the size of the basis used to model these processes will need to increase, and these time-domain models present some computational efficiencies compared to Fourier-domain models.

astro-ph.IM