arXiv ScienceSearch

arXiv subjects

Jia Liu

Publications and source records attributed to Jia Liu.

At least 19 recordsLinked to original sources

Hyperparameter Scaling Laws Across MoE Sparsity

Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone. To characterize this dependence, we conduct 1,800 pre-training runs spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens at a cost of 200,000 equivalent H800 GPU-hours. Our results reconcile conflicting findings in prior work by revealing two scaling regimes. At fixed sparsity, the optimal batch size follows a power-law relationship with training tokens $D$, whereas the optimal learning rate scales with training compute $C$ and remains robust to the allocation between model size and data. Across sparsity levels, the activation ratio $A$ enters both relationships as an additional multiplicative power-law factor. These observations lead to unified hyperparameter scaling laws that transfer across MoE sparsity levels. Large-scale evaluation shows that the scaling form outperforms alternative functional forms. On a held-out ultra-sparse MoE with 12B total parameters and only 1/64 of its experts activated, the predicted hyperparameters remain close to the observed optima, supporting joint extrapolation across model scale and sparsity. Further experiments demonstrate transfer across expert granularities and isolate the effect of activation ratio from that of total expert count.

cs.LG

Microscopic Calculation of Electric Quadrupole Effective Charges in Exotic Nuclei

Electric quadrupole ($E2$) effective charges are evaluated based on the self-consistent relativistic Hartree-Fock single-particle states, with core-polarization corrections resummed to all orders using the Tamm-Dancoff approximation (TDA). Configuration-interaction relativistic Hartree-Fock (CI-RHF) calculations employing the TDA effective charges well reproduce the $B(E2)$ strength for neon isotopes from stability to the neutron drip line. We find that polarization charges associated with continuum states are significantly quenched due to their extended density distributions and weak coupling to the core, underscoring the critical role of continuum effects in $E2$ transition evaluations for exotic nuclei. Moreover, the CI-RHF model predicts a suppressed $B(E2; 2^+_2 \to 0^+_2)$ in $^{30}$Ne, together with strong in-band $B(E2)$ strengths of the yrast band, suggesting the coexistence of a nearly spherical excited $0^+_2$ state and a deformed ground state within the $N=20$ "island of inversion".

nucl-th

B(E2) Serves as a Robust Signature of N = 32,34 Shell Evolution

Electric quadrupole transition probabilities $B(E2)$ serve as key probe of nuclear shell evolution, yet anomalous $B(E2)$ values in exotic nuclei complicate the identification of new magic numbers. In this letter, employing the configuration-interaction relativistic Hartree-Fock model, we demonstrate that effective charges are sensitive to orbital radii, and this orbital dependence is significantly amplified by the halo structure of valence nucleons. This mechanism is critical for reliably describing $E2$ transitions and understanding the unusual behavior of $B(E2)$ in exotic nuclei. Our calculations predict reduced $B(E2; 2^+_1 \rightarrow 0^+_1)$ values in $^{52,54}\text{Ca}$, signaling the emergence of subshell closures at $N=32$ and 34. Furthermore, the suppressed $B(E2; 7/2^{-}_{1} \rightarrow 11/2^{-}_{1})$ transition in $^{53}\text{Sc}$ underscores the robustness of the $N=32$ new magic number, whereas the enhanced transition strength in $^{55}\text{Sc}$ indicates the rapid erosion of the $N=34$ shell gap with the occupancy of the proton orbital $\pi1f_{7/2}$.

nucl-th

Future of Artificial Intelligence for Science in Japan 2024 Community Report

This white paper summarizes scientific challenges and AI/ML research opportunities identified through the FAIRS Japan 2024 unconference process. The discussion focuses on three major physics domains: accelerator physics, cosmology and astrophysics, and neutrino physics. Although each domain has distinct scientific goals and experimental constraints, several common technical themes emerge: high-dimensional reconstruction, fast and accurate simulation, uncertainty propagation, simulation-to-data mismatch, anomaly detection, real-time decision-making, and shared infrastructure.

hep-ph

A Synthetic Iterative Scheme for Non-Gray Phonon Boltzmann Transport Equation with Dual Relaxation Times

Solving the non-gray Callaway phonon Boltzmann transport equation allows a dual-relaxation-time approximation of separate normal and resistive scatterings and resolving the mode-dependent spectrum. The conventional iterative scheme (CIS) for deterministic solutions avoids a monolithic phase-space inversion, but its collision-source iteration can become prohibitively slow at a large characteristic length of a material. Existing synthetic acceleration schemes address either non-gray single-relaxation models or gray dual-relaxation models, leaving mode-resolved dual-relaxation transport without a dedicated acceleration framework. We develop a general synthetic iterative scheme (GSIS) for the stationary, linearized, non-gray Callaway equation, where synthetic approximations for the normal-process pseudo-temperature and phonon drift velocity are provided by exact energy and quasi-momentum balance laws closed with first-order Chapman-Enskog constitutive relations and non-equilibrium terms evaluated from the kinetic solution. The resistive-process pseudo-temperature is retrieved from the two quantities. Each iteration couples an upwind nodal discontinuous Galerkin kinetic sweep and a hybridizable discontinuous Galerkin solution of the synthetic equations to achieve high-order spatial discretization. A branch- and frequency-resolved Fourier analysis identifies the deterioration of CIS and shows that the GSIS contraction factor remains bounded away from unity for the considered graphene material. Asymptotic analysis indicates that GSIS reduces to a consistent discretization of a Guyer-Krumhansl-like equation and Fourier's law of heat conduction in the hydrodynamic and diffusive limits, respectively.

physics.comp-ph

An Axial $U_A(1)_{L_\mu-L_\tau}$: UV Completion and Experimental Searches

We propose an anomaly-free and renormalizable axial $U_A(1)_{L_\mu-L_\tau}$ model and study its experimental signatures for $A'$ masses from the MeV scale to the TeV scale. The opposite charges of the left- and right-handed charged leptons forbid the usual muon and tau Yukawa interactions. Their masses are instead generated by a singlet scalar and heavy vector-like leptons through a universal-seesaw mechanism. We focus on heavy vector-like leptons, small light--heavy mixing, and $m_s\gtrsim10~\mathrm{GeV}$. In this limit, the observables considered here depend mainly on $(m_{A'},g_X)$, while the other model parameters are restricted by mixing and perturbativity. We confront this benchmark with current experimental searches. For neutrino trident production, our finite-$m_\mu$ calculation shows that $A'$ modifies the axial weak coefficient, rather than the vector coefficient relevant to the usual $L_\mu-L_\tau$ model. The longitudinal mode enhances muon bremsstrahlung and gives a negative contribution to $(g-2)_\mu$; the latter dominates over the scalar contribution in our benchmark. Combining these results with invisible meson decays and four-muon resonance searches, we summarize the phenomenological constraints in the $(m_{A'},g_X)$ plane. For $m_{A'}\gg m_\mu$, vector and axial final-state-radiation rates become nearly identical, so the corresponding collider limits can be obtained by rate matching. At a future muon collider, the total rate alone does not fully resolve the interaction structure, whereas angular distributions, especially the forward--backward asymmetry in $\mu^+\mu^-\to\tau^+\tau^-$, retain direct sensitivity to chirality.

hep-ph

Parallelizable Gradient-Based Optimization For Multi-Objective MaxCut

Multi-objective combinatorial optimization arises in a wide range of problems and applications, including the canonical multi-objective MaxCut problem. Differentiable single-instance quadratic methods have recently achieved remarkable performance in single-objective combinatorial optimization. In this paper, we develop a differentiable framework for multi-objective MaxCut by combining an adjacency-based quadratic formulation with linear scalarization, thereby reducing the problem to a preference-conditioned single-objective signed-weight MaxCut problem. Theoretically, we characterize the stationary points of the resulting signed-weight formulation and show how they induce preference-conditioned fixed points on the Pareto front. Computationally, unlike conventional heuristics and branch-and-bound methods, our approach is GPU-parallelizable and can therefore benefit from substantial performance speedups. We term our algorithm Multi-objective QUadratic Combinatorial Optimization (MO-QUCO) and its parallelized variant pMO-QUCO. Empirically, across different multi-layered (and weight distributions) graphs, we show that both our CPU-only and GPU-based algorithms outperform SOTA exact and heuristic methods in terms of wall-clock runtime and objective quality. Despite operating under different computational settings, MO-QUCO also outperforms the SOTA quantum method.

cs.DM

Topological phase rectification via Aharonov-Bohm interference in a Majorana--quantum-dot interferometer

We propose and theoretically investigate a topological superconducting rectifier based on a quantum-dot--Majorana interferometer. The Aharonov-Bohm phase, controlled by a magnetic flux threading the interferometer loop, tunes the quantum interference between a trivial $2\pi$-periodic quantum-dot channel and a topological $4\pi$-periodic Majorana channel. At non-integer flux, this interference generates a persistent current background $I_{\rm off}$ that shifts the current-phase relation into a unipolar regime, in which the supercurrent flows strictly in one direction. We introduce a signed unipolarity factor $\eta_u$, with $|\eta_u|>0.5$ defining the unipolar regime, and establish its quantitative relationship to the conventional diode efficiency $\eta$. The unipolarity proves robust against variations of the quantum-dot level, spin polarization, and Majorana hybridization, is enhanced by stronger Majorana coupling and Rashba spin-orbit interaction, and persists at realistic temperatures and under quasiparticle poisoning. We further propose a topological diode figure of merit $\mathcal{Z}_{\rm TD}$, defined from the Fourier spectrum of $\eta_u$, whose nonzero value provides a model-independent signature of the $4\pi$-periodic Majorana channel and distinguishes topological from trivial rectification mechanisms. Our findings establish the quantum-dot--Majorana interferometer as a promising route toward high-performance topological superconducting diodes with clear experimental signatures accessible via standard dc transport measurements.

cond-mat.mes-hall

Flavor--Kinetic Entanglement Production from Decay and Scattering at Finite Density

We extend the scattering-entanglement dictionary to finite-density environments by investigating the flavor--kinetic bipartition of the Hilbert space. We show that tracing over kinematic degrees of freedom maps the total branch-changing transition probability directly onto the leading flavor--kinetic linear entanglement entropy. At finite density, the vacuum branch-changing probability is replaced by an occupation-weighted collision probability, built from the same directed reaction-density kernel that enters the integrated Boltzmann equation. The resulting observable is the bath-averaged flavor--kinetic entanglement entropy of a pair sampled from the medium. As a proof of principle, this framework is applied to an $O(N)$ singlet-scalar extended model to probe thermal phase transitions. In the examples studied, the resulting entanglement entropy serves as a collision-based phase-transition-type diagnostic, exhibiting a finite discontinuity across a first-order phase transition and a nonanalytic temperature derivative for continuous transitions. These examples suggest a novel way to characterize thermal phase structures, distinct from traditional thermodynamic order parameters.

hep-ph

Collider Spin Tomography with Missing Neutrinos

Missing neutrinos need not destroy collider spin tomography. We formulate the visible measurement under kinematic ambiguities arising from invisible particles as a coarse-grained positive-operator-valued measure on the production spin density matrix. We show that information loss is governed by the null space of the resulting visible-data map, not by the number of kinematic solutions. In $e^+e^-\to\tau^+\tau^-\to\pi^+\pi^-+\nu\bar\nu$, the twofold ambiguity leaves only the antisymmetric spin-correlation combination $C_{nr}-C_{rn}$ unidentifiable, while the differential production rate and the remaining fourteen spin coefficients are identifiable. For practical reconstruction under kinematic ambiguities, we develop a self-consistent fixed-point unfolding method using only visible data, without assuming a theoretical production template. Closure tests in Standard Model and anomalous tau-dipole benchmarks show that the method reproduces the truth-level differential production rate and all identifiable spin coefficients, whereas the usual flat average over kinematic folds gives significantly biased reconstructions. When a nontrivial null space is present, the reconstructed identifiable subspace together with positivity yields controlled ranges for concurrence and the CHSH parameter.

hep-ph

Generation of high-fluence and high-intensity hard x-ray attosecond pulses at European XFEL

By combining hard x-ray attosecond pulses from the European XFEL with total-reflection focusing x-ray optics, we generated nanofocused hard x-ray attosecond pulses with intensities and fluences comparable to the highest values attained in the hard x-ray regime. A peak intensity on the order of 10$^{20}$ W/cm$^2$ is confirmed through the observation of saturation in amplified spontaneous emission from copper atoms. These x-ray pulses enable new scientific opportunities, including the exploration of higher-order nonlinear light--matter interactions, damage-free structure determination, and coherent control of atoms and molecules.

physics.optics

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.

astro-ph.IM

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.

cs.CL

Accuracy Analysis of VLBI Universal Time Measurement Based on a GNSS Single-Station Regional Ionospheric Model

Universal Time (UT1) is a key parameter characterizing Earth's rotation, and very long baseline interferometry (VLBI) is the mainstream technique for measuring UT1. To address the limitations in the timeliness and accuracy of existing global ionospheric models for single-frequency VLBI UT1 measurements, we construct a single-station regional ionospheric model using GNSS data from the VLBI stations on the Jilin-Kashi baseline. We apply this model to VLBI observations and compare its correction performance with that of a global predictive model and a global post-processed model. The results show that the line-of-sight ionospheric delays and baseline corrections calculated with the single-station regional model have precision close to that of the global post-processed model and are substantially better than those of the global predictive model. After correction with the single-station regional model, the derived UT1 values differ from the US Naval Observatory (USNO) reference values by a mean bias of -15.6 us and an RMS deviation of 82.3 us, both better than the results obtained with the other two model classes. A single-station regional ionospheric model constructed independently from GNSS data available at VLBI stations can effectively correct single-frequency VLBI observations and support quasi-real-time high-precision UT1 measurements. It therefore has important value for improving the timeliness of independent UT1 products.

astro-ph.EP

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.

cs.LG

Optimus: A Generic Operator-Level PyTorch Model Transformation Framework

In large-scale industrial applications, deep learning models that power recommendation and ranking have complex and diverse model architectures. These models are continuously developed and refined by large teams of machine learning engineers, rendering manual optimization infeasible. Consequently, graph-based optimization techniques have become an industry standard for boosting performance, with PyTorch FX transformations leading the charge. These transformations typically rely on a set of human-engineered module-level rewrite rules which are not scalable to diverse model architectures. To address this limitation, we introduce Optimus, a general-purpose model transformation framework built in the PyTorch 2.x (PT2) machine learning compiler. With a concise set of predefined patterns, Optimus applies an efficient greedy search algorithm for pattern matching and replacement, while preserving model semantic. It is designed and implemented as a highly customizable and extensible framework integrated into the PT2 stack. Our evaluation shows that the framework can achieve up to 63% speedup, 6% peak memory reduction, and over 400 second compile time decrease for our industry-scale recommendation models compared to baselines. Optimus is open-sourced together with PyTorch 2.x as a customizable model transformation layer.

cs.PF

On strong algebrability and spaceability of continuous functions and fractal dimensions

In this paper, we investigate the strong algebrability and $(\alpha,\beta)$-lineability/spaceability of continuous functions with prescribed fractal dimensions. For $1< s< r< t\leq2$, we define $$H_s[0,1]=\{f\in C[0,1]:{\dim}_HG_f([0,1])=s\},$$ $$\underline{B}_r[0,1]=\{f\in C[0,1]:\underline{{\dim}}_BG_f([0,1])=r\}$$ and $$\overline{B}_t[0,1]=\{f\in C[0,1]:\overline{{\dim}}_BG_f([0,1])=t\}.$$ We prove that $H_s[0,1]\cap\underline{B}_r[0,1]\cap\overline{B}_t[0,1]$ is both strongly $\mathfrak{c}$-algebrable and spaceable. This complements recent findings of Bonilla et al. \cite{BFBS}, Esser et al. \cite{EMVVS}, and Liu et al. \cite{LZS}. We prove that for any $1<s\leq t\leq2$, $H_s[0,1]\cap\overline{B}_t[0,1]$ is $(p,\mathfrak{c})$-spaceable for $p=1,2$. We also prove that $H_s[0,1]\cap\overline{B}_t[0,1]$ is $(n,m+n)$-lineable for any $m,n\in\mathbb{N}$, thus complementing the recent work of Liu et al. \cite{LS}.

math.FA

Ferromagnetic broadband sensing of axionlike dark matter

Levitated particles have demonstrated ultrahigh sensitivity to magnetic fields and accelerations owing to their extremely low dissipation. Such systems have strong potential for fundamental physics research, particularly for the detection of axions and axionlike particles, well-motivated dark matter candidates spanning a broad mass range. In this context, both high sensitivity and large bandwidth are essential. Here, we demonstrate a levitated magnet magnetometer based on an engineered double-resonance mode, achieving an effective linewidth at its optimal sensitivity that is approximately three orders of magnitude broader than those of previous approaches. Together with a hard-magnet array that enhances the axion-induced signal and soft-ferromagnetic shielding that suppresses environmental magnetic noise, this system constitutes a hybrid ferromagnetic platform for axionlike dark matter searches. We search for axionlike dark matter through its photon coupling $g_{a\gamma}$ over the $40$-$3000\,\mathrm{Hz}$ frequency range and establish new direct limits in this frequency band. The best sensitivity is achieved near the upper resonance around $276\,\mathrm{Hz}$, where the magnetometer reaches a magnetic-field resolution of $0.7\,\mathrm{fT}$, corresponding to a limit of $g_{a\gamma}\sim10^{-7}\,\mathrm{GeV}^{-1}$. At this frequency, this result improves upon previous direct limits by more than four orders of magnitude. The demonstrated high-bandwidth levitated sensor may also enable a broad range of applications, including biological sensing and precision measurements.

hep-ex