arXiv ScienceSearch

arXiv subjects

Naoki Sato

Publications and source records attributed to Naoki Sato.

At least 19 recordsLinked to original sources

Compact Core--Shell Equilibria in Gravitational Vlasov--Poisson Systems with Positive and Negative Mass

We construct regular, spherically symmetric, compact energy-cutoff equilibria for gravitational Vlasov--Poisson systems sourced by positive and negative mass distributions in the context of stellar dynamics. Motivated by Bondi's notion of negative mass and its relation to the weak field limit of Einstein gravity with a signed mass--energy source, we show that the internal structure of the steady states depends qualitatively on the type of negative mass present. In the Bondi convention, the signs of inertial, passive gravitational, and active gravitational mass are reversed together, preserving the equivalence principle. The resulting equilibria consist of a central overlap core containing both mass species, surrounded by a finite positive mass shell and an exterior vacuum. Thus, although a pure Bondi negative mass gas cannot form a compact steady state, Bondi negative mass can be spatially confined by a suitable positive mass distribution. In the gravitational-charge convention, only the passive and active gravitational masses change sign, so that the two species obey opposite free-fall laws and unlike masses repel. The resulting equilibria are spatially segregated into a negative mass core, a vacuum gap, a positive mass shell, and an exterior vacuum. For cutoff exponent $n=-1/2$, the radial matching is explicit, while for the regular cutoff $n=1/2$ the matter regions satisfy Lane--Emden-type equations. These results provide a kinetic framework for investigating the distribution of negative mass in astrophysical and cosmological settings.

gr-qc

Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization

Stochastic gradient descent (SGD) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla SGD, particularly with momentum, perform in the presence of heavy-tailed noise? In this paper, we refine existing convergence results for vanilla SGD and, more importantly, provide the first comprehensive convergence analysis of vanilla SGD with momentum for strongly convex, convex, and nonconvex objectives, without employing any gradient control mechanisms. Our results demonstrate that the obtained convergence rates are inferior to the optimal rates achieved by clipped or normalized variants of SGD, thereby revealing inherent limitations of vanilla methods under heavy-tailed noise. The theoretical findings are supported by experiments on synthetic functions.

cs.LG

Admissible Invariant-Torus Foliations for Steady Euler Flows

In 1965, V. I. Arnold established a structure theorem guaranteeing the existence of a foliation by invariant surfaces for general three-dimensional steady Euler flows with non-constant pressure. In this paper, we investigate what foliation structures can arise in steady Euler flows. We consider a toroidal domain foliated by the level sets of a flux function $Ψ$, and prove that every $C^{1}$ steady Euler flow $(\boldsymbol{u},p)$ satisfying the assumptions $ι_{\boldsymbol{u}}dΨ=0$ and $p=p(Ψ)$ admits the tangential flow representation \[\boldsymbol{u}=c_1(Ψ)\boldsymbolξ^{1}+c_2(Ψ)\boldsymbolξ^{2},\] for some lifted solenoidal vector fields $\boldsymbolξ^{1}$ and $\boldsymbolξ^{2}$ associated with a natural basis of weighted harmonic one-forms on the toroidal leaves. Moreover, the flux function $Ψ$ satisfies a single scalar equation, referred to as the normal flux equation. These characterizations reveal the general foliation structure of steady Euler flows, with the Clebsch representation and the Grad--Shafranov equation recovered as the axisymmetric special case.

math.AP

Structural Requirements for Ion-Acoustic Double Layers: A Parametric Perturbation Analysis of the Maxwellian Limit

Standard Maxwellian plasmas exhibit a mathematical \textit{rigidity}, possessing insufficient degrees of freedom to support electrostatic double layers (DLs) and yielding only soliton solutions. This study investigates the hypothesis that the formation of DLs is a generic consequence of breaking this structural rigidity through parametric perturbation. By introducing two independent continuous control parameters, $δ_1$ and $δ_2$, into the electron distribution, we demonstrate that DLs are a structural property of any plasma model that relaxes the strict Maxwellian constraint. Through a Gardner small-amplitude expansion, we analytically prove that a perturbation must modify both the quadratic and cubic density coefficients to decouple the nonlinear structure and generate physical, supersonic double layers, deriving small-amplitude acoustic-limit threshold conditions of $δ_1 > 1$ and $δ_2 > 7/3$. We show that these theoretical boundaries broaden for large-amplitude, nonlinear structures. By mapping the exact existence regions of DLs in phase space, we demonstrate how higher-order terms relax the weak-amplitude limits, confirming that the Maxwellian state represents a singular point where the DL solution collapses.

physics.plasm-ph

Minus one Homogeneous Euler Flows are Geodesible

In this paper, we study $(-1)$-homogeneous steady solutions to the Euler equations on $\mathbb{R}^n \setminus \{0\}$. In low dimensions $n=2,3$, such flows are known to be essentially trivial. In contrast, we show that in higher dimensions $n \ge 4$, every $(-1)$-homogeneous Euler flow is a geodesible vector field with constant Bernoulli function. Moreover, any $(-1)$-homogeneous geodesible field is induced by a geodesible field on the sphere $\mathbb{S}^{n-1}$. In particular, in the case $n=4$, every $(-1)$-homogeneous Euler flow is obtained as an extension of a Beltrami field on $\mathbb{S}^{3}$.

math.AP

Convergence Bound and Critical Batch Size of Muon Optimizer

Muon, a recently proposed optimizer that leverages the inherent matrix structure of neural network parameters, has demonstrated strong empirical performance, indicating its potential as a successor to standard optimizers such as AdamW. This paper presents theoretical analysis to support its practical success. We provide convergence proofs for Muon across four practical settings, systematically examining its behavior with and without the inclusion of Nesterov momentum and weight decay. We then demonstrate that the addition of weight decay ensures almost-sure boundedness of the parameter and gradient norms -- without relying on the commonly imposed bounded-gradient assumption -- and clarify the interplay between the weight decay coefficient and the learning rate. Finally, we derive a lower bound on the critical batch size for Muon -- the batch size that minimizes the stochastic first-order oracle (SFO) complexity of training. Because the resulting formula involves problem-dependent quantities that are not directly observable (gradient variance, target precision, effective rank), it does not predict the critical batch size in absolute terms; rather, it reveals how the hyperparameters $β$ (momentum) and $λ$ (weight decay) govern the qualitative scaling of this value. Our experiments validate these hyperparameter-dependent predictions across workloads including image classification and language modeling.

cs.LG

Quantum--Fluid Correspondence for Systems of Nonrelativistic Spin-$\frac{1}{2}$ Particles

We show that a charged fluid endowed with an internal spin degree of freedom naturally satisfies the Pauli equation for a nonrelativistic spin-1/2 particle, and that a collection of n such interacting fluids can be reformulated as an Euler flow in 3n dimensions, thereby providing a natural representation of a system of n Pauli particles. These results provide a fluid-mechanical derivation of the Pauli equation and extend the Madelung, or quantum-hydrodynamic, picture to many-particle quantum systems. In particular, they imply that an n-qubit quantum computer can, at least in principle, be realized as a suitable combination of n fluids, or equivalently as a 3n-dimensional Euler flow.

quant-ph

Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach

Deep equilibrium models (DEQs) achieve infinitely deep network representations without stacking layers by exploring fixed points of layer transformations in neural networks. Such models constitute an innovative approach that achieves performance comparable to state-of-the-art methods in many large-scale numerical experiments, despite requiring significantly less memory. However, DEQs face the challenge of requiring vastly more computational time for training and inference than conventional methods, as they repeatedly perform fixed-point iterations with no convergence guarantee upon each input. Therefore, this study explored an approach to improve fixed-point convergence and consequently reduce computational time by restructuring the model architecture to guarantee fixed-point convergence. Our proposed approach for image classification, Lipschitz multiscale DEQ, has theoretically guaranteed fixed-point convergence for both forward and backward passes by hyperparameter adjustment, achieving up to a 4.75$\times$ speed-up in numerical experiments on CIFAR-10 at the cost of a minor drop in accuracy.

cs.LG

Degenerate Soft Modes and Selective Condensation in BaAl$_2$O$_4$ via Inelastic X-ray Scattering

BaAl$_2$O$_4$ is a ferroelectric material that exhibits structural quantum criticality through chemical composition tuning. Although theoretical calculations and several diffraction experiments have suggested the involvement of a soft mode in its ferroelectric structural phase transition, direct experimental verification is still lacking. In this study, we successfully observed two soft modes of BaAl$_2$O$_4$ using x-ray inelastic scattering, providing direct experimental evidence for their role in the structural phase transition. Furthermore, we reveal that the soft modes at the M and K points are nearly degenerate in energy, indicating a delicate balance in which either mode could potentially freeze. The K-point mode simultaneously softens toward the transition temperature ($T_{\rm C}$) in a manner nearly identical to the M-point mode. However, the phase transition condenses only at the M point, with the M-point mode stabilizing as an acoustic mode in the low-temperature structure and the K-point mode hardening as temperature decreases.

cond-mat.mtrl-sci

Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization

The graduated optimization approach is a method for finding global optimal solutions for nonconvex functions by using a function smoothing operation with stochastic noise. This paper makes three contributions regarding graduated optimization. First, we extend the definition of function smoothing that is traditionally achieved through convolution with Gaussian noise and characterize for the first time function smoothing with heavy-tailed noise. Second, we show that light- or heavy-tailed stochastic noise in stochastic gradient descent (SGD) has the effect of smoothing the objective function, the degree of which is determined by the learning rate, batch size, and the moment of the stochastic noise. Using this finding, we propose and analyze a new graduated optimization algorithm that varies the degree of smoothing by varying the learning rate and batch size. Third, we relax the $σ$-nice property, a standard but restrictive condition in the analysis of graduated optimization. Our refinement enables convergence guarantees for a broader class of non-convex functions, thereby bridging the gap between theoretical assumptions and practical optimization landscapes.

cs.LG

A Collision Operator for Field-Mediated Interactions in General Relativistic Kinetic Theory

We develop a Hamiltonian framework for general relativistic kinetic theory on the cotangent bundle $T^{\ast}M$ of a Lorentzian (pseudo-Riemannian) manifold. Starting from the geodesic Hamiltonian $H$, we derive a Landau-type collision operator for self-gravitating particles undergoing binary interactions mediated by an arbitrary potential energy $V$, and couple the resulting kinetic stress-energy to the Einstein field equations to obtain the Landau-Einstein system. In the presence of a coordinate-time Killing symmetry we find a family of stationary states of the form $f \propto γ\exp[-β(H+Φ)]ζ(p_0)$, where $Φ$ is the mean field, $γ=dt/dτ$, $β$ is an inverse-temperature parameter, and $ζ$ encodes symmetry-induced degeneracy.

gr-qc

Topological Invariants in Higher-Dimensional Magnetohydrodynamics

It is well known that the three-dimensional ideal magnetohydrodynamics (MHD) equations possess three magnetic invariants: (M) magnetic helicity, (C) cross helicity, and (P) the mean-square magnetic potential, in addition to the fundamental invariants of fluid motion. In this paper we construct higher-dimensional generalizations of these invariants for ideal MHD. Specifically, we identify generalized magnetic helicity and generalized cross helicity in all odd spatial dimensions $n=2m+1$, and families of invariants given by integrals of arbitrary functions of the scalar density $B^m/ν$ of the magnetic field $2$-form $B$, where $B^m$ denotes its $m$-fold wedge product and $ν$ the fluid-density top form, in all even spatial dimensions $n=2m$. We further establish the existence of invariants for symmetric solutions in arbitrary dimensions, generalizing the mean-square magnetic potential and showing that this invariant arises from symmetry rather than from even dimensionality, in contrast to the enstrophy invariant of the two-dimensional Euler equations.

math-ph

Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent

For nonconvex objective functions, including those found in training deep neural networks, stochastic gradient descent (SGD) with momentum is said to converge faster and have better generalizability than SGD without momentum. In particular, adding momentum is thought to reduce stochastic noise. To verify this, we estimated the magnitude of gradient noise by using convergence analysis and an optimal batch size estimation formula and found that momentum does not reduce gradient noise. We also analyzed the effect of search direction noise, which is stochastic noise defined as the error between the search direction of the optimizer and the steepest descent direction, and found that it inherently smooths the objective function and that momentum does not reduce search direction noise either. Finally, an analysis of the degree of smoothing introduced by search direction noise revealed that adding momentum offers limited advantage to SGD.

cs.LG

Scattering Theory in Noncanonical Phase Space: A Drift-Kinetic Collision Operator for Weakly Collisional Plasmas

After developing a scattering theory for grazing collisions in general noncanonical phase spaces, we introduce a guiding center collision operator in five-dimensional phase space designed for plasma regimes characterized by long wavelengths (relative to the Larmor radius), low frequencies (relative to the cyclotron frequency), and weak collisionality (where repeated Coulomb collisions induce cumulatively small changes in particle magnetic moment). The collision operator is fully determined by the noncanonical Hamiltonian structure of guiding center dynamics and exhibits a metriplectic structure, ensuring the conservation of particle number, momentum, energy, and interior Casimir invariants. It also satisfies an H-theorem, allowing for deviations from Maxwell-Boltzmann statistics due to the nontrivial kernel of the noncanonical guiding center Poisson tensor, spanned by the magnetic moment. We propose that this collision operator and its underlying mathematical structure may offer valuable insights into the study of turbulence, transport, and self-organizing phenomena in both laboratory and astrophysical plasmas.

physics.plasm-ph

MHS equilibria in the non-resistive limit to the randomly forced resistive magnetic relaxation equations

We consider randomly forced resistive magnetic relaxation equations (MRE) with resistivity $κ>0$ and a force proportional to $\sqrtκ\ $ on the flat $d$-torus $\mathbb{T}^{d}$ for $d\geq 2$. We show the path-wise global well-posedness of the system and the existence of the invariant measures, and construct a random magnetohydrostatic (MHS) equilibrium $B(x)$ in $H^{1}(\mathbb{T}^{d})$ with law $D(B)=μ$ as a non-resistive limit $κ\to 0$ of statistically stationary solutions $B_κ(x,t)$. For $d=2$, the measure $μ$ does not concentrate on any compact sets in $H^{1}(\mathbb{T}^{2})$ with finite Hausdorff dimension. In particular, all realizations of the random MHS equilibrium $B(x)$ are almost surely not finite Fourier mode solutions.

math.AP

Axisymmetric Dynamos Sustained by a Modified Ohm's Law in a Toroidal Volume

This work tackles a significant challenge in dynamo theory: the possibility of long-term amplification and maintenance of an axisymmetric magnetic field. We introduce a novel model that allows for non-trivial axially-symmetric steady-state solutions for the magnetic field, particularly when the dynamo operates primarily within a ``nearly-spherical'' toroidal volume inside a fluid shell surrounding a solid core. In this model, Ohm's law is generalized to include the dissipative force, arising from electron collisions, that tends to align the velocity of the shell with the rotational speed of the inner core and outer mantle. Our findings reveal that, in this context, Cowling's theorem and the neutral point argument are modified, leading to magnetic energy growth for a suitable choice of toroidal flow. The global equilibrium magnetic field that emerges from our model exhibits a dipolar character. The central insight of the model developed here is that if an additional force is incorporated into Ohm's law, symmetric dynamos become possible.

astro-ph.EP

Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks

A scaled conjugate gradient method that accelerates existing adaptive methods utilizing stochastic gradients is proposed for solving nonconvex optimization problems with deep neural networks. It is shown theoretically that, whether with constant or diminishing learning rates, the proposed method can obtain a stationary point of the problem. Additionally, its rate of convergence with diminishing learning rates is verified to be superior to that of the conjugate gradient method. The proposed method is shown to minimize training loss functions faster than the existing adaptive methods in practical applications of image and text classification. Furthermore, in the training of generative adversarial networks, one version of the proposed method achieved the lowest Frechet inception distance score among those of the adaptive methods.

cs.LG

Explicit and Implicit Graduated Optimization in Deep Neural Networks

Graduated optimization is a global optimization technique that is used to minimize a multimodal nonconvex function by smoothing the objective function with noise and gradually refining the solution. This paper experimentally evaluates the performance of the explicit graduated optimization algorithm with an optimal noise scheduling derived from a previous study and discusses its limitations. It uses traditional benchmark functions and empirical loss functions for modern neural network architectures for evaluating. In addition, this paper extends the implicit graduated optimization algorithm, which is based on the fact that stochastic noise in the optimization process of SGD implicitly smooths the objective function, to SGD with momentum, analyzes its convergence, and demonstrates its effectiveness through experiments on image classification tasks with ResNet architectures.

cs.LG