arXiv ScienceSearch

SEARCH · arXiv Science

Results for “math.OA”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

929 records · Page 6Linked to original sources

DOFFO_TR: a Decentralized Objective Function-Free Optimization method with Trust-Region

In this paper, we propose a novel objective function-free trust-region method designed to solve optimization problems over decentralized networks. Unlike traditional approaches that often rely on stepsize tuning, our framework employs a function-free trust-region procedure that enables adaptive selection of the step length. Our approach accommodates first- and second-order models and eliminates the need to share local function values and gradients among agents, thereby enhancing privacy and computational efficiency. On the theoretical side, we establish provable iteration complexity guarantees that, for some variants, match those established for classical centralized trust-region methods. Numerical evaluations demonstrate that our approach achieves a favorable trade-off between performance and efficiency, requiring only moderate communication overhead compared to state-of-the-art methods in the literature.

math.OC

Rigorous Error Certification for Neural PDE Solvers: From Empirical Residuals to Solution Guarantees

Uncertainty quantification for partial differential equations is traditionally grounded in discretization theory, where solution error is controlled via mesh/grid refinement. Physics-informed neural networks fundamentally depart from this paradigm: they approximate solutions by minimizing residual losses at collocation points, introducing new sources of error arising from optimization, sampling, representation, and overfitting. As a result, the generalization error in the solution space remains an open problem. Our main theoretical contribution establishes generalization bounds that connect residual control to solution-space error. We prove that when neural approximations lie in a compact subset of the solution space, vanishing residual error guarantees convergence to the true solution. We derive deterministic and probabilistic convergence results and provide certified generalization bounds translating residual, boundary, and initial errors into explicit solution error guarantees.

cs.LG

Why Multi-Layer Message Passing Works: Completeness Theory for Graph Neural Network Interatomic Potentials

We prove that the Hypergraph Neural Network, an invariant architecture with 3-body message passing, is a universal approximator for potential energy surfaces. Our main contribution is a multi-layer completeness theory. We show that $L$ layers of message passing on sparse, cutoff-based graphs achieve the same representational power as having access to the full $L$-hop neighborhood, provided the configurations are generic, satisfy an overlap condition and a connectivity condition. This provides the first rigorous justification for the common practice of using multi-layer message passing with a per-layer cutoff smaller than the physical interaction range, the setting used by virtually all practical graph neural network based machine-learned interatomic potentials. As immediate consequences, we show that both DPA3 and CHGNet architectures inherit universal approximation.

cs.LG

Numerical Ergodicity and Uniform Estimate of Monotone SPDEs Driven by Multiplicative Noise

We analyze the long-time behavior of numerical schemes for a class of monotone stochastic partial differential equations (SPDEs) driven by multiplicative noise. By deriving several time-independent a priori estimates for the numerical solutions, combined with the ergodic theory of Markov processes, we establish the exponential ergodicity of these schemes with a unique invariant measure, respectively. Applying these results to the stochastic Allen--Cahn equation indicates that these schemes always have at least one invariant measure, respectively, and converge strongly to the exact solution with sharp time-independent rates. We also show that these numerical invariant measures are exponentially ergodic and thus give an affirmative answer to a question proposed in (J. Cui, J. Hong, and L. Sun, Stochastic Process. Appl. (2021): 55--93), provided that the interface thickness is not too small.

math.NA

Bernstein--von Mises theorems for Bayesian probabilistic numerics

We study probabilistic numerical methods for solving nonlinear PDEs from a Bayesian nonparametric perspective. Given noisy evaluations at random collocation points, we place a truncated Gaussian series prior on the unknown solution and establish contraction at the minimax nonparametric rate, up to a logarithmic factor. Our main results give Gaussian approximations of the posterior in positive-order Sobolev spaces and, under suitable conditions, in the uniform topology. This contrasts with classical ill-posed inverse problems, where Bernstein--von Mises theorems typically require substantially weaker topologies. Here, the observation operator is differential rather than smoothing, and inversion of its linearisation gains regularity, making these strong-topology results possible. The posterior may be centred at either the posterior mean or the posterior mode. We further prove that the Gaussian Laplace approximation is asymptotically equivalent to the true posterior at a $\sqrt{N}$-scale.

math.ST

Eleven, twelve, and thirteen lonely runners

Wills conjectured that, for any non-zero integers $u_1,\ldots,u_k$, there is a real number $t$ such that, for all $i=1,\ldots,k$, \[\lVert tu_i\rVert\geq\frac{1}{k+1},\] where $\lVert x\rVert$ is the distance from $x$ to the closest integer. This statement is known as the Lonely Runner Conjecture. A computational method developed by Rosenfeld and the second author verified the conjecture for $k\leq9$. We further refine this method with new sieving techniques and employ a polynomial method argument to show that any $(u_1,\ldots,u_k)\equiv(1,2,\ldots,k)\pmod{p}$ with $\gcd(u_1,\ldots,u_k)=1$ satisfies the conjecture when $k+1$ and $p > k^2+k$ are both odd primes. Ultimately, we provide a computer-assisted proof of the Lonely Runner Conjecture for $k\in\{10,11,12\}$.

math.CO

On the Equality of the ELBO to a Sum of Entropies at Stationary Points of Learning

The variational lower bound (a.k.a. ELBO or free energy) is the central objective for many established as well as for many novel algorithms for unsupervised learning. Such algorithms usually increase the bound until parameters have converged to values close to a stationary point of the learning dynamics. Here we show that (for a very large class of generative models) the variational lower bound is at all stationary points of learning equal to a sum of entropies. Concretely, for standard generative models with one set of latents and one set of observed variables, the sum consists of three entropies: (A) the (average) entropy of the variational distributions, (B) the negative entropy of the model's prior distribution, and (C) the (expected) negative entropy of the observable distribution. The obtained result applies under realistic conditions including: finite numbers of data points, at any stationary point (including saddle points) and for any family of (well behaved) variational distributions. The class of generative models for which we show the equality to entropy sums contains many standard as well as novel generative models including standard (Gaussian) variational autoencoders. The prerequisites we use to show equality to entropy sums are relatively mild. Concretely, the distributions defining a given generative model have to be of the exponential family, and the model has to satisfy a parameterization criterion (which is usually fulfilled). Proving equality of the ELBO to entropy sums at stationary points (under the stated conditions) is the main contribution of this work.

stat.ML

A multi-class kinetic traffic flow model: discrete-velocity formulation and diffusively-corrected macroscopic limits

This paper introduces a multi-class extension of a discrete-velocity kinetic traffic flow model based on a non-local Prigogine-Herman framework. We derive a hyperbolically scaled system of equations from a continuous kinetic formulation describing interactions between different vehicle classes through braking and relaxation terms. The model is then discretized with respect to the velocity variable for an arbitrary number of vehicle classes, and the structural properties of the resulting formulation are analyzed. In particular, we prove hyperbolicity and total linear degeneracy. Due to the non-conservative structure of the model, we employ a path-conservative finite volume scheme for the numerical approximation of the system. Finally, we derive the corresponding diffusively-corrected macroscopic multi-class model, investigate its stability and present numerical simulations on a single-lane road to illustrate the theoretical findings.

math.NA

Asymptotic Bounds on Generalized Covering Radii of Binary Primitive BCH Codes

Fix integers $e\ge2$ and $r\ge1$. In this paper we study the $r$-th generalized covering radius $ρ_r\left(BCH(e,m)\right)$ of the binary primitive $e$-error-correcting BCH code $BCH(e,m)$. By using an algebraic-geometric reformulation of the covering problem together with an explicit Lang-Weil estimate, we prove that \[ρ_r\bigl(\BCH(e,m)\bigr)\le(r+1)e-1\] for all sufficiently large $m$. For $e\ge7$, this improves a recent result of Belinsky--Zabokritskiy. Our proof gives a substantially simpler geometric approach to this upper bound. In particular it implies that \[ρ_2\bigl(BCH(e,m)\bigr)=3e-1\] for all sufficiently large $m$. Previously it was only known that \[ρ_2\bigl(\BCH(e,m)\bigr) \in \left\{3e-1,3e\right\}\] for all sufficiently large $m$.

cs.IT

Moment-enhanced shallow-water equations with an effective wall closure for no-slip bottoms

Shallow-water equations and low-order shallow-water moment models use vertically coarse representations and therefore cannot, in general, resolve the thin wall-affected region produced by a no-slip bottom. Enforcing the pointwise wall value on a low-order global polynomial reconstruction can introduce stiff relaxation and distort the resolved interior velocity profile. Starting from the incompressible Navier--Stokes equations with Navier bottom friction, we derive a bottom-to-mean relation in a distinguished regular-friction regime and use it to define an endpoint-consistent effective wall-traction closure for the shallow-water equations and the hyperbolic shallow-water moment equations. The closure represents the momentum effect of unresolved near-wall dynamics; it neither resolves the physical boundary layer nor imposes the pointwise no-slip trace on the reconstructed polynomial. It recovers the perfect-slip wall contribution when the friction coefficient vanishes. Because only source terms are changed, the homogeneous principal matrices and their established two-dimensional hyperbolicity classification remain unchanged. We compare the standard and modified reduced models with two-phase incompressible Navier--Stokes computations in OpenFOAM for wet-bed dam-break and three-dimensional collapse tests. In the cases considered, the modified closure reduces the excessive damping of the classical low-order wall source and improves agreement in depth-averaged and resolved-interior velocity diagnostics, but it does not uniformly improve front-propagation speed. The regular-friction asymptotic remainder is not uniform in the large-friction numerical regime; there the effective coefficient is used as a wall-model continuation and assessed empirically.

math.NA

Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU

In this work, we study the nonlinear dynamics of a shallow neural network trained with mean-squared loss and leaky ReLU activation. Under Gaussian inputs and equal layer width k, (1) we establish, based on the equivariant gradient degree, a theoretical framework, applicable to any number of neurons k>= 4, to detect bifurcation of critical points with associated symmetries from global minimum as leaky parameter $α$ varies. Typically, our analysis reveals that a multi-mode degeneracy consistently occurs at the critical number 0, independent of k. (2) As a by-product, we further show that such bifurcations are width-independent, arise only for nonnegative $α$ and that the global minimum undergoes no further symmetry-breaking instability throughout the engineering regime $α$ in range (0,1). An explicit example with k=5 is presented to illustrate the framework and exhibit the resulting bifurcation together with their symmetries.

math.OC

Breakdown of Edgeworth Expansion in Finite-Blocklength Regime and Exact Absorption via $q$-Deformation

This paper addresses the structural breakdown of the Edgeworth expansion in the finite-blocklength (FBL) regime, where conventional asymptotic approximations yield unphysical negative probabilities in the deep-tail region. We propose a $q$-deformed framework that resolves this inconsistency by replacing additive polynomial perturbations with a geometric deformation of the information density space. Motivated by the linearization of nonlinear dynamics, we prove that dynamically scaling the $q$-logarithmic parameter exactly absorbs the third-order skewness while preserving global nonnegativity. We establish a universal asymptotic matching, demonstrating that the framework encapsulates higher-order asymptotic scales. Numerical results confirm that the proposed method matches the state-of-the-art precision of the Cornish-Fisher bound without the risk of negative probabilities. The framework offers a robust and computationally stable foundation for evaluating operational limits in ultra-reliable communications such as 6G and URLLC.

cs.IT

Prove2Me: An Open Collaborative Platform for Scaling Math Formalization

Proof assistants such as Lean 4 promise the paradigm of formally verified mathematics, but large-scale formalization projects have faced major barriers to entry, including the need for expertise in formal verification (as well as the underlying mathematics) and the significant time required for writing formal proofs. AI coding agents have dramatically reduced these barriers; human users can now use natural language to prompt agents to write complex proofs in Lean. This opens up the intriguing possibility of internet-scale mathematical collaboration involving both humans and AI agents, where correctness is machine-checked. To realize this possibility, we introduce Prove2Me (https://prove2.me), an open collaborative platform for formalizing mathematics. Users launch formalization "missions", to which AI agents contribute formal proofs toward completion. We designed mechanisms and a specialized harness in Prove2Me that enable large-scale collaboration so that agents can build on one another's work and freely reuse existing results. In doing so, Prove2Me aims to turn math formalization into a scalable, crowd-sourced effort open to anyone with an agent.

cs.AI

Universal Approximation of Nonlinear Operators and Their Derivatives

Establishing Universal Approximation Theorems (UATs) for nonlinear operators and their derivatives is a foundational open problem in Operator Learning (OL) and raises delicate questions in Nonlinear Functional Analysis. We prove the first UATs for $k$-times differentiable nonlinear operators and their derivatives via OL architectures, uniformly on compact sets and in weighted Bastiani--Sobolev spaces for general finite input measures. In full Banach-space generality, these are the first complete generalizations of the corresponding influential classical UATs in [Hornik, 1991] to infinite-dimensional spaces and OL, {and launch Derivative-Informed Operator Learning (DIOL) (i.e. learning nonlinear operators and their derivatives)} on general Banach spaces. Based on our UATs, we formulate Bastiani--Sobolev training in DIOL. We present open frontiers where DIOL and our UATs find applications: high-order accuracy in OL; fast constrained optimization in Banach spaces (e.g. optimal control of PDEs, inverse problems) via Learn-Then-Optimize; numerical methods for infinite-dimensional PDEs (e.g. HJB PDEs on Banach spaces from infinite-dimensional optimal control via Optimize-Then-Learn, such as optimal control of PDEs, SPDEs, path-dependent systems, partially observed systems, mean-field control). We parameterize nonlinear operators via Encoder-Decoder Architectures, classical OL architectures. These include DeepONets, Deep-H-ONets, and PCA-Nets, which our UATs cover. Our UATs are based on (i) Approximation Properties of Banach spaces; (ii) continuous Bastiani differentiability (weaker than continuous Fréchet differentiability); (iii) $C^k_B$ (Bastiani) compact-open topologies; indeed, UA in $C^k$ (Fréchet) compact-open topologies (induced by operator norms) fails; (iv) construction of weighted Bastiani--Sobolev spaces, generalizing classical Gaussian Sobolev spaces on Banach spaces.

cs.LG

Continuous data assimilation in steady Navier-Stokes equations with unknown viscosity: robust and efficient solvers and fast parameter recovery

Recent advances in equation discovery methods such as SINDy have highlighted the growing interest in identifying governing parameters and models directly from data. In this work, we take a complementary approach grounded in analysis and numerical PDE methods: we recover an unknown viscosity in steady Navier-Stokes equations (NSE) from partial incompressible flow observations using continuous data assimilation (CDA). We propose a simple and efficient parameter recovery algorithm and also a nonlinear solver for CDA-NSE. Together, this creates a highly efficient technique for recovering an unknown viscosity from partial solution data. Our analysis establishes the well-posedness of steady CDA-NSE, quadratic convergence of the parameter recovery algorithm, and quadratic convergence of a CDA-Picard + CDA-Newton nonlinear solver. Numerical experiments illustrate that the methods are very effective in restoring parameters quickly, even with poor initial guesses.

math.NA

Improved $\ell_0$-Isoperimetry for Convex Bodies via Mass Transport

We study $\ell_0$ isoperimetry for a convex body $K\subset \mathbb{R}^n$, $n\ge2$. For a Borel set $S\subset K$, let $\partial_0^K S$ be the set of points in $K \setminus S$ that can be reached from $S$ by changing at most one coordinate (i.e. the $\ell_0$ boundary of $S$). Suppose that, for some unconditional convex body $Q \subset \mathbb{R}^n$, numbers $r,R>0$, and possibly different centers $x_0,y_0$, \[ x_0+rQ \subset K\subset y_0+RQ. \] Writing $s=\text{vol}(S)/\text{vol}(K)$, we prove that whenever $0 0$ is an absolute constant. Consequently, the associated $\ell_0$-isoperimetric coefficient is at least $cr/(n^2R)$. Previous direct lower bounds were only known for $\ell_2$ and $\ell_\infty$ regularity whereas our lower bound holds directly for any $Q$-regularity, where $Q$ is an unconditional convex body. Compared to $\ell_2$ and $\ell_\infty$ regularity, our lower bound result improves upon the previously best known lower bounds, for any $s$, by a factor of $n$. As an application of our result, we give improved mixing time bounds for the Coordinate Hit and Run walk (CHAR). Our proof of the lower bound is based on a modification of the method of canonical paths applied to a continuous Hamming graph over our convex body. Our construction of canonical paths can be viewed as a suitable coordinate discretization of certain mass transport maps from $S$ to $S^c$. We also give complementary upper-bounds for any $Q$-regularity, with an overall factor of $n$ gap between the two.

math.FA

Accelerated primal--dual dynamics and algorithms for convex optimization with nonlinear inequality constraints

We consider convex optimization with nonlinear inequality constraints and develop a primal--dual multiplier framework that is consistent in continuous and discrete time. We first propose continuous-time dynamics with Nesterov-type vanishing damping $α/t$, together with suitable extrapolations of the dual variable and the nonlinear constraint mapping. Under convexity assumptions and $α\geq3$, we establish $\mathcal O(t^{-2})$ convergence rates for both nonlinear feasibility and the objective residual. We then derive an inexact accelerated primal--dual algorithm through a compatible discretization of a perturbed version of the dynamics. For composite convex objectives, a weighted summability condition on the primal inexactness yields the $\mathcal O(k^{-2})$ rates for feasibility and the objective residual, thereby matching the accelerated rates of their continuous-time counterparts. To the best of our knowledge, this is the first Nesterov-type primal--dual multiplier framework for convex optimization with nonlinear inequality constraints.

math.OC

The Prime Clockwork: A Dynamic Representation of Modular and Multiplicative Arithmetic

The way numbers are represented strongly influences which arithmetic structures are easy to see. The \emph{prime clockwork} is a recursively growing discrete dynamical system: a list of autonomous two-hand clocks driven by one common $+1$ signal. No primes or primality labels are supplied. Starting empty, the process appends a clock of period $n$ whenever none already present rings; the primes are generated internally as its growth times. For each installed prime $p$, the seconds reading $R_p$ advances through $0,\ldots,p-1$, and each return to zero increments the minutes reading $M_p$, which counts completed $p$-cycles. The hands use only increment, comparison, reset, and carry, without explicit \texttt{mod} or \texttt{div} operations. At time $n$, $n=pM_p(n)+R_p(n)$. The valuation readout $V_p(n)=ν_p(n)$ is generated locally: it is zero when the seconds counter is non-zero (silent state) and otherwise (when the p-clock rings) one plus the earlier valuation addressed by the current minutes reading. The valuation vector gives the integer in unique prime-factorized form. Its coordinates add and subtract under multiplication and division, representing every positive rational uniquely; divisibility becomes weak componentwise order, and unique factorization is natural in this representation. Finite seconds arrays form Cartesian-product state spaces whose common orbit visits every joint state once before repeating; this \emph{grand cycle} is the order-sensitive dynamical counterpart of the Chinese remainder theorem. The same coordinates expose gcd, lcm, perfect powers, Bézout's identity, and Euler's totient. Rational valuation levels reach certain positive algebraic irrationalities, but not algebraic numbers in general.

math.HO