arXiv ScienceSearch

SEARCH · arXiv Science

Results for “math.CA”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,337 records · Page 4Linked to original sources

From Tsallis to KL: Convergence and Error Estimates for Tsallis-Regularized Optimal Transport

We study the Tsallis-to-Kullback--Leibler (KL) limit for entropy-regularized optimal transport with nonnegative bounded continuous costs. Fixing the regularization parameter $\varepsilon > 0$, we first derive an exact variational reformulation of Tsallis-regularized optimal transport in terms of the Tsallis information projection onto the set of couplings. The formula isolates an explicit correction term and thereby explains why, unlike in the KL case, the regularized transport problem and the corresponding information projection problem do not coincide exactly. We also establish existence and uniqueness for the Tsallis information projection. We then prove, with respect to the narrow topology, the $Γ$-convergence of the Tsallis-regularized functionals to the KL-regularized functional as $q\downarrow1$, together with narrow convergence of their unique minimizers. Finally, we obtain explicit error estimates of order $O(q-1)$ for both the regularized optimal transport values and the associated information projection values. These results quantify the passage from Tsallis regularization to the classical KL setting and clarify the relation between entropic regularization and information projection for $1 < q \leq 2$.

cs.IT

A Complete Characterization of Tensorizable $f$-divergences

Csiszar's formulation of the $f$-divergence introduced a vast family of functionals for quantifying dissimilarity between probability distributions. However, many applications in statistics and information theory rely only on a few $f$-divergences, such as the Kullback-Leibler divergence, the $χ^2$-divergence, and the squared Hellinger distance. These divergences are especially useful because they admit simple compositional formulas under product measures, a property sometimes referred to as tensorization. In this work, we refine a formalism of tensorization previously introduced in the literature. Then, we show that any possible tensorization formula has a multi-affine form characterized by a single parameter, and identify all tensorizable $f$-divergences under our adopted notion of tensorization.

cs.IT

No information transmission through quantum channels above capacity

We show that the capacity of a quantum channel demarcates a phase transition: while reliable transmission below capacity is always possible, any attempt to transmit information above it fails catastrophically. Specifically, we prove exponential strong converse theorems for unassisted quantum and classical communication over arbitrary finite-dimensional memoryless quantum channels. At rates beyond the respective capacity, the entanglement-generation fidelity and the success probability for classical communication decay exponentially with the number of channel uses. This rules out transmission above capacity even when one tolerates arbitrarily large errors. Our proof follows the classical Arimoto strategy, augmented by a crucial new ingredient: integral representations of Rényi information measures that lead to asymptotic continuity bounds for Rényi capacities.

quant-ph

Spatial symmetry invariance of solution of Kolmogorov flow

We prove a mathematical theorem that solution for all $t > 0$ of the two-dimensional (2D) Kolmogorov flow governed by Navier-Stokes (NS) equations with periodic boundary condition keeps the same spatial symmetry as its smooth initial condition. The proof of a similar theorem for the three-dimensional NS equations is given in the appendix. These mathematical theorems can be used to check the correctness and reliability of numerical simulations of NS turbulence. For example, they support the corresponding CNS (clean numerical simulation) results of the 2D and 3D turbulent Kolmogorov flows [1-3] that remain the same spatial symmetry in the whole time interval of simulation, but do not support the corresponding DNS (direct numerical simulation) results that lose the spatial symmetry quickly. In other words, these DNS results violate these mathematical theorems. Thus, these mathematical theorems rigorously confirm that the spatiotemporal trajectories of NS turbulence given by DNS are indeed quickly polluted by numerical noises badly. All of these indicate that CNS can indeed provide helpful enlightenments to deepen our understanding about turbulence and besides approach some mathematical truths about NS equations.

physics.flu-dyn

A simple derivation of the Kalman filter

In this lecture note, we present a concise and self-contained derivation of the discrete-time Kalman filter equations that requires only a basic understanding of least squares estimation. The treatment is designed to minimize mathematical overhead while preserving both rigor and generality.

math.OC

Efficient Polynomial-Time Decoding of Simplicial Anticodes with Near-Optimal Performance

In this work, we propose an efficient decoding algorithm for codes arising from simplicial complexes, a family of binary linear codes for which no decoding method of this type was previously known. Although the algorithm does not always attain the maximum theoretical error-correcting capability, it provides an explicit bound that can be computed directly from the structure of the complex. Moreover, this bound is asymptotically optimal: the ratio between the guaranteed correcting capability and the theoretical maximum converges to $1$ as the code length increases, under natural assumptions on the dimension of the maximal faces. The correction capability is also presented in specific examples. Finally, we introduce specific families of simplicial complexes where the algorithm successfully reaches this theoretical bound.

cs.IT

Neural operators approximate strongly continuous convex monotone semigroups

We approximate strongly continuous convex monotone semigroups by learning their Chernoff-type one-step operators with neural operators. First, we introduce the general class of so-called Chernoff-neural operators and show in a universal approximation theorem that they can approximate the Chernoff one-step operators arbitrarily well. By using stability estimates between weighted Hölder spaces, the one-step approximation error can be propagated through the iterations which yields universal approximation of the corresponding semigroup. Second, we introduce the more specialized class of envelope-neural operators for envelope semigroups which allows us to derive quantitative approximation rates. Finally, we illustrate the effectiveness of these neural operators in several numerical examples arising from non-linear partial differential equations, stochastic optimal control and stochastic processes under model uncertainty.

math.NA

Summing the sum of digits

We revisit and generalize inequalities for the summatory function of the sum of digits in a given integer base. We prove that several known results can be deduced from a theorem in a 2023 paper by Mohanty, Greenbury, Sarkany, Narayanan, Dingle, Ahnert, and Louis, whose primary scope is the maximum mutational robustness in genotype-phenotype maps.

math.NT

On the Gram matrix of standard inner products of asymmetrically-weighted Hermite functions

Let A denote the infinite Gram matrix associated with the standard L2 inner product of asymmetrically-weighted (AW) Hermite functions. We derive an explicit representation of its entries and its Cholesky factorization. We further show that this factorization admits a natural interpretation on a scaled Bargmann-Fock basis. An explicit formula for the inverse of A is also obtained. We then consider the corresponding finite Gram matrix and analyze its asymptotic property, as well as that of its Schur complement. The analysis is motivated by numerical methods for plasma physics, in particular Galerkin spectral methods applied to the Vlasov-Poisson (VP) system. As an application, we demonstrate how the derived Gram matrix formulas and asymptotic results can be exploited in the analysis and implementation of a Galerkin spectral method for the VP system.

math.NA

Oracle-free Boltzmann Sampling for Powersets

We propose an approach for sampling powersets under the Boltzmann distribution in an oracle-free way, i.e. without numerically evaluating the associated generating function. Our approach relies on a Poissonised infinite occupancy model and thinning. It yields an explicit sampler for bounded counting sequences and extends under mild growth conditions. We implement the sampler and find runtimes comparable to existing Boltzmann samplers.

cs.DM

Momentum-based gradient descent methods for Lie groups

Polyak's Heavy Ball (PHB; Polyak, 1964), a.k.a. Classical Momentum, and Nesterov's Accelerated Gradient (NAG; Nesterov, 1983) are well-established momentum-descent methods for optimization. Although the latter generally outperforms the former, primarily, generalizations of PHB-like methods to nonlinear spaces have not been sufficiently explored in the literature. In this paper, we propose a generalization of NAG-like methods for Lie group optimization. This generalization is based on the variational one-to-one correspondence between classical and accelerated momentum methods (Campos et al., 2023). We provide numerical experiments for chosen retractions on the group of rotations based on the Frobenius norm and the Rosenbrock function to demonstrate the effectiveness of our proposed methods, and that align with results of the Euclidean case, that is, a faster convergence rate for NAG.

math.OC

A Note on Sphere Packing Bounds for Tuple Lattice Sieving

A finite set of unit vectors is $k$-irreducible if every signed sum of between two and $k$ distinct elements has norm greater than one. Let $\mathcal{R}_k$ be the maximal asymptotic rate of such sets, and let $κ(α)$ be the maximal asymptotic rate of spherical codes with pairwise inner products at most $α$. For $k \ge 2$ we show: \begin{align} \mathcal{R}_k \le \min_{1 \le r \le \lfloor k/2 \rfloor} \frac{1}{r} \, κ\!\left(1 - \frac{1}{2r}\right) \, . \end{align} Combining this with standard sphere packing bounds, for large $k$ we obtain an almost-tight asymptotic comparison with the known lower bounds: \begin{align} \left(\tfrac{1}{2}-o(1)\right) \, \frac{\log_2 k}{k} \le \mathcal{R}_k \le (1 + o(1)) \, \frac{\log_2 k}{k} \, . \end{align}

math.CO

Degenerating orbits of the Longest Edge Bisection process

We study the Longest Edge Bisection (LEB) process as a dynamical system on the projective shape space of simplices. A long-standing conjecture going back to Adler and Rivara-Levin and motivated by finite-element mesh refinement, often taken as a standing assumption, is that this procedure is non-degenerate and, in fact, in a certain way periodic. We prove: \begin{itemize} \item There are 3-dimensional simplices such that the longest edge-bisection algorithm degenerates. \item There is an open set of 4-dimensional simplices on which the longest edge-bisection algorithm degenerates. \item If parametrizing the space of $d$-dimensional simplices by independent standard Gaussian vectors, then as $d$ increases, a random simplex degenerates asymptotically almost surely. \end{itemize} This is realized through exhibiting hyperbolic behaviour of the LEB process. We also exhibit elliptic behaviour that is nonperiodic.

math.DS

A \(3\times 3\) counterexample to Lin and Wimmer's rank-minimization conjecture associated with Roth's similarity theorem

We give a \(3\times 3\) counterexample, valid over every field, to a rank-minimization conjecture of Lin and Wimmer (Bull. Aust. Math. Soc., 84 (3) (2011), 441--443) related to Roth's similarity theorem for the Sylvester matrix equation. We also prove that, over the complex field, no counterexample can occur when one of the two matrix sizes is less than \(3\). Hence the example is dimensionally minimal over the complex field.

math.RA

Riccati Stability Without Auxiliary Matrices

In the 2004 collection \emph{Unsolved Problems in Mathematical Systems and Control Theory}, Erik Verriest posed the problem of characterizing Riccati stability ``without invoking additional matrices.'' We give such a characterization through a scalar invariant of the resolvent family. The invariant is defined by covariance balances using at most $n^2+1$ frequency--direction pairs. A finite-dimensional separation argument shows that this covariance radius equals the optimal common ellipsoidal norm of the resolvent family. Combined with the strict bounded real lemma, this yields a necessary and sufficient condition for Riccati stability involving only the Hurwitz property of the first matrix and a single scalar inequality. The invariant reduces to the ordinary spectral radius for one matrix, lies between the pointwise spectral-radius and unscaled small-gain levels of the resolvent family, and coincides with the operator-space spectral radius of Shalit and Shamovich for the natural resolvent function space.

math.OC