arXiv ScienceSearch

SEARCH · arXiv Science

Results for “math.AT”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

3,733 records · Page 5Linked to original sources

Multiplicative comparisons of Rényi entropies for weighted Bernoulli sums

We establish improved multiplicative bounds relating the Rényi entropies of different orders for weighted sums of independent Bernoulli random variables. In particular, we prove a logarithmic bound between the zeroth-order and infinity-order Rényi entropies, which yields a polynomial improvement over the square-root bound of Jain, Sah, and Sawhney. Additionally, we obtain explicit constant-factor bounds for comparisons among Rényi entropies of nonzero orders.

math.PR

Logarithmic-Free Moment and Generalization Bounds for Uniformly Stable Algorithms

Uniform stability is a classical tool for controlling the generalization error of a learning algorithm. Bousquet, Klochkov, and Zhivotovskiy (2020) showed that the problem can be reduced to a moment inequality for a sum of weakly interacting functions of independent random variables. Their bound contains an additional factor $\log n$, and they asked whether this factor can be removed. We answer this upper-bound question affirmatively. More specifically, let $Z=(Z_1,\ldots,Z_n)$ have independent coordinates and let $g_i(Z)$ satisfy $\mathbb E[g_i(Z)\mid Z_{-i}]=0, \ \left| \mathbb E[g_i(Z)\mid Z_i]\right|\le M, \ \text{for every } i = 1, \dots, n, $ where $Z_{-i}$ denotes all coordinates except $Z_i$. Assume additionally that changing any coordinate $Z_j$, $j\neq i$, changes $g_i$ by at most $β$, we prove that, for every $p\ge2$, for every $p\ge2$, $$ \left\| \sum_{i=1}^n g_i(Z)\right\|_p \le 16pnβ+M\sqrt{2pn}. $$ This removes the $\log n$ factor from the previous bound and matches the lower bound of Bousquet, Klochkov, and Zhivotovskiy up to universal constants in the range covered by their construction. Our proof first establishes the required estimate on the Rademacher cube, then transfers it to arbitrary product distributions by a two-copy randomization argument.

stat.ML

Estimating quantum relative entropies on quantum computers

Quantum relative entropy, a quantum generalization of the renowned Kullback-Leibler divergence, serves as a fundamental measure of the distinguishability between quantum states and plays a pivotal role in quantum information science. Despite its importance, efficiently estimating quantum relative entropy between two quantum states on quantum computers remains a significant challenge. In this work, we propose the first quantum algorithm for directly estimating quantum relative entropy and Petz Renyi divergence from two unknown quantum states on quantum computers, addressing open problems highlighted in [Phys. Rev. A 109, 032431 (2024)] and [IEEE Trans. Inf. Theory 70, 5653-5680 (2024)]. Notably, the circuit size of our algorithm is at most $2n+1$ with $n$ being the number of qubits in the quantum states and it is directly applicable to distributed scenarios, where quantum states to be compared are hosted on cross-platform quantum computers. We prove that our loss function is operator-convex, ensuring that any local minimum is also a global minimum. We validate the effectiveness of our method through numerical experiments and observe the absence of the barren plateau phenomenon. As an application, we employ our algorithm to investigate the superadditivity of quantum channel capacity. Numerical simulations reveal new examples of qubit channels exhibiting strict superadditivity of coherent information, highlighting the potential of quantum machine learning to address quantum-native problems.

quant-ph

Local Identifiability of Networks with Nonlinear Node Dynamics

We study the identifiability of nonlinear network systems with partial excitation and partial measurement when the network dynamics is linear on the edges and nonlinear on the nodes. We assume that the graph topology and the nonlinear functions at the node level are known, and we aim to identify the weight matrix of the graph. Our main result is that, for almost all static analytic nonlinearities that cross the origin, directed graphs are generically locally identifiable if and only if at least one node is excited in every source component of the condensation graph and at least one node is measured in every sink component. This holds even when all other nodes remain unexcited and unmeasured and stands in sharp contrast to most findings on network identifiability requiring measurement and/or excitation of each node. The result applies to homogeneous feed-forward and recurrent artificial neural networks and generalizes previous literature by considering a broader class of activations and architectures.

math.OC

Set Theory in the Foundation of Math; Internal Classes and External Sets

Usual math sets have special types: countable, compact, open, occasionally Borel, rarely projective, etc. Each such set is described by a single set theory formula with parameters unrelated to formulas. Exotic expressions involving sets related to formulas of unbounded quantifier depth appear mostly in esoteric or foundational studies. Recognizing the internal to math (formula-specified) and external (parameter-based) aspects of math objects greatly simplifies foundations. I postulate that external sets (not internally specified, constituting the domain of quantifiable variables) are hereditarily countable and independent of purely formula-defined classes, i.e. with finite algorithmic information about them. Variables for classes are not explicitly quantified. This opens a way to eliminate all non-integer quantifiers in set theory sentences. The restrictions seem to require almost no changes in math papers, only reinterpreting some formalities.

cs.LO

Harmonic higher weight distributions, Simonis' approach of MacWilliams identity and moments

We present a combinatorial proof of Simonis type MacWilliams identity for harmonic higher weight distributions of linear codes. Furthermore, we investigate the statistical moments of the harmonic higher weight enumerators for random linear codes. Defining the enumerators via rank functions of the generator matrices of linear codes, we prove that its expectation vanishes for all non-trivial harmonic functions due to the inherent symmetry of random matrices, and we also derive an explicit, non-trivial formula for the covariance.

math.CO

The half-rate linear programming bound for binary codes is $\frac12-\frac1π$

In their work on sphere packing and the conformal bootstrap, Afkhami-Jeddi, Cohn, Hartman, de Laat, and Tajdini conjectured the exact high-dimensional exponent of the Cohn--Elkies sphere-packing linear program. OpenAI's Chapter 1 subsequently proved their conjecture by establishing that both Fourier sign-uncertainty radii are $(1/π+o(1))\sqrt d$. We prove the binary coding analogue: the half-rate point of the asymptotic binary Delsarte linear program is $1/2-1/π$; equivalently, \[ R_D\!\left(\frac12-\frac1π\right)=\frac12. \] We also formulate the two Krawtchouk sign-uncertainty problems and determine both of their asymptotics. If $A^{\mathrm K}_{\pm}(n)$ denotes the first radial layer after which an origin-vanishing Krawtchouk $(\pm1)$-eigenfunction can be nonnegative, then \[ \frac{A^{\mathrm K}_{\pm}(n)}n\longrightarrow \frac12-\frac1π. \] The common lower bound is the Hamming space counterpart of the mass-concentration principle in the Chapter 1 proof. The upper bound has a different source. It is the binary-code counterpart of the final spherical-code construction in OpenAI's Chapter 2. Gay, Jeronimo, and Liu improved the resulting binary bound and suggested the functional $Φ$ used here, but explicitly evaluated only a few low levels of the corresponding hierarchy. We construct and evaluate compatible binary certificates at every level, attaining the upper bound in the limit. The construction uses an $N$-qubit generalization of the pure-state channel of Alrabiah and Guruswami.

math.CO

Counterexamples to Charpin's Conjecture on BCH codes

We construct an infinite family of $q$-ary primitive narrow-sense BCH codes whose minimum distance strictly exceeds the Bose distance; in fact, the gap between the two can be arbitrarily large as the length of the code tends to infinity. The key idea is to embed these BCH codes in a suitably large punctured generalized Reed--Muller code, whose codeword weights obey divisibility conditions supplied by Ax's theorem. This divisibility forces the minimum distance of the BCH codes far above the Bose distance. In particular, our family disproves a longstanding conjecture of Charpin asserting that this difference is at most four.

cs.IT

A Compositional Kernel Model for Feature Learning

We study a compositional variant of kernel ridge regression in which the predictor is applied to a coordinate-wise reweighting of the inputs. Formulated as a variational problem, this model provides a tractable setting for studying feature learning in compositional architectures. From the perspective of variable selection, we show how relevant variables are recovered while noise variables are eliminated. We prove that both global minimizers and stationary points discard noise coordinates when the noise variables are Gaussian distributed. A central finding is that $\ell_1$-type kernels, such as the Laplace kernel, succeed in recovering features contributing to nonlinear effects at stationary points, whereas Gaussian kernels recover only linear ones.

cs.LG

Shannon's problem on the monotonicity of entropy and a Conjecture of Tao

Let $X_1,X_2,\ldots$ be i.i.d. finitely supported random variables in a torsion-free abelian group, and write $S_k=X_1+\cdots+X_k$, and $H(S_k)$ is the Shannon entropy $S_k$, for all $k \ge 1$. We prove that, for every fixed $n\geq1$, \[ H(S_{n+1})-H(S_n) \geq \frac12\log\frac{n+1}{n} -o_{H(X_1)\to\infty}(1), \] uniformly over the ambient group and the input law. This proves a conjecture of Tao [29] in 2010.

math.PR

Geometric integrators for adiabatically closed simple thermodynamic systems

A variational formulation for non-equilibrium thermodynamics was developed by Gay-Balmaz and Yoshimura. In a recent article, the first two authors of the present paper introduced partially cosymplectic structures as a geometric framework for thermodynamic systems, recovering the evolution equations obtained variationally. In this paper, we develop a discrete variational principle for adiabatically closed simple thermodynamic systems, which can be utilised to construct numerical integrators for the dynamics of such systems. The effectiveness of our method is illustrated with several examples.

math-ph

A Temperature-Coupled Cahn-Hilliard-Stokes-Heat Model for Thermally Driven Phase Separation

We study a diffuse-interface model for thermally driven phase separation in viscous incompressible mixtures. The system couples a convective Cahn-Hilliard equation for the order parameter with a Stokes subsystem for the velocity-pressure field and a heat equation for the temperature. Temperature enters the bulk free energy through a Landau-type coefficient, while the phase field affects the flow through concentration-dependent density and viscosity. The model serves as a proxy for temperature-triggered condensation-like phase separation; humidity, latent heat, vapor pressure, and capillary forcing are absorbed into the choice of the threshold temperature $Θ_S$. We motivate the chemical potential through a temperature-dependent Landau free energy and use a regularized auxiliary formulation to prove local-in-time existence of weak solutions. For the numerical analysis, we employ a first-order sequential finite-element discretization of a simplified quasi-static formulation. The heat equation is advanced by implicit diffusion, the variable-coefficient Stokes problem is treated by a Taylor-Hood discretization, and the Cahn-Hilliard bulk derivative is evaluated at the previous time level, so each algebraic subproblem is linear. An isothermal diffusive test confirms mass conservation to roundoff and exhibits monotone discrete-energy decay for the tested parameters. Time-step and mesh-refinement studies show first-order temporal and approximately second-order spatial behavior. The remaining computations provide qualitative, parameter-specific illustrations; no global discrete energy law is claimed for the non-isothermal sequential scheme.

math.AP

Energy-Consistent Splitting and Decomposition Approaches for Coupled port-Hamiltonian ODEs

Operator splitting provides an attractive approach for the numerical integration of (coupled) port-Hamiltonian systems, as it allows the underlying system structure to be exploited at the level of the individual subproblems. However, the choice of the decomposition is not unique and may strongly affect both the computational efficiency and the preservation of the energy behavior of the original system. In this work, we investigate this interplay systematically and introduce energy consistency as a criterion for assessing splitting methods for port-Hamiltonian ordinary differential equations. We derive sufficient conditions under which a splitting based on a given decomposition inherits the energy behavior of the continuous system and use these conditions to analyze several decomposition strategies for coupled port-Hamiltonian systems. In particular, we compare decompositions that preserve the structure with approaches that exploit lower-dimensional subsystem dynamics or separated time scales. The analysis is complemented by numerical experiments using Strang splitting and its multiple-time-stepping extension. A scalable electro-thermal benchmark with fast electrical and slow thermal dynamics is employed to assess accuracy, energy behavior, and computational efficiency. The results demonstrate that preserving the port-Hamiltonian structure of the subflows is essential for energy-consistent splitting, whereas decompositions that exploit subsystem structure or time-scale separation can provide substantial computational advantages. In particular, the time-scale decomposition yields significant efficiency gains for systems with pronounced multirate characteristics, while structure-destroying decompositions may lead to undesirable energy behavior.

math.NA

Countable Graphs with Finite Path-width: Characterisation and Universality

We study path-width and the closely related parameter line-width in countably infinite graphs. Our first result characterises the graphs of finite path-width: they are the graphs that do not have infinitely many vertices of infinite degree, do not have infinitely many pairwise disjoint infinite paths, and contain no subdivision of some finite tree of maximum degree 3. We then investigate universality under the subgraph relation for graphs of bounded path-width or line-width. In particular, we prove that there exists a universal graph with line-width $\mathcal{O}(k^2)$ for the class of graphs with line-width at most $k$. In contrast, we show that no graph of finite path-width is universal for the class of locally finite graphs with path-width $1$. Finally, we show that for each $k\geq 2$, every universal graph for the class of graphs with path-width at most $k$ has line-width at least $k + 1$.

math.CO

Iterative Semantic Decoding for Short Block Codes

This paper proposes an iteratively enhanced semantic receiver for natural-language text transmission over noisy wireless channels using multiple short block codes. At the transmitter, each sentence is permuted by a character-level interleaver, partitioned into segments, and independently encoded by short block codes. At the receiver, we develop an iterative decoder consisting of a channel decoder and a language model, where a de-interleaver between them disperses the burst decoding errors within each segment across the sentence. In each iteration, the language model denoises the channel decoding output, and the denoised characters verified to be consistent with the channel observations are fed back to the channel decoder as semantic information for the next iteration. Simulation results on the Stanford Natural Language Inference (SNLI) corpus over the additive white Gaussian noise (AWGN) channel show that the proposed receiver achieves approximately 1.5 dB block error rate (BLER) gain over conventional short-block coding, while maintaining BLEU and ROUGE scores above 99% at SNRs beyond 1.0 dB.

cs.IT

Spectra of Non-Self-Adjoint Almost Mathieu Matrices and the Scottish Flag Operator

For $N\geq 3$ and a potential phase $\vartheta\in\mathbb{R}$, we study the non-self-adjoint almost Mathieu matrix obtained by multiplying the discrete Laplacian by a complex phase with angle $φ\in\mathbb{R}$, $A_N(φ,\vartheta)=e^{iφ}(S+S^{-1})/2+\operatorname{diag}(\cos(2πj/N+\vartheta))_{j\in\mathbb{Z}/N\mathbb{Z}}$, where $S e_j=e_{j+1}$ is the periodic shift on $\mathbb{C}^N$. We derive a Chambers formula and isolate the part $Q_{N,φ}$ of the characteristic polynomial that depends only on $N$ and $φ$, but not on $\vartheta$ or on a change of boundary conditions for the shift operator. We then show, for every $N$, that the zeros of $Q_{N,φ}$ lie on the two perpendicular lines $e^{iφ/2}\mathbb{R}\cup e^{i(φ/2+π/2)}\mathbb{R}$. For even $N$, the same property holds for the matrices $A_N(φ,\vartheta)$ with $\vartheta\in 2π\mathbb{Z}/N$, and we compute their limiting eigenvalue measure explicitly. For $φ\in[-π,π]$, the eigenvalue distribution approximates elliptic-integral densities with masses $1-|φ|/π$ and $|φ|/π$, and maximal radii $2|\cos(φ/2)|$ and $2|\sin(φ/2)|$, respectively. At $φ=π/2$, the central polynomial $Q_{N,φ}$ factors into positive quartic factors. This proves that the Scottish flag matrix, after Trefethen and Chapman, has its spectrum on the two diagonal lines of the saltire.

math.SP

Fast Gauss Sums via Flash Attention

Gaussian kernel sums are the computational core of maximum mean discrepancies (MMDs), kernel gradient flows, Stein variational gradient descent (SVGD), and many other kernel methods. At the same time, softmax attention has received an extraordinary amount of hardware-aware code engineering, culminating in flash attention. We show that Gauss kernel sums with arbitrary, signed weights can be evaluated via flash attention: two small input augmentations turn the normalized softmax reduction into the unnormalized Gauss sum, without writing a single line of custom GPU code. For feature dimension D>8 in fp16, this approach beats compiled PyTorch code as well as PyKeOps kernels (often significantly) in speed, memory-overhead and accuracy. Indeed, its memory scaling remains linear.

cs.LG