arXiv ScienceSearch

SEARCH · arXiv Science

Results for “math.AT”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

5,006 records · Page 6Linked to original sources

The Non-Orientable Topology of Condorcet's Paradox

Preference cycles are prevalent in problems of decision-making, and are contradictory when preferences are assumed to be transitive. This contradiction underlies Condorcet's Paradox, a pioneering result of social choice theory, wherein intuitive and seemingly desirable constraints on decision-making necessarily lead to contradictory preference cycles. Topological methods have since broadened social choice theory and elucidated existing results. However, characterisations of preference cycles in topological social choice theory are lacking. In this paper, we address this gap by introducing a framework for topologically modelling preference cycles that generalises Baryshnikov's existing topological model of strict, ordinal preferences on 3 alternatives. In our framework, the contradiction underlying Condorcet's Paradox topologically corresponds to the non-orientability of a surface homeomorphic to either the Klein bottle or real projective plane, depending on how preference cycles are represented. These findings allow us to reformulate Arrow's Impossibility Theorem in terms of the orientability of a surface as well.

math.AT

GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts

Mathematical reasoning is a key capability for vision-language models (VLMs), yet current benchmarks mainly evaluate text-based or explicitly symbolic visual inputs. It remains unclear whether VLMs can reason mathematically when information must be perceived and inferred from images rather than read from explicit symbols. We introduce GSM8K-V, a benchmark transforming GSM8K into multi-image sequences with semantic equivalence preserved. By mapping text-based problems into visual form via an automated pipeline and human verification, we curate 1,319 high-quality samples. In GSM8K-V, quantities must be extracted through visual perception, and reasoning chains must be reconstructed by integrating implicit cues across scenes. Evaluation of 34 VLMs reveals a striking modality gap: while most models exceed 90\% on text, the best model achieves only 59\% on GSM8K-V, far below the 91\% human accuracy. Notably, models enhanced for visual math reasoning show no improvement on GSM8K-V despite large gains on existing benchmarks, confirming that it evaluates a distinct capability. Error analysis shows that the primary bottleneck lies in Implicit Visual Inference Error (IVIE), where models fail to recover visual semantics that are implied rather than explicitly stated. Our code and data are released at https://github.com/ZJU-REAL/GSM8K-V.

cs.CV

Isogeometric $C^1$ mortar method

We present an isogeometric mortar method for the discretization of the biharmonic equation posed on multi-patch domains. We assume only $C^0$-conformity at interfaces and employs a mortar approach to weakly enforce $C^1$-continuity across patch interfaces. Discrete inf-sup stability is ensured by selecting a Lagrange multiplier space consisting of splines of degree reduced by two compared to the primal space, with increased smoothness or merged elements near vertices. We prove optimal a priori error estimates and confirm the theoretical findings with a series of numerical experiments.

math.NA

Distributed Linear Programming on GPU Clusters at Extreme Scale

Large linear programs can exceed the memory of a single compute node. Although first-order methods replace sparse factorizations with GPU-suited matrix-vector products, other solver phases can reintroduce a single-node memory limit. We present SHARDLP, a distributed GPU LP solver that keeps the matrix and primal-dual state partitioned from sharded input through solution output. On the Google PDLP benchmark, SHARDLP reaches the published criterion on nine of eleven instances, compared with eight in the published CPU PDLP study. On the largest benchmark, eight H200 GPUs solve a 1.185-billion-variable, 6.338-billion-nonzero LP in 9.9 minutes; the published CPU experiment reports 21.06 hours on different hardware. Beyond this benchmark, separately checked multi-node solves reach up to 13.604 billion variables and 40.807 billion nonzeros, while validated executions span up to 76 GPUs across 29 compute nodes. For column-partitioned solves, support-aware communication skips GPUs that store no coefficients for a row; on an LP with 2.76 billion nonzeros, it cuts modelled communication by 92.97% and improves solver time by 1.27x-1.52x

math.OC

A Borel Concept Class of VC Dimension One with a Non-PAC Consistent Learner in ZFC

The fundamental theorem of statistical learning states that, under suitable measurability assumptions, finite Vapnik--Chervonenkis (VC) dimension guarantees that every proper consistent learning rule is probably approximately correct (PAC). Blumer, Ehrenfeucht, Haussler, and Warmuth showed, assuming the Continuum Hypothesis, that the "well-behavedness" condition of the concept class cannot be omitted: they constructed a concept class of Borel sets of VC dimension one admitting a consistent learning rule that is not PAC. We show that the Continuum Hypothesis is unnecessary. Working in Zermelo--Fraenkel set theory with the Axiom of Choice (ZFC) alone, we construct a concept class of Borel sets on $[0,1]$ of VC dimension one and a proper consistent learning rule that is not PAC. More precisely, for a suitable Borel probability measure and target concept, the rule has true risk one at every sample size on a set of samples of outer probability one. Consequently, finite VC dimension and Borel measurability of the individual concepts do not suffice to guarantee that every proper consistent learning rule is PAC. The result shows, with no need of extra set-theoretical assumptions, that the additional regularity assumption in the fundamental theorem cannot in general be omitted.

math.LO

Rényi Entanglement of Purification Is Non-additive

Entanglement of purification is a fundamental measure of total correlations whose additivity remains unresolved. We study its additivity for classical states on two qubits at different Rényi orders. For every $α\in[0,1)$, we prove nonadditivity within this family, witnessed by two copies of a single state. We first solve the one-copy optimization exactly for the entire family at every Rényi order. We then restrict the two-copy optimization to a natural finite set of purifications and exhibit one whose entropy is strictly below the product value. In contrast, for $α\in[2,\infty]$ we prove additivity under tensor products within this family. The interval $α\in[1,2)$, including the von Neumann case $α=1$, remains open, and we conjecture additivity there throughout the same family.

quant-ph

A Note on Binary Quadratic Systems and their relation to complexity theory

Deciding whether a system of multivariate quadratic equations over $\mathbb F_2$ has a solution is a classical NP-complete problem, and remains so for square systems, with as many equations as variables. The hardness of this problem is one of the cornerstones of nowadays post-quantum cryptography. Let $\MQ_0(n)$ and $\MQ_1(n)$ denote the sets of square quadratic systems in $n$ variables having respectively no solutions and exactly one solution. $\cup_{n\geq 2} \MQ_0(n)$ is a coNP-complete language, while $\cup_{n\geq 2} \MQ_1(n)$ lies in DP. It is known that $\lim_{n\to \infty} |\MQ_1(n)|/|\MQ_0(n)|=1$. Here we prove the explicit finite-$n$ bounds \[ |\MQ_0(n)|<|\MQ_1(n)| \le \left(1+\frac{1}{2^n-1}\right)|\MQ_0(n)|, \] More generally, let $Q_d$ be the space of polynomial functions $(\FF_2)^n\to\mathbb F_2$ of degree at most $d$, and let $α_k$ count square systems in $(Q_d)^n$ having exactly $k$ solutions. Then \[ α_0<α_1 \le \left(1+\frac{1}{2^n-1}\right)α_0\,, \qquad 2\le d\le n \,. \] The proof combines matroid and coding-theoretic methods. We interpret $(\FF_2)^n$ as the ground set of the evaluation matroid of $Q_d$, express $α_0$ and $α_1$ through characteristic polynomials, and use a Whitney-type sign-reversing involution to show that the only terms that can push $α_1-α_0$ below $α_1/2^n$ come from the elements of a matroid port. These are identified with minimal-support words of the Reed--Muller code $\RM(n-d-1,n)=\RM(d,n)^\perp$; the required estimate then follows from the MacWilliams identity, the minimum-distance bound $2^{d+1}$, and the even-weight structure of the code.

cs.IT

Eigenvalues and eigenfunctions of the fractional Laplacian on the interval

We prove a three-term asymptotic formula for the eigenvalues of the fractional Laplacian on the bounded interval $(-1,1)$. This improves the eigenvalue asymptotics of Kulczycki--Kwaśnicki--Małecki--Stós and Kwaśnicki, and confirms the conjectural $O_α(n^{-2})$ remainder suggested by the numerical simulations of Kaleta--Kwaśnicki--Małecki. Moreover, we prove that the normalized eigenfunctions are bounded uniformly in the eigenvalue index $n$ and the fractional order $α$. This settles the conjecture proposed by Kwaśnicki through numerical experiments. Furthermore, we prove that the $n$-th eigenfunction has exactly $n-1$ zeros in the interval $(-1,1)$ and every zero is simple, and hence there are exactly $n$ nodal domains. A key ingredient in the proof is an explicit representation of the eigenfunction.

math.CA

A note on the $Σ_2^P$-completeness of the Frobenius number

Given a finite set $A$ of natural numbers whose greatest common divisor is one, the Frobenius number $g(A)$ is the largest integer that is not a non-negative integer combination of the numbers in $A$. In a 2016 preprint, Matsubara states that given $A$ and $k$, deciding if $g(A) \geq k$ is $Σ_2^P$-complete. A decade has passed since without peer-reviewed publication of this result. At the same time, the community has found it difficult to verify this result. In this note, we give a write-up of the completeness proof based on Matsubara (2016).

cs.CC

Robust topology optimization with non-Gaussian material fields using polygonal finite elements

We present a computational framework for robust topology optimization that integrates polygonal finite-element discretizations, spatially correlated non-Gaussian material modeling, and non-intrusive polynomial-chaos surrogates. Spatial uncertainty in Young's modulus is represented as a homogeneous non-Gaussian random field obtained via a memoryless transformation of a truncated Karhunen-Loève expansion, ensuring physical admissibility through positivity of stiffness while preserving the prescribed autocovariance. Polygonal finite elements provide a stable discretization for density-based optimization on unstructured meshes and mitigate checkerboard artefacts and mesh bias, while the sparse polynomial-chaos expansion enables efficient estimation of low-order statistical moments required by the robust objective at a fraction of the cost of intrusive or Monte Carlo approaches. Numerical studies on a cantilever and a curved beam show that introducing non-Gaussian material variability leads to systematic load-path redistribution and a reallocation of 6-12% of the structural volume, together with a reduction in compliance scatter. The non-intrusive surrogate reproduces intrusive reference results within 3% using an order of magnitude fewer full finite-element analyses. These results demonstrate that the proposed framework offers a physically consistent and computationally efficient route to topology-optimized designs that remain reliable under realistic material uncertainty.

cs.CE

Logarithmic-Free Moment and Generalization Bounds for Uniformly Stable Algorithms

Uniform stability is a classical tool for controlling the generalization error of a learning algorithm. Bousquet, Klochkov, and Zhivotovskiy (2020) showed that the problem can be reduced to a moment inequality for a sum of weakly interacting functions of independent random variables. Their bound contains an additional factor $\log n$, and they asked whether this factor can be removed. We answer this upper-bound question affirmatively. More specifically, let $Z=(Z_1,\ldots,Z_n)$ have independent coordinates and let $g_i(Z)$ satisfy $\mathbb E[g_i(Z)\mid Z_{-i}]=0, \ \left| \mathbb E[g_i(Z)\mid Z_i]\right|\le M, \ \text{for every } i = 1, \dots, n, $ where $Z_{-i}$ denotes all coordinates except $Z_i$. Assume additionally that changing any coordinate $Z_j$, $j\neq i$, changes $g_i$ by at most $β$, we prove that, for every $p\ge2$, for every $p\ge2$, $$ \left\| \sum_{i=1}^n g_i(Z)\right\|_p \le 16pnβ+M\sqrt{2pn}. $$ This removes the $\log n$ factor from the previous bound and matches the lower bound of Bousquet, Klochkov, and Zhivotovskiy up to universal constants in the range covered by their construction. Our proof first establishes the required estimate on the Rademacher cube, then transfers it to arbitrary product distributions by a two-copy randomization argument.

stat.ML

Turing complete Navier-Stokes steady states via cosymplectic geometry

In this article, we construct stationary solutions to the Navier-Stokes equations on certain Riemannian $3$-manifolds that exhibit Turing completeness, in the sense that they are capable of performing universal computation. This universality arises on manifolds admitting nonvanishing harmonic 1-forms, thus showing that computational universality is not obstructed by viscosity, provided the underlying geometry satisfies a mild cohomological condition. The proof makes use of a correspondence between nonvanishing harmonic $1$-forms and cosymplectic geometry, which extends the classical correspondence between Beltrami fields and Reeb flows on contact manifolds.

math.DG

Local Identifiability of Networks with Nonlinear Node Dynamics

We study the identifiability of nonlinear network systems with partial excitation and partial measurement when the network dynamics is linear on the edges and nonlinear on the nodes. We assume that the graph topology and the nonlinear functions at the node level are known, and we aim to identify the weight matrix of the graph. Our main result is that, for almost all static analytic nonlinearities that cross the origin, directed graphs are generically locally identifiable if and only if at least one node is excited in every source component of the condensation graph and at least one node is measured in every sink component. This holds even when all other nodes remain unexcited and unmeasured and stands in sharp contrast to most findings on network identifiability requiring measurement and/or excitation of each node. The result applies to homogeneous feed-forward and recurrent artificial neural networks and generalizes previous literature by considering a broader class of activations and architectures.

math.OC

Geometric mean and Lebesgue-type decomposition of completely positive maps

We introduce the geometric mean and the parallel sum of completely positive (CP) maps between von Neumann algebras, based on the Pusz--Woronowicz theory of positive sesquilinear forms. We provide a concrete characterization via a block matrix positivity condition and establish their fundamental properties, including the AM--GM--HM inequality with respect to the CP order. In finite-dimensional settings, our construction is compatible with the Choi--Jamiolkowski correspondence, under which the geometric mean of CP maps corresponds to the Kubo--Ando geometric mean of their Choi matrices. This yields a natural operator-theoretic framework for interpolating quantum channels. As an application, we obtain index-type inequalities for conditional expectations in subfactor theory. Finally, we establish a Lebesgue-type decomposition of CP maps via a parallel sum construction, thereby providing a unified framework that simultaneously generalizes Ando's decomposition of bounded positive operators and Kosaki's decomposition of normal positive functionals on von Neumann algebras.

math.OA

Estimating quantum relative entropies on quantum computers

Quantum relative entropy, a quantum generalization of the renowned Kullback-Leibler divergence, serves as a fundamental measure of the distinguishability between quantum states and plays a pivotal role in quantum information science. Despite its importance, efficiently estimating quantum relative entropy between two quantum states on quantum computers remains a significant challenge. In this work, we propose the first quantum algorithm for directly estimating quantum relative entropy and Petz Renyi divergence from two unknown quantum states on quantum computers, addressing open problems highlighted in [Phys. Rev. A 109, 032431 (2024)] and [IEEE Trans. Inf. Theory 70, 5653-5680 (2024)]. Notably, the circuit size of our algorithm is at most $2n+1$ with $n$ being the number of qubits in the quantum states and it is directly applicable to distributed scenarios, where quantum states to be compared are hosted on cross-platform quantum computers. We prove that our loss function is operator-convex, ensuring that any local minimum is also a global minimum. We validate the effectiveness of our method through numerical experiments and observe the absence of the barren plateau phenomenon. As an application, we employ our algorithm to investigate the superadditivity of quantum channel capacity. Numerical simulations reveal new examples of qubit channels exhibiting strict superadditivity of coherent information, highlighting the potential of quantum machine learning to address quantum-native problems.

quant-ph

An Inverse Problem for Determining the Piston Speed from a Given Lipschitz Leading Shock

We analyze an inverse problem for determining the piston speed and the associated flow field from a prescribed leading shock and the initial data in a shock tube. The gas flow is described by the isentropic Euler equations (i.e., the $p$-system), while the trajectory of the leading shock is prescribed as a given Lipschitz curve. Under an Oleĭnik-type entropy condition on the leading shock, we develop a modified wavefront tracking scheme to construct the flow field behind the shock. This construction enables us to determine the corresponding piston speed and the associated flow field.

math.AP

Any-Dimensional Learning by Sampling

Many machine learning models are defined for inputs of different sizes, such as point clouds containing different numbers of points, sequences of tokens of different lengths, and graphs on different numbers of nodes. Such models are trained on finitely many examples of necessarily limited sizes. How well do these models generalize from inputs of small size to larger inputs of size not seen during training? Furthermore, evaluating such models on large inputs is often expensive. How can we sketch large inputs to obtain smaller ones on which the model takes similar values? At the heart of both questions is the need to compare inputs of different sizes and to approximate large inputs by small ones. We present a unified approach to address these questions by using random sampling maps to compare inputs of different sizes. The sampling maps we consider are generalizations of sampling with replacement, random binning, and species sampling. We characterize the application domains in which each type of sampling is appropriate in terms of the symmetries and relations between problem instances of different sizes in the domain. Our framework yields explicit generalization and sketching rates for function classes continuous with respect to a chosen notion of sampling, encompassing large families of functions defined on sequences, graphs, and tensors of different sizes. Specific examples include moment polynomials on measures, homomorphism densities and numbers of graphs, permutation-invariant transformers, and graph neural networks.

math.ST