arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.

cs.CV↗

Wave-Growth-Limited Particle Confinement in a Near-Parallel Interplanetary Shock Observed by Parker Solar Probe

Collisionless shocks confine the particles they accelerate with the upstream wave field they themselves drive. Streaming particles resonantly amplify Alfvén waves, and the waves then confine the same particles by pitch-angle scattering. The rigidity a shock precursor can hold is therefore capped by how fast those waves grow, and that coupling has not previously been constrained in-situ at any interplanetary shock. Here we report Parker Solar Probe measurements of a fast, near-parallel coronal mass ejection-driven shock at 0.24~AU on 2023 March 13 that constrain this wave-particle coupling through the energetic-particle population it accelerates. We measure the spectral folding energy, the boundary between the confined and the upstream-escaping protons, and find that it falls from $\sim$6 to $\sim$1.5~MeV as $E_\mathrm{fold}\proptoΔt^{\,α}$ with $α\simeq0.34$, where $Δt$ is the time remaining before shock arrival. One precursor convective time before arrival this boundary is at $E_\mathrm{c}\simeq 2.46$~MeV, a characteristic confinement energy of a precursor of scale $L_\mathrm{EP}$. Read through the quasi-linear marginal-stability condition, the same evolution implies a resonant wave-growth timescale $τ_\mathrm{growth}(E)\propto E^{1/α}\simeq E^{2.92}$ that exceeds the precursor convective time above $E_\mathrm{c}$. The post-shock pitch-angle anisotropy reverses sign across the same confinement--escape transition at 2--5~MeV. The rigidity to which a shock precursor can confine particles is therefore set, in-situ, by how quickly the streaming population builds its own confining waves.

astro-ph.SR↗

Efficient Provably Private Classification with a Tabular Foundation Model

Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specific optimisation, suffers substantial utility loss under strong privacy, and is often difficult to apply correctly. Tabular foundation models adapt rapidly to new datasets, but existing models lack formal privacy guarantees, and are highly vulnerable to membership-inference attacks, limiting their use on sensitive data. Here we introduce PrivTab, an easy to use tabular foundation model for differentially private classification that embeds a privacy mechanism within its architecture. Pretrained on simulated datasets, PrivTab uses in-context learning to transform sensitive rows into compact, provably private summaries---effectively learning how to learn under privacy. PrivTab outperforms private linear and neural-network baselines under moderate-to-strong privacy, shows negligible membership leakage, maintains well-calibrated predictions under strong privacy, and reduces dataset fitting time by 10,000 times, requiring only a single forward pass. By combining formal privacy, speed, and easy of use, PrivTab brings recent advances in AI to applications where sensitive individual-level data have limited their adoption.

cs.LG↗

Demand Models for Market-Level Data with Closed-Form Inverses

We introduce a class of demand models for market-level data. The models can be estimated by linear instrumental variables regression while accommodating substitution patterns far richer than the logit and nested logit models they embed. They are built from closed-form inverse market share functions through a generator analogous to McFadden's generalized extreme value generating function, but acting on market shares. Constructive results allow arbitrary nesting structures, including overlapping nests and partial membership, yielding inverse-share analogs of generalized extreme value models. The class is consistent with utility maximization and strictly larger than the class of regular additive random utility models.

econ.GN↗

Exotic structures on hyperbolic manifolds via the EO-theory of Projective Spaces

Farrell and Jones showed that negatively curved manifolds in dimension $\geq 5$ are topologically rigid in the sense that homotopy equivalence implies homeomorphism. In the smooth category, the result does not hold even for hyperbolic manifolds, and there are negatively curved manifolds $N$ homotopy equivalent to a given hyperbolic manifold $M$ but not diffeomorphic. For complex hyperbolic manifolds, the analogous construction is demonstrated only in dimensions of the form $8n+2$. In this paper, we prove this result in many other dimensions. The main approach involves the use of $EO$-theory of complex projective spaces to construct suitable examples of exotic spheres. These theories may be viewed as a version of real $K$-theory, and are defined at each prime $p$. The computation also yields positive results about the existence of free smooth $S^1$ and $S^3$ actions on exotic spheres which do not bound parallelizable manifolds.

math.AT↗

HGP:An on-device personalized agent memory via hybrid graph storage

LLM-based agents face challenges in personalized interactive tasks due to heterogeneous, multi-typed, and implicitly constrained long-term traces. Existing memory mechanisms struggle with accurate routing and retrieval, especially on-device where personalization is critical. Most methods use single-vector representations, blurring type distinctions and relational structure. We propose HGP, a hybrid graph memory framework. HGP employs a lightweight self-enhancement classifier for personalized memory routing and constructs episodic, semantic, and procedural memories as graphs. It also extracts working memory as a state trajectory to capture current state and implicit constraints, ensuring reliable decision-making. The classifier reduces large-model calls, enabling on-device deployment, while graph storage enables accurate retrieval and incremental user profile refinement. Experiments on two benchmarks show that on PAL-Set solution selection, HGP achieves an S-score of 35.58, nearly 7 points above the strongest baseline. Code and data are at https://github.com/Ouan6/HGP-.git.

cs.AI↗

Invariant Primes in Lubin-Tate Space and Hovey-Strickland at Every Height

Let $H_n$ be the one-dimensional Honda formal group of height $n$ over $\mathbf F_{p^n}$ and let \[ R_n=W(\mathbf F_{p^n})[[u_1,\ldots,u_{n-1}]],\qquad A_n=R_n/(p). \] We prove, for every height $n$ and every prime $p$, that the prime ideals of $A_n$ stable under an open subgroup of the Morava stabilizer group are exactly the height ideals $(u_1,\ldots,u_j)$. We also prove that a stable prime of $R_n$ avoiding $p$ is zero. The resulting radical-ideal classification implies the Hovey-Strickland classification of thick tensor ideals in the category of dualizable $K(n)$-local spectra via the forward implication of Barthel-Heard-Naumann. The special-fiber argument is local and geometric. After cutting an invariant prime by a one-parameter curve, we construct from the Cartier structure equation a smooth formal quotient $\mathcal{Q}$ and a distinguished subgroup $\mathcal{H}$. The generic fiber of $\mathcal{H}$ is identified with the deformation space of connected-étale extensions. A Cartier obstruction map from the full Honda endomorphism order is compared with evaluation on Tate vectors through completed universal covers. Fargues-Fontaine vector bundles give a period-detection statement. Chai's rigidity theorem then promotes detection by homomorphisms to formal Zariski density after a renormalization of the valuation. A uniform fixed-jet argument transfers this density to Morava-stabilizer orbits. The generic-fiber assertion is proved separately from the Gross-Hopkins period map.

math.NT↗

Tradeoff between Wigner negativity and decoherence time for cubic Gaussian states

Wigner negativity has been shown to be a resource in quantum computing, raising the question of how to efficiently produce states with a large Wigner negativity. Cubic Gaussian states, obtained when a cubic gate is applied to a Gaussian state, have been proposed as candidates for this purpose. We show that cubic Gaussian states with a large Wigner negativity $\mathcal N$ necessarily have a short decoherence time $τ$ that is bounded above by $\mathcal N^{-2}$. In other words, a large Wigner negativity comes at a cost: the decoherence time $τ$ of cubic Gaussian states decreases at least quadratically in the negativity $\mathcal N$. Maximizing $τ$ at fixed Wigner negativity $\mathcal N$, we show that this upper bound is reached by optimal cubic Gaussian states. They have the property that for large negativity, the optimal squeezing scales as $\ln\mathcal N$ and the optimal cubicity as $\mathcal N^{-1}$. Optimal cubic Gaussian states with large negativity can therefore be constructed with small cubicity, provided sufficient squeezing is applied. However, their decoherence time decreases with growing negativity. In addition, we show that their preparation is extremely sensitive to thermal noise. Finally, we compare, for each integer $n$, the optimal cubic Gaussian state with the $n$th Fock state that has the same negativity: despite their pronounced differences in photon number distribution and Wigner function, we show that they have similar, be it slightly larger, decoherence times and entanglement generating potential through a beam splitter.

quant-ph↗

Finite Chains of Two-Complexes and Acyclic Covers

A negative answer to Whitehead's asphericity problem would follow from an infinite chain of two-complexes, starting with a non-aspherical one, in which every inclusion induces the zero map on second homotopy groups. We prove that a connected two-complex $K$ is the first term of such chains of every \emph{finite} length if and only if $K$ has a connected acyclic regular cover. In particular, there is a finite non-aspherical two-complex $K$, a presentation complex of $\mathrm{SL}(2,5)$, such that for every $n$ there is a strictly increasing chain of finite two-complexes $K=K_0\subset K_1\subset\dots\subset K_n$ in which all inclusions are zero on $π_2$.

math.GT↗

BetweenCut: Private Heavy-Node Classification with Doubly Logarithmic Error in Tree Height

Finding heavy nodes in a tree---those whose counts exceed a given threshold---is a building block for analysis and learning over structured data. Achieving record-level differential privacy (DP) without sacrificing accuracy is challenging because each record contributes to counts along an entire root-to-leaf path, allowing privacy costs to accumulate across levels. Existing methods account for the multiple threshold comparisons for each record incur additive error margins of $Ω_{\varepsilon,δ}(\log h)$ or $Ω_{\varepsilon,δ}(\sqrt{\log h})$ for tree height $h$. We introduce \textsc{BetweenCut}, an $(\varepsilon,δ)$-DP algorithm with an additive error margin of $O_{\varepsilon,δ}(\log\log h)$, improving the existing bounds for deep trees. This error holds simultaneously for all nodes and is independent of the input database size.

cs.CR↗

Towards Calibrated Probabilistic Forecasts for Events of Interest via Outcome-Conditional Recalibration

Calibration is an essential requirement for probabilistic predictions to be useful for decision making. While state-of-the-art prediction methods often yield miscalibrated predictive distributions, several post-hoc recalibration schemes have been proposed to generate calibrated predictions. However, popular recalibration schemes can conceal miscalibration in specific regions of the outcome space. Since particular outcomes, such as extreme events, often matter most for decision making, probabilistic predictions should be calibrated when evaluation is restricted to these outcomes. Hence, in this paper, we introduce outcome-conditional recalibration, a post-hoc method to recalibrate probabilistic predictions on user-defined regions of the outcome space. The method is simple, easy to implement, and can be applied to arbitrary predictive distributions. It works by applying the quantile recalibration approach of Kuleshov et al. (2018) to forecast conditional distributions, before rescaling these conditional distributions so that forecast event probabilities match empirical occurrence frequencies. This produces valid and continuous predictive distributions that are calibrated within each region of interest. Across regression benchmarks, we demonstrate that existing recalibration schemes do not necessarily yield calibrated predictions when interest is on particular outcomes, and that our approach improves outcome-conditional calibration relative to existing conditional and unconditional recalibration methods, while retaining competitive calibration overall. In an application to day-ahead electricity price forecasting, the approach substantially improves calibration when predicting negative prices, at negligible cost to forecast accuracy.

stat.ML↗

Nonexpansive Bijections in $T_0$-Quasi-Metric Spaces and Their $q$-Hyperconvex Hulls

We study bijective nonexpansive maps in $T_0$-quasi-metric spaces, distinguishing preservation of directed distances from pointwise rigidity. Our focus is their behaviour under passage to the $q$-hyperconvex hull. The upper quasi-metric integers are $q$-plastic and bicomplete, but their hull is the non-$q$-plastic upper real line. An explicit hull calculation and an established metric dense-rigidity theorem give the same failure of inheritance for a $q$-rigid original space. Tightness propagates distance preservation from the canonical copy throughout the hull. Under compactness and a bounded lattice structure for the zero-distance order, identifying that copy with the complemented elements gives a sufficient condition for inheritance of $q$-rigidity. A four-vertex asymmetric rectangle verifies this criterion. We also record metric transfer mechanisms and show that a directed interval condition gives an endpoint-coordinate representation, automatic $q$-plasticity, and $q$-rigidity under uniqueness of the ordered diametral pair.

math.GN↗

Exact fillings using universal covers in Floer theory

We study exact fillings of Legendrians and contact manifolds using Floer theory of universal covers. We prove vanishing results on wrapped Floer homology of the universal cover and symplectic (co)homology with $k[π_1]$ coefficients by using lifted versions of Ritter's TQFT operations. We use these results to show that an exact Maslov zero Lagrangian filling of the standard Legendrian $Λ_{std}$ in any Liouville filling of $(S^{2n-1},ξ_{std})$ is diffeomorphic to a ball in dimension at least 5. We show that the existence of a topological simple exact filling $W$ with vanishing symplectic homology for a dynamically convex contact manifold $M^{2n-1}$ reveals information about homotopy groups of $W$ and $M.$ We also strengthen a theorem of Zhou for ADC-contact manifolds as an application of universal covers.

math.SG↗

RealtimeWAM: How Fast Can I Run My World Action Model?

World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural design. More general caching strategies exploit feature redundancy, but redundancy alone does not capture the changing computational demands of closed-loop control. To address these challenges, we present RealtimeWAM, a general, training-free framework that coordinates parallel execution with adaptive computation for low-latency inference across diverse WAM architectures. We exploit layerwise dependencies to overlap observation processing with prediction. However, concurrent branches still compete for GPU resources, limiting the benefit of parallel execution. We therefore adapt computation throughout the pipeline through selective reuse, caching observation features in visually stable regions and reusing Transformer residuals while reserving additional refinement for small predicted adjustments. We evaluate RealtimeWAM on FastWAM and OpenWAM across RoboTwin, LIBERO, and LIBERO-Plus. On an RTX 4090, measured mean inference latencies are 24.09 and 63.09 ms, corresponding to average speedups of 8.90$\times$ and 10.67$\times$. Average success rates are 82.75% and 87.41%, respectively, within 0.02 and 0.53 percentage points of native inference. Across five real-world tasks, RealtimeWAM improves average success rates over native inference by 17.2 and 37.2 percentage points on FastWAM and OpenWAM, respectively.

cs.RO↗

Sharp Asymptotic Theory of Maximum Likelihood Estimation for Gaussian Processes with an RBF Kernel

Gaussian processes (GPs) are widely used across machine learning, spatial statistics, time-series analysis, optimization, Bayesian statistics, and scientific applications. A central component of a GP model is its kernel, which is typically specified through a parametric family. Among the most widely used choices is the radial basis function (RBF), also known as the squared exponential or Gaussian kernel, owing to its simple form, smoothness, and flexibility. In practice, the kernel parameters are routinely estimated by the maximum likelihood estimators (MLEs), as implemented by standard GP software. Despite this widespread use, the asymptotic behavior of the MLEs remains poorly understood under fixed-domain asymptotics, even for the RBF kernel. The main difficulty arises from the increasingly strong dependence among densely sampled observations and the nonlinear dependence of the covariance matrix on the kernel parameters. In this paper, we address this gap by providing, to the best of our knowledge, the first complete asymptotic characterization of the joint MLE of the spatial variance, lengthscale, and nugget variance under fixed-domain asymptotics. We establish consistency, derive convergence rates for all three parameters, prove joint asymptotic normality, and show that these rates are minimax optimal.

math.ST↗

Quantification of plastic deformation in anorthite using micropillar compression

Plagioclase is the most abundant mineral group of the crust and will therefore either control or strongly influence crustal rheology. Yet, we still know little about crystal plasticity in plagioclase as brittle deformation dominates under most laboratory conditions. To overcome past challenges in quantifying crystal plastic deformation in plagioclase, we used micropillar compression to quantify the strength of specific slip systems as well as the stresses necessary for mechanical twinning in anorthite. Pillars with diameters of around 1 micrometre were milled from three differently oriented cuts of the same anorthite single crystals using focused-ion beam. Uniaxial deformation of 36 pillars in total was then conducted inside a scanning-electron microscope under either 25, 300, or 800 degrees Celsius. Most pillars were deformed with a constant displacement rate that corresponds to a strain rate of around 0.001 1/s on the pillars. We find that the critical resolved shear stress (CRSS) for mechanical (pericline) twinning lies at around 300 MPa and is, as expected from previous studies on other materials, insensitive to changes in temperature or strain rate. Pillar deformation in crystals oriented to activate slip exhibit a strong decrease in yield stress with increasing temperature. Among the deformation mechanisms active during pillar deformation, twinning is the easiest under the applied conditions, followed by slip on (011)[100] regardless of temperature. Slip on (011)[0-11] or (0-11)[011] is consistently the hardest mechanism to activate in our anorthite pillars.

physics.geo-ph↗

A scanning cavity for large area coherent coupling in atom interferometry

Optical cavities can enhance atom-light coupling in atom interferometry, but free-fall geometries require large transverse modes. Marginally stable resonators provide such modes, yet aberrations and alignment imperfections turn their near-degenerate response into detuning-dependent transverse ring modes, strongly limiting the usable interaction volume under static excitation. Here, we turn this limitation into a resource. By chirping the interrogation frequency across the cavity resonance while dynamically shaping the input amplitude, we map laser detuning onto the transverse position and scan the cavity-enhanced coupling across the atomic cloud. In a cavity-enhanced Bragg diffraction experiment with $^{87}$Rb atoms, this produces a nearly uniform effective interaction region of about 15 mm and increases the transfer efficiency from about 5% to 28%. We also demonstrate phase-coherent operation of a three-pulse atom interferometer using chirped cavity-enhanced Bragg pulses. These results establish scanning-cavity interrogation as a route to large-area, cavity-enhanced atom interferometry.

physics.atom-ph↗

Two dimensional Villain model in the BKT phase

We establish sharp two-point function asymptotics for the two dimensional Villain model at sufficiently low temperatures, confirming a central prediction in the Berezinskii-Kosterlitz-Thouless theory. We characterise the power law exponent through the renormalisation group flow of a marginal line defect observable in the dual high-temperature Discrete Gaussian model.

math.PR↗