arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,549 records · Page 86Linked to original sources

SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs

Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatially-oriented GQA questions with scene-graph-grounded reasoning chains, retained only when the generated answer matches the symbolic ground truth, and a two-axis evaluation combining objective chain-overlap metrics with a scene-graph-aware LLM judge that scores faithfulness and completeness independently of the final answer. Applied to nine thinking-enabled VLMs, the protocol surfaces three findings invisible to standard accuracy: (i) four of nine models achieve $\geq$79% VQA accuracy while exhibiting shortcut rates above 39%, i.e., correct answers whose reasoning the judge marks as unfaithful; (ii) chain quality significantly predicts answer correctness for seven of nine models, but the two exceptions (Claude Sonnet 4.6, InternVL3.5-8B) reveal qualitatively distinct failure modes, terse output vs. verbose-decorative reasoning, that benchmark accuracy alone conflates; (iii) SFT on SpatialChain improves Qwen3-VL-8B by +6.2 pp in-domain and reduces its shortcut rate to 22%, while a stylistic specialization effect on external benchmarks motivates replay-augmented training as mitigation. The faithfulness judge is validated against 198 human-annotated items, where judge-human agreement matches human-human agreement, and against a second judge from a different provider, which preserves the model ranking ($ρ$ = 0.88). Data, generation scripts, and evaluation code are released at https://github.com/spatialchain/SpatialChainBenchmark.

cs.CV↗

A Mixed-Precision Model for Matrix-Splitting Methods

The performance of stationary iterative methods is memory-bound, due to the low arithmetic intensity. So, we propose a mixed-precision technique to reduce the amount of data movement in matrix-splitting iterative methods. We demonstrate both theoretically and experimentally that this technique will generally converge at a similar rate as an exact-precision iteration, depending on the precision used and the conditioning of the $M$ matrix. Furthermore, we demonstrate that iterative refinement can be used to refine the solution to full double-precision accuracy. Using this approach, we mixed half and double precision to achieve an average speedup over uniform double precision of $1.60\times$ for scalar and block Jacobi and $1.25\times$ for scalar and block Gauss-Seidel. Furthermore, we reduced the precision in the SSOR iteration used by OVERFLOW, a NASA code for compressible fluid dynamics simulations, and achieved a $1.14\times$ speedup to the overall application.

math.NA↗

Training-Free Transformer Merging via Sequential Local Operator Alignment

Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different activations than they do in the original model, thus, the merged and original execution paths no longer match. In this paper, we introduce Sequential Local Operator Alignment, a training-free method that merges transformers along the execution path of the partially merged model. Our method uses calibration data to estimate the local behavior of each functional component, aligns operators sequentially under the intermediate activation of the partially merged model, and subsequently factorizes the merged operators back into valid transformer parameters. We empirically show that this sequential step reduces error accumulation across layers. Furthermore, the proposed operator factorization step enables rank expansion, providing a principled mechanism for increasing multi-task capacity. We demonstrate that our approach generalizes across modalities, model scales, and varying numbers of tasks, from CLIP and RoBERTa to billion-parameter LLMs, and further extends naturally to the merging of LoRA-fine-tuned models. The results indicate improvements over strong merging baselines without requiring rank expansion, while optional expansion provides a further accuracy-inference-cost trade-off. Project link: https://akansh12.github.io/SLOA-Merge/

cs.LG↗

Dynamics around non-spherical symmetric bodies - III. The case of a spherical body with a crater

Motivated by the peculiar features of the trans-Neptunian object (TNO) Máni and the discovery of ring systems around Centaurs and trans-Neptunian bodies, we investigate the stability of particles around a small irregular body hosting a deep equatorial crater. This study is the third contribution in a sequence devoted to the dynamics around non-spherically symmetric bodies. Using Máni as a reference model, whose crater depth exceeds 10% of its radius, we map the system stability through complementary methods: Poincaré surfaces of section (PSS), survival maps, and the finite-time Lyapunov exponent (FTLE). In the nominal case, the $1\!:\!1$, $2\!:\!1$, $3\!:\!1$, and $4\!:\!1$ spin-orbit resonances (SORs) are consistently identified by all techniques. The prominent $3\!:\!1$ SOR exhibits a bifurcated structure at high eccentricities, matching structural transitions observed at lower Jacobi constant values in the PSSs. Increasing the rotation rate ($λ$) and crater mass ratio ($μ$) enlarges resonance widths and promotes overlap, making the $2\!:\!1$ and $4\!:\!1$ SORs progressively dominant. Lower rotation rates allow stable trajectories to persist at higher eccentricities, whereas the $1\!:\!1$ SOR is highly sensitive to stronger perturbations and disappears in the most extreme cases. In contrast to mass-anomaly models, which can clear the region interior to the $2\!:\!1$ SOR, the crater configuration considered here preserves stable regions in this vicinity, suggesting alternative scenarios for the formation and maintenance of debris structures around irregular minor bodies.

astro-ph.EP↗

Lindblad equation for the quantum Rabi model in the deep strong coupling regime

We derive a quantum master equation in Lindblad form for the open quantum Rabi model in the perturbative deep strong coupling regime, where the spectrum is close to degenerate and harmonic. The Lindblad form is obtained by combining secular and singular-coupling approximations, working in a Schrieffer-Wolff transformed frame. We benchmark our equation against a Bloch-Redfield one, analyzing the time-evolution and steady state properties of several observables and find excellent agreement. Analytical results are obtained in limiting cases.

quant-ph↗

Unbounded orbits in $C^1$-generic centrally symmetric strictly convex outer billiards

Let $\mathcal H$ be the space of support functions of centrally symmetric, strictly convex, compact planar bodies whose boundary is a $C^1$ curve, endowed with the $C^1$ topology. We prove that for a residual subset of $\mathcal H$ the outer billiard maps about the associated bodies admit unbounded orbits escaping to infinity. In particular, this residual set includes support functions arbitrarily $C^1$-close to those of disks or ellipses. This provides an affirmative answer, in the strictly convex $C^1$ setting, to a famous question of Moser and Neumann concerning the existence of unbounded outer-billiard orbits. The key ingredient is a rigidity theorem: if $h\in\mathcal H$ has a purely singular radius-of-curvature measure, then every continuous invariant tangent graph consists entirely of periodic points. We then show that for a generic table in $\mathcal H$ these periodic invariant graphs are absent, allowing us to construct escaping orbits via Mather's diffusing mechanism in a Birkhoff region of instability.

math.DS↗

Magic positivity for polar duals of pseudo-symmetric smooth Fano polytopes

Motivated by Gal's conjecture, Ferroni and Higashitani conjectured that the $h^*$-polynomial of any Gorenstein lattice polytope admitting a quadratic triangulation is $γ$-positive. Together with conjectures predicting quadratic properties of toric ideals of smooth lattice polytopes, this suggests that smooth Gorenstein lattice polytopes should have $γ$-positive $h^*$-polynomials. In this paper, we prove that the Ehrhart polynomial of the polar dual of every pseudo-symmetric simplicial reflexive polytope is magic positive. For pseudo-symmetric smooth Fano polytopes, we prove the stronger statement that all magic coefficients are strictly positive. Consequently, for every pseudo-symmetric simplicial reflexive polytope, its polar dual is Ehrhart positive and has a real-rooted and $γ$-positive $h^*$-polynomial. We also prove that the Ehrhart polynomial of the polar dual of the symmetric edge polytope of every cycle is magic positive. This gives an affirmative answer to a question of Konoike, who had previously proved partial positivity results for this family.

math.CO↗

Efficient Secure Federated Learning via Information-Theoretically Secure Key Distribution: A Medical Imaging Case Study

Federated Learning (FL) enables collaborative training of models across institutions without centralizing sensitive data, making it well-suited for privacy-concerned applications, such as medical imaging. To protect FL model updates during secure aggregation, additive masking is commonly employed. However, its underlying classical key establishment is only computationally secure. On the other hand, physics-based Information-Theoretically Secure (ITS) key exchange introduces practical constraints: finite key generation rates and time-limited storage severely limit throughput and sustained training of uncompressed models. In this work, we address this bottleneck by developing an FL framework that integrates frozen backbones, knowledge distillation, and quantization. These techniques reduce communication payload and, consequently, key material consumption. Moving beyond simulation, we benchmark this framework on a real physics-based key distribution testbed involving a chest X-ray classification application. Our results show that key usage can be reduced by $\sim$35$\times$ while maintaining predictive accuracy. This prevents buffer depletion and key expiration, enabling sustainable FL training under physical key generation constraints.

cs.LG↗

On the Comparison of Optimizers for Imbalanced Learning

Data imbalance is pervasive in machine learning, from rare words and anomalies to underrepresented patterns in heterogeneous or cross-tabulated data. We study idealized optimizers geometries in continuous time to model small-step training in deep learning. We assume that the source of imbalance is unobserved: the optimizer has only access to the aggregate training loss ignoring the exact contributions of the majority and minority groups. In this setting, we characterize a region where majority losses are optimized regardless of the admissible minority structure. We derive explicit equations of this zone and bounds on the time needed to leave it. These bounds exhibit a milder dependence on minority amplitude for sign, spectral, and Newton descent than for Euclidean gradient descent. Experiments with AdamW and Muon suggest similar advantages over SGD across language, tabular, and image tasks.

stat.ML↗

How close can rational points get to a manifold?

We prove the heuristically predicted lower bound for the number of rational points of height at most $Q$ lying within $ε/Q$ of a fixed analytic nondegenerate manifold in $\mathbb{R}^n$, provided that $ε\asymp Q^{-τ}$ for some $τ\leq 3/(2m+1)$, where $m$ is the codimension of the manifold. Our result establishes the lower bound well beyond the previously conjectured range $τ\leq 1/m$, and improves upon a recent result of Schindler, Srivastava, and Technau, who established the same lower bound for $τ\leq 3/(2n-1)$.

math.NT↗

Toward Reliable Infant Pose Estimation: A Training-Dynamics Approach to Noisy Annotation Detection

Spontaneous movement analysis in preterm infants relies increasingly on markerless pose estimation (PE) to derive clinically relevant motion biomarkers directly from video recordings. Training accurate infant PE models requires large sets of manually annotated keypoints, and human annotation is inherently prone to error. Noisy keypoints (i.e., keypoints mislocalized with respect to their true anatomical position) are especially problematic in this clinical setting, since they can propagate as artificial artifacts into the reconstructed joint trajectories. Building on the small-loss hypothesis and training-dynamics-based sample selection established in the noisy-label learning literature, we propose a novel framework for detecting noisy keypoint annotations. A hybrid convolutional-attention model is trained to predict the anatomical category of each keypoint from its spatial coordinates and local visual features; the resulting cross-entropy training dynamics are then used to derive per-keypoint descriptors, which are partitioned into clean and noisy subsets via unsupervised clustering. We validate the approach on NeoPose, a newly collected dataset of 65 hospitalized preterm infants, under two realistic noise scenarios (random positional perturbation and left-right swapping) across multiple noise levels. Results show that the proposed approach achieves an F1-score of up to 91.9% in noisy-keypoint detection. The framework further generalizes to the heterogeneous COCO benchmark, where filtering CE-detected noisy keypoints from the training set also yields measurable improvements (up to 7.4 AP points) in downstream pose estimation accuracy at moderate-to-high noise levels.

cs.CV↗

SoK: Semantic Decision Engines in Network Control Loops

A semantic decision engine such as Jev can return a valid answer and still miss a network deadline, select an infeasible action or leave the service unverified. We systematize 139 paper families by decision interface, execution path and check ownership. Fifty families claim that their engine fits a control loop or time budget, but only four support the claim with matched measurement. Across all 139, four report deadline attainment. The gap concentrates where the decision has no deterministic computation step. Those 72 families make 22 of the claims, none supported, and name a coverage owner in only two. Bounded tests under one event model show that each gap can reverse an admission verdict. A decision that meets a 10 s budget for every isolated request meets it for none once decisions queue ahead of replayed execution times. The same engine passes one coverage check and fails another. We derive a minimum reporting record, design rules and a research agenda for admitting decision engines to control loops.

cs.NI↗

Reflexive ideals of a numerical semigroup

We study the reflexivity property of ideals in numerical semigroups. In particular, we characterize irreducible reflexive ideals and determine the Apéry set of a reflexive ideal. Moreover, we study how reflexivity allows us to characterize some classes of numerical semigroups and the interaction of reflexivity with the property of being a trace ideal.

math.AC↗

A Taxonomy on Collective Awareness

As robotic teams tackle increasingly complex tasks in dynamic and unstructured environments, effective coordination requires agents to maintain accurate, aligned representations of their environment, teammates, and mission state. We argue that Mutual Awareness, Shared Situational Awareness, and Team Situational Awareness --- three concepts widely invoked in the multi-robot systems literature --- are not interchangeable: they form a containment hierarchy in which each type subsumes the previous in scope, and span a distribution spectrum from fully individualized understanding (MA) to fully uniform understanding (SSA), with TSA occupying a mixed position. We establish this through a two-dimensional taxonomy organized along the object of awareness and awareness distribution axes. Treating these terms as synonyms obscures the precise coordination requirements each imposes on a robotic system. A Search and Rescue case study with heterogeneous aerial and ground robots grounds each taxonomy position in concrete coordination requirements. These findings provide a conceptual foundation for principled specification, design, and comparison of collective awareness in multi-robot systems.

cs.RO↗

Choice via Ownership Rules

We consider a model of collective choice where for each menu there is an agent who is accountable for the choice made. Specifically, we assume that two agents (individuals or selves) allocate choice tasks over different menus using an ownership rule: the owner of the decision over a menu picks what maximizes his own preference and is accountable for it. We characterize three distinct ownership rules under different (nested) choice environments.

econ.TH↗

Do Speech Representations Preserve Regional Accent Across Read and Spontaneous Speech?

Regional accent cues can be captured under matched conditions, but it remains unclear whether they persist between read and spontaneous speech. We study RVG1, with 500 German speakers from nine regions, comparing ten speech representations on regional classification and continuous geolocation under matched conditions and speaker-independent read--spontaneous transfer. Whisper performs best under matched conditions, reaching 0.489 nine-way UAR and 148 km median geolocation error, but drops to 0.11/0.18 UAR across transfer directions and 363 km geolocation error. Self-supervised models show a similar degradation, whereas speaker embeddings are less discriminative in-domain but more robust under transfer. This contrast is consistent across classification and geolocation. Across representations, robustness is associated with how little a representation shifts between styles (style-invariance), for which crossstyle speaker retrieval is an interpretable proxy. Age, sex, sentence-overlap, and duration controls do not account for the gap, although channel characteristics contribute. These results show that strong matched-condition performance does not indicate robust regional information.

cs.CL↗

Direct Vertex-Derivative Coefficient Propagation forHomogeneous Numerical Integration (HNI)

Homogeneous numerical integration (HNI) combines Euler's identity and Stokes' theorem to integrate homogeneous polynomials over a polytope $P$. Existing algorithms return scalar integrals or reusable moment families. Any resulting moment table can be recoded afterward: once all degree-$q$ moments are known, integration on $\mathcal H\_q$ (degree-$q$ forms) has a trivial representation by a pure order-$q$ distribution supported at any single arbitrary point. Our contribution is instead to propagate functionals through the face complex, once for $(P,q)$ and without first forming the moment vector, to obtain a signed vertex-derivative coefficient functional on $\mathcal H\_q$, specific to $P$ and determined directly by its supplied face geometry and chosen anchors. A finite-dimensional Riesz framework describes its nonuniqueness, gives an intrinsic $L^2(P)$ error for restricted formulas, and supports weighted post-processing. The independently generated degree-$q$ and degree-$2q$ coefficient vectors satisfy an explicit compatibility relation. Vertex anchors can prune the propagated representation as shown in some polygon examples and an offline benchmark quantify this effect without claiming global cardinality optimality or improved conditioning. The construction covers convex polytopes and oriented non-convex polyhedral domains in arbitrary dimension.

math.OC↗