arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,333 records · Page 74Linked to original sources

Spin Birefringence and Spin Polarized Transmission across p-wave Altermagnetic Heterostructure

This paper investigates electron transport properties across N/AM interface or N/AM/N heterojunction. Based on Hamiltonians in the three regions, general wave functions are proposed and combined linear equations are constructed. Reflection and transmission coefficients as well as spin polarizatioin of the altermagnet heterostructure are obtained. When electron incident from normal metal to altermagnet, it is scattered and split into + branch and _ branch transmitting into altermagnet zone with separate refractive angles. Total internal reflection might happen to + branch. Critical angle of total reflection is determined by incident conditions and properties of AM material. In the heterostructure case, spin-up and spin-down electrons transmitted with same propagation directions but different transmission probabilities. Transmission and polarization periodically change with length of the junction. Exchange field and spin-splitting strength have pronounced effect on spin-polarized transmittance while the effect of spin-orbit coupling is quite small. Optimized system can be found by tuning parameters of the AM material. This study will lay the theoretical groundwork for experimental research on spin-selective outputs and spintronic devices.

cond-mat.mes-hall↗

Tunability of the structural and magnetic transition in kagome material: PrIr$_3$B$_2$

We report the temperature and pressure tunability of an unusual structural transformation associated with a two-step metal-insulator-metal (MIM) transition in the kagome lattice compound PrIr$_3$B$_2$ using synchrotron X-ray powder diffraction. At ambient conditions of temperature and pressure, the monoclinic ($C2/m$) and the hexagonal ($P6/mmm$) phases coexist as twinned structure in the crystal. As the temperature (pressure) is decreased (increased), PrIr$_3$B$_2$ converts fully to monoclinic structure at $T =$ 280 K (at ambient pressure) and $P =$ 1.2 GPa (at room temperature). Temperature dependence of the monoclinic structure at ambient pressure presents complex evolution of lattice parameters with weak but clearly discernable anomalies at \SI{\sim 250} {\K} and \SI{\sim 110} {\K}, which are correlated with the second MIM transition and the linear to nonlinear temperature-dependent resistivity crossover, respectively. These anomalies are likely due to some charge order state causing a partially gapped Fermi surface. The magnetic phase diagram of PrIr$_3$B$_2$ is also investigated from anisotropic measurements. At 10 K, a superzone gap opens near the antiferromagnetic transition, which does not close even in the polarized state. From the tunability of the crystal structure and magnetic and electronic ground state, promising electronic orders are indicated in this kagome metallic magnet.

cond-mat.str-el↗

Computing electron overlap integrals for Gausslet orbitals on cubic lattices

Gausslet orbitals on a cubic lattice, introduced in [Steven R. White, J. Chem. Phys. 147, 244102 (2017)], represent a localized, smooth, and systematically refineable basis set featuring a "diagonal" approximability of the electron repulsion integral tensor. The present work develops efficient algorithms for evaluating overlap integrals required for electronic-structure simulations using these Gausslets, specifically the kinetic, nuclear, and electron repulsion integrals (both with and without the diagonal approximation). Computational efficiency improvements rest on the exploitation of translation, permutation, and octahedral symmetries, a reordering and precomputation of nested sums, as well as an early truncation of small coefficients. Our algorithms reduce the number of electron repulsion integrals on a $5 \times 5 \times 5$ grid from $125^4 = 244140625$ to $324275$ due to symmetries, and achieve a wall-clock runtime for evaluating the remaining integrals with a truncation tolerance of $10^{-5}$ in under 2 seconds on a laptop computer. We apply the developed methodology to compute the ground state of the hydrogen atom and molecule as a demonstration.

physics.chem-ph↗

Whitening Improves Robustness to Spurious Correlations in Linear Probes

Deep neural networks tend to rely on simple features that may be spurious and thus fail to generalize. We study this problem in the setting of linear probes, where a (generalized) linear model is fitted on the representations of a (pretrained) model. We use the connection of these models to the max-margin classifier, and show they favor directions associated with large eigenvalues of the covariance matrix. Whitening removes this preference by equalizing the eigenvalues of the covariance matrix. This observation motivates whitening as a preprocessing step that can reduce reliance on spurious correlations without requiring prior knowledge of their presence or labeled data. We examine the effect of whitening on a synthetic data-generating process and standard spurious correlation benchmarks, and find that it improves robustness. We also find that whitening can improve robustness when added to existing approaches.

cs.LG↗

Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics

Recently, Vision-Language-Action (VLA) models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understanding, and action generation in an end-to-end learning framework. However, since these models are designed to interact directly with the physical world and humans, their security is critical, and even small vulnerabilities can lead to catastrophic failures. In this work, we propose the Universal Adversarial Object, a sphere with optimized surface texture that significantly degrades task success rates when placed within the robot's field of view. Specifically, our approach introduces a multi-level attack framework that jointly disrupts trajectory planning, task execution, and action control. We validate our method in both simulated and real-world robotic settings. Experimental results demonstrate that the adversarial object reduces the average task success rates by 31.2%-39.9% for two representative VLA models (Pi0 and RDT), with success rates dropping to near zero in complex scenarios. Index Terms--Vision-Language-Action models, adversarial attack, robotic security, universal adversarial object

cs.RO↗

LocoWM: High-Precision Locomotion through World-Model-Guided Residual Adaptation

High-precision locomotion combines motion-command tracking with precise regulation of task-relevant physical states, enabling robots to interact reliably with their surroundings during motion. Joint end-to-end optimization can leave precision objectives insufficiently optimized, while reactive residual control adjusts actions only after deviations become observable. We present \textbf{LocoWM}, a world-model-guided preactive residual adaptation framework for high-precision locomotion. A base policy provides command-following locomotion, while an action-conditioned world model predicts a sequence of future physical states from proprioceptive history and the proposed base action. A residual adapter conditions on this predicted sequence to generate additive action corrections that compensate for anticipated deviations. Two-stage training first learns locomotion and action-conditioned dynamics, then freezes both modules while training the adapter, separating locomotion acquisition from precision adaptation. Experiments spanning terrain leveling, acceleration compensation, and push recovery demonstrate improved control precision and disturbance robustness over end-to-end and reactive residual baselines. Demos and code are available at: https://zhaozijie2022.github.io/LocoWM

cs.RO↗

Multi-feed Plane Wave Generator for Compact OTA Multi-Target Emulation With Joint Angle-Delay-Doppler Control

Over-the-air (OTA) radar target emulation is an important method for evaluating the sensing capabilities of integrated sensing and communication (ISAC) and radar devices. Conventional single-feed compact antenna test ranges (CATRs) are generally limited to a single target direction, while multi-feed angular-synthesis methods commonly used in radar testing still require far-field illumination over the device-under-test (DUT) aperture. To overcome these limitations, this paper proposes a radar target OTA emulation framework based on multi-feed plane wave generator (PWG). The key capability is the simultaneous emulation of spatially close targets from different directions with joint angle--range--Doppler control. Target range is emulated through the corresponding delay, while Doppler and the complex echo coefficient are controlled by a channel emulator. Meanwhile, a multi-input amplitude and phase matrix (APM) maps each independent target-signal path to a dedicated PWG excitation vector that synthesizes the corresponding incident plane wave over the same test zone. This enables multiple target directions to be reproduced as local plane waves at a test distance substantially shorter than the conventional far-field requirement. The framework is experimentally validated at 3~GHz using a real $1\times11$ phased-array DUT. The PWG--DUT separation is 1.6~m, substantially shorter than the 7.2~m Fraunhofer distance of the DUT. At this separation, the synthesized $0^{\circ}$ and $10^{\circ}$ incident waves each exhibit an amplitude peak-to-peak error of 1.0~dB. Across two four-target angle--delay cases and one two-target angle--velocity case, the estimated target parameters show angular deviations within $1^{\circ}$, interpolated delay errors within 2.4~ns, and velocity errors within 0.15~m/s.

eess.SP↗

The Minkowski $?(x)$ function and Salem's problem. II

In 1943, R. Salem asked whether the Fourier-Stieltjes transform of the Minkowski question-mark function vanishes at infinity. This problem was answered affirmatively by Jordan and Sahlsten in 2016 as a consequence of their general results on Gibbs measures for the Gauss map. In this note we give a self-contained proof for ?(x). Several aspects of our argument are substantially different, making a very short proof possible in the ?(x) case. Moreover, we obtain an explicit polynomial decay estimate, with exponent at least 0.0472. We then prove the existence of the correlation dimension of the Minkowski measure, thereby improving the previous upper bound for the optimal Fourier decay exponent from 0.4373 to 0.4223.

math.CA↗

MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models

World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.

cs.CV↗

Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift

This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.

cs.CV↗

Fiber-Resolved Microstructure Quantification from Multi-Shell Diffusion MRI using Detection Transformers

Fiber orientation and compartmental microstructure are central to the characterization of white matter tissue in diffusion MRI, yet existing methods either resolve fiber orientations without quantifying microstructure, or quantify microstructure while assuming a fixed number of compartments and a single fiber direction. Nonparametric approaches that recover both require tensor-valued diffusion encoding and computationally expensive Monte-Carlo inversion of an ill-posed inverse Laplace transform. We propose to reframe this problem as an object detection-like task, adopting the Detection Transformer (DETR) architecture to jointly predict mean diffusivity (MD), fractional anisotropy (FA), main fiber direction, and signal fraction for a variable number of compartments per voxel from standard multi-shell diffusion MRI with linear encoding. Hungarian matching during training resolves permutation invariance across compartments. We introduce mean Average Precision as a reproducible benchmark metric. Evaluated on synthetic test data with up to five compartments per voxel, our model achieves $R^2=0.95$ for MD, $R^2=0.88$ for FA, and a median angular error of 4.2°, with performance scaling naturally with compartmental signal fraction.

cs.CV↗

Low-Discrepancy Dither for Quantized Recurrent State Caches

Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems round the state stochastically; we ask which rounding rule such caches should use. We find that a deterministic golden-ratio Weyl dither, which needs no random numbers, consistently brings the quantized model closer to the full-precision one than stochastic rounding, across pure and hybrid models, storage formats, and long decoding horizons, at no extra cost. Round-to-nearest behaves differently: because it discards small updates, its error keeps growing, so it can look best in short evaluations yet falls far behind over long generations. A discrepancy analysis explains this ordering, and we document implementation pitfalls that silently remove the benefit.

cs.LG↗

Black Holes With Complete Asymptotic Throats

We investigate whether a spherically symmetric black hole can have a complete inner end at finite, nonzero areal radius while satisfying the standard energy conditions sufficiently far outside its horizon. Retaining two independent metric functions, we treat both null-type and spacelike-type asymptotic throats and derive a necessary and sufficient integral criterion for causal geodesic completeness. Under boundedness and finite-limit assumptions on the independent curvature components, null-type ends have limiting curvature fixed by the throat radius, whereas spacelike-type ends admit an additional nonnegative curvature parameter. We obtain criteria for bounded curvature in timelike and null parallel-propagated frames. Complete spacelike-type examples can have bounded scalar invariants and bounded timelike-frame curvature but unbounded null-frame curvature at infinite affine parameter. We also prove that the radial null energy condition must fail arbitrarily close to every complete inner end in the stated class. Nevertheless, explicit real-analytic constructions with either type of throat satisfy the null, weak, dominant, and strong energy conditions throughout a sufficiently distant exterior region. Smooth constructions can become exactly Schwarzschild beyond a finite radius. These results establish the compatibility of complete asymptotic throats with exterior energy conditions while distinguishing completeness from different levels of curvature regularity.

gr-qc↗

A limit theorem linking continuous plate models to the topology of the plate boundary network

Plate tectonics is described in two registers: a discrete one, in which rigid plates tile the sphere and meet at trivalent junctions, and a continuous one, in which Bercovici and Wessel (1994) replace hard plate outlines by smooth shape functions of finite boundary half-width $δ^*$. We prove that the two registers are connected by a limit theorem. Generalizing the shape functions to a signed-geodesic-distance construction, we show that they converge almost everywhere, exponentially in $1/δ^*$, to the plate indicator functions; that the sharp plate mosaic is a finite regular trivalent CW decomposition of the sphere under an explicit structural hypothesis; and that a component-counted weighted Euler characteristic $χ_w^{δ^*}(τ)$, built from the super-level sets of the shape functions, equals $V-E+F=2$ for all $δ^*$ below an explicit threshold $δ_0(τ)$, with its face, edge and vertex terms converging separately to the numbers of plates, boundary arcs and triple junctions. The theorem is proved unconditionally on the signed-distance offset cover under positive-reach, separation and bounded-sector hypotheses, and transferred to the normalized partition-of-unity field for thresholds $τ<1/3$ by a comparison lemma whose junction-ball step, a planar single-crossing property, is certified numerically. The residual $|χ_w-2|$ defines a topological diffuseness index for diffuse plate boundaries. For the present-day PB2002 network the hypotheses are measured to be non-vacuous, with $δ_0(0.15)\approx2.8$ km. The theorem is the mathematical foundation of the topological audit of plate reconstructions presented in a companion paper (Kim, 2026, submitted to Geoscience Frontiers).

physics.geo-ph↗

Attention Function as an Intrinsic Inductive Bias: How Models' Behavior Diverges in Novel Contexts

Developmental psychology holds that certain priors are given to infants prior to experience rather than induced from data, and that the influence of such priors is suppressed under strong, well-constrained conditions but reasserts itself under weak ones. We ask whether an analogous principle holds for the Transformer: can the activation function given to attention heads serve as an intrinsic inductive bias? We propose Mixture of Function Attention (MoFA), a parameter-free modification to multi-head attention that fixes a ratio of softmax and sigmoid heads before training. Across five ratios, a 124M-parameter GPT-2 model, and five seeds, we find that this given ratio has little effect in-distribution -- differences between ratios are statistically negligible for moderate mixtures and remain small even at the extremes -- but its influence re-emerges sharply under zero-shot distribution shift across 15 out-of-distribution domains. Perplexity gaps between ratios widen by more than an order of magnitude on several domains, and the best-performing ratio tracks a single axis of domain structure, separating short, informal text (softmax-favoring) from technical, long-form text (sigmoid-favoring), that explains 78.3% of the variance in domain response. This reorganization is visible at the head level: sigmoid heads show an accelerating drop in attention entropy as their ratio increases, while softmax heads respond more modestly, yielding a consistent division of labor between the two head types. Our results suggest that activation choice functions as a given prior whose influence is masked in-distribution and re-emerges out-of-distribution.

cs.LG↗

ViLegalExpert: A Large-Scale Benchmark for Vietnamese Legal Retrieval and Question Answering from Real-World Consultations

Trustworthy Legal AI requires systems that can answer legal questions while grounding their responses in authoritative sources. However, existing Vietnamese legal benchmarks provide limited coverage of real-world legal consultations. We introduce \textbf{ViLegalExpert}, a large-scale benchmark constructed from authentic citizen--lawyer consultations, containing over \textbf{172K} questions across \textbf{34 legal domains}, together with professional answers and expert-verified legal evidence. ViLegalExpert supports legal information retrieval, extractive QA, and abstractive QA. Experiments with representative retrieval methods and language models reveal substantial challenges in evidence retrieval and grounded answer generation. While pretrained models perform strongly on QA, hybrid retrieval achieves the best retrieval performance. These results demonstrate the difficulty of mapping naturally expressed legal questions to authoritative provisions and establish ViLegalExpert as a challenging benchmark for reliable Vietnamese Legal AI.

cs.CL↗

Dynamics to decision: A mathematical theory of Lyapunov spectra and decision boundaries in deep classifiers

A deep classifier is defined not only by the decision it produces, but also by the sequence of transformations through which that decision is formed. Treating this evolution as a dynamical system across layers provides a natural framework for asking how decision geometry emerges through depth and how far back we can trace a boundary's dynamical signature. We model a feed-forward classifier as a finite, nonautonomous discrete dynamical system, with layers playing the role of discrete time steps. We study the Finite-Time Maximum Lyapunov Exponent (FTMLE) of the data samples' dynamical trajectory through depths of the classifier. The FTMLE measures the rate of convergence/divergence of nearby trajectories. We move the observation endpoint backward from probabilities to logits and then to hidden representations. For Gaussian classes, we prove that probability-level FTMLE carries a clear geometric signature of the decision boundary, with its dominant direction aligned with the boundary normal. Moving one step backward to the logits, we prove this relationship is no longer universal but depends critically on how the classifier is trained, particularly on the choice of loss function. Moving further backward to the hidden representation, the connection becomes more conditional: boundary-related FTMLE can persist, but only under identifiable structural conditions. We propose geometry-aware fine-tuning for restructuring the classifier's hidden FTMLE, and propose conditions for guaranteed concentration of high hidden FTMLE near the decision boundary. Through our numerical results, we show the generality and validity of our theoretical results. Understanding the evolution of data samples as traveling through the layers of classifier provides a principled foundation for identifying where boundary-relevant sensitivity emerges and for developing layer-aware regularization strategies.

cs.LG↗

Combining standard chemoradiotherapy with CAR-T cell therapy in malignant gliomas: Insights from an impulsive mathematical framework and virtual trials

The extreme therapeutic resistance of malignant gliomas motivates the development of multimodal treatment strategies. We present a mechanistic mathematical framework to investigate the coupled dynamics of radiotherapy, temozolomide, and chimeric antigen receptor (CAR) T-cell therapy. The model is formulated as a nine-dimensional impulsive dynamical system that captures proliferative-quiescent tumor states, treatment-specific resistance, tissue damage, TMZ pharmacokinetics, and the lymphotoxic interferences exerted by conventional regimens on engineered T cells. Mathematical analysis establishes the existence of biologically meaningful solutions, maps invariant structures, and identifies threshold conditions under which sustained therapy locally stabilizes the tumor-free equilibrium. After study-matched benchmarking against clinical trials, including the standard Stupp protocol, we deploy heterogeneous virtual patient cohorts to evaluate five combined CAR-T--Stupp sequences. At the population level, the immunotherapeutic adjunct yields a consistent but modest survival extension, increasing median overall survival by $0.8$--$0.9$ months with minimal sensitivity to temporal ordering or scheduling perturbations. Crucially, these population-level medians mask profound individual heterogeneity: under combined protocols, $19\%$--$22\%$ of virtual patients extend survival by more than $60$ days, and $\sim 3\%$ gain over a year. A low intrinsic tumor proliferation rate $r_1$ emerges as the primary determinant of response, bolstered by robust T-cell expansion and attenuated microenvironmental suppression. These findings suggest that the clinical promise of multimodal integration depends on identifying responsive tumor phenotypes for personalized treatments rather than searching for a single universally optimal timeline.

math.DS↗