arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,351 records · Page 75Linked to original sources

The twisted convolution identity and ghost $r$-SICs from finite quantum dilogarithms

Radchenko and Wheeler (RW) recently proved a finite pentagon relation for real quadratic special values of the modular quantum dilogarithm and used it to establish the rank-$1$ twisted convolution identity conjectured by the current authors. RW gave an explicit argument in the principal case and remarked that their proof holds for all rank-$1$ admissible tuples. We extend their proof to all rank-$r$ admissible tuples and provide an explicit dictionary between the modular quantum dilogarithm and the Shintani-Faddeev modular cocycle conventions in the respective papers. Thus, we establish that, if $d,r$ are positive integers such that $r<\frac{d-1}{2}$ and $\frac{d^2-1}{r(d-r)} \in \mathbb{Z}$, then there exist ghost $r$-SICs: i.e., configurations of $d^2$ rank-$r$ subspaces in $\mathbb{C}^d$ that satisfy a non-Hermitian equichordal condition. Under the Stark conjecture, these configurations are Galois conjugate to Hermitian equichordal configurations called $r$-SICs (or rank-$r$ SIC-POVMs).

math.NT↗

Heavy Dark Baryons as Self-Interacting Dark Matter: A GeV-Scale Coincidence

We study self-interacting dark matter composed of heavy baryons in a confining dark $SU(N)$ gauge theory with a single heavy vector-like quark, $m_Q \gg Λ_d$. In this regime, the dark baryons are compact Coulombic bound states whose long-distance interactions can be described by a multipole expansion. We derive the leading baryon-baryon interaction by adapting the quarkonium operator-product-expansion formalism to heavy baryons. The resulting force contains non-retarded London dispersion, retarded Casimir-Polder, and Yukawa contributions from dark glueball exchange. We find that the self-interaction phenomenology depends strongly on the number of colors. For $N=2,3$, the attractive interaction is too weak to overcome the short-distance Pauli repulsion, and no phenomenologically relevant velocity dependence is obtained within our minimal set-up. For sufficiently large $N$, the attractive interaction remains competitive with the repulsive core and can generate strong velocity dependent cross sections. Taking $SU(10)$ as a representative finite-$N$ realization, we find regions compatible with self-interaction requirements inferred from dwarf galaxies and clusters, while allowing much larger cross sections at low velocities relevant for gravothermal evolution. These regions correspond to GeV-scale dark baryons and a confinement scale close to that of QCD, which naturally motivate an asymmetric dark matter origin.

hep-ph↗

Importance-Aware Feature Sparsification for Wireless Split Learning

Wireless split learning (SL) reduces on-device computation by offloading upper layers to a server, yet transmitting high-dimensional intermediate features at each iteration remains a major communication bottleneck. Existing methods select features at the client side using task-agnostic criteria such as magnitude, statistics, or clustering, which increases client-side processing and often degrades accuracy under non-independent and identically distributed (non-i.i.d.) client data. We propose importance-aware class-balanced sparsification (ICS), a lightweight approach in which the server ranks feature channels using Grad-CAM-based scores obtained from the true-class logit during backpropagation. The per-class scores are aggregated into a class-balanced, label-agnostic importance vector that mitigates head-class bias under label skew, and each client reuses this vector in the next round to retain the top-$N$ feature channels, incurring no additional client-side forward or backward passes. We further derive a non-asymptotic convergence bound that isolates the sparsification-induced error and characterizes how the sparsification ratio and mini-batch size jointly affect convergence under a fixed communication budget, and we analyze the communication and computational overhead of ICS against representative baselines. Beyond sequential CNN-based SL, we extend ICS to parallel split learning and to transformer-based models. Experiments show that ICS consistently outperforms the baselines, with larger gains under severe non-i.i.d. partitions.

cs.LG↗

Uruqi: Learning Spatial Cognition from Visual Experience

Spatial intelligence requires maintaining a coherent understanding of the world as the embodied agent moves. Like humans, the agent must use its own motion to interpret changes across observations and update object locations and spatial relations accordingly. Despite spatial post-training having substantially broadened the spatial intelligence of vision-language models (VLMs), they still struggle with two atomic spatial capabilities: tracking self-motion and mapping the surrounding world during motion. To address this gap, we provide dense multi-turn supervision over interleaved atomic capabilities within each training episode, mimicking the visual experience of a continuously moving agent that reasons as it observes. To scale this up, we synthesize 11,738 motif-driven camera trajectories over a broad range of 3D scenes, supporting self-motion tracking, persistent object mapping, and rich spatial operations within each visual experience. By training models to reason over these atomic questions, our URUQI$_{\mathrm{Syn}}$-8B improves accuracy from 15.84% to 47.73% on our Uruqi benchmark comprising 52k questions across 2.7k episodes. URUQI-SI-Mix-8B further reaches 50.41%, comparable to the 50.08% achieved by GPT-6 Astra. Trained solely on our synthesized data, URUQI$_{\mathrm{Syn}}$-8B achieves an average relative accuracy improvement of 17.13% over its InternVL3-8B backbone across three external spatial benchmarks. These results highlight continuous visual experience as a scalable source of supervision for developing spatial cognition in VLMs.

cs.CV↗

Particle track reconstruction in high-density collider environments with an Evolving-Geometry Transformer

Charged particle track reconstruction at the High-Luminosity Large Hadron Collider is challenged by dense hit environments and combinatorial ambiguities. Existing graph-based and Transformer approaches commonly rely on static geometric neighbourhoods, which cannot adapt as track-level representations evolve. We propose the Evolving-Geometry Transformer (EGT), a graph Transformer that reconstructs the hit-connectivity graph at successive encoder layers and incorporates relative hit coordinates as learnable attention biases. Across five simulation benchmarks spanning linear, helical, and reduced TrackML events with up to 200--500 simultaneous trajectories, EGT achieves the highest FitAccuracy on three of five settings under the shared benchmark protocol, reaching 79.6\% on the most complex setting and exceeding the strongest baseline by 1.6 percentage points. Replacing the evolving topology with a static graph reduces the perfect-track rate from 81\% to 29\%, while removing the geometric bias reduces it to 73\%. Under synthetic background injection equal to the number of signal hits, FitAccuracy decreases by 3.4 percentage points. The neural-network forward pass requires 31~ms per event on an NVIDIA A100 GPU. These results indicate that iterative refinement of hit associations improves complete-track recovery when local geometry alone is insufficient to determine trajectory membership.

hep-ex↗

Improved Finite-Time Lyapunov Exponent framework for Comprehensive Synchronization Analysis in Multilayer Networks

We propose an Improved version of the Lagrangian Finite Time Lyapunov Exponents(IFTLE) framework as a new quantifier for Comprehensive Synchronization analysis of multilayer networks undergo adaptive coupling scheme in order to incorporate the effects of deformation in synchrony due to the dynamic variations in the coupling strengths of intra-layer and inter-layer links. Theoretical formulation of the IFTLE along with its numerical validation has been provided using a three layer network of Chaotic Rössler oscillators. The numerical illustration clearly shows the efficacy of the proposed method revealing the need for comprehensive synchrony deformation analysis in adaptive coupling scenarios. Also we have stated a few properties of IFTLE useful for the quantification purpose.

nlin.CD↗

DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction

While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack temporal motion cues; the \textbf{latency gap}, where inference delays render actions obsolete; and the \textbf{control gap}, caused by the open-loop action chunk execution without real-time adjustment. In this work, we propose \textbf{DSDyn-VLA}, a Slow-Fast \textbf{D}ual-\textbf{S}tream \textbf{Dyn}amic manipulation framework that integrates motion-aware foresighted planning with real-time residual correction. The slow \textbf{Flow-Planner} serves as a macro-planner. By enhancing the VLA with optical flow for temporal perception and a future state awareness mechanism to preemptively offset inference latency, it produces globally consistent, motion-aware action chunks. Complementing this, the fast \textbf{Res-Refiner} employs a lightweight RL policy to inject high-frequency, closed-loop corrections into the planned action chunks based on real-time observations. In addition, we introduce \textbf{DynBench}, a MuJoCo-based benchmark for dynamic object manipulation that comprises nine tasks. Extensive experiments demonstrate that DSDyn-VLA reduces the failure rate by over 76\% compared to current SOTA method in high-latency setting on the Kinetix dynamic benchmark, while achieving about 6$\times$ the success rate of PI0.5 in real-world dynamic settings and about 5$\times$ on DynBench. We will open-source all the code and weights.

cs.RO↗

UniAE-MoE: A Unified Audio Encoder via Mixture of Experts

Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific data scaling (TSDS) technique to enhance \tool's understanding capabilities. On the XARES-LLM benchmark, UniAE-MoE attains a score of 0.802, achieving state-of-the-art performance. It also delivers top-tier performance in the official Interspeech 2026 Audio Encoder Capability Challenge, further demonstrating robust generalization across diverse audio tasks. Together, these results validate the effectiveness of \tool for unified audio understanding across speech, music, and general audio domains.

cs.SD↗

Fast and Stable Harmonic Approximation from Ill-distributed Data using Moving Least Squares

We consider harmonic approximation on the torus $\mathbb{T}^d$, the sphere $\mathcal{S}^d$, and the rotation group $\mathrm{SO}(3)$. While several global approaches are available for computing a harmonic expansion from scattered data, their stability and, especially, their runtime depend strongly on the geometry of the nodes, in particular on a sufficiently small fill distance. Even a single large hole in the data may cause instability and long runtimes. We propose harmonic approximation via moving least squares (HAMLS), a global harmonic approximation scheme that avoids the ill-conditioned global solve and works well in both the overdetermined and the underdetermined settings. The main idea is an intermediate transition from the scattered data to a quadrature grid, which is realized via moving least squares (MLS). This replaces one large global problem by many small local ones that can be regularized individually. The global harmonic approximation is then obtained via quadrature using a fast Fourier transform. This makes the global stage fast, stable and non-iterative. For sufficiently smooth functions sampled at well-distributed nodes with small fill distance $h$, we bound the resulting $\mathrm{L}^2$-error by the error of the harmonic approximation obtained from exact values on the quadrature grid, plus a term of order $h^{K+1}$, where $K$ is the polynomial degree employed in the MLS step. Numerical experiments on $\mathcal{S}^2$ with ill-distributed nodes show that HAMLS achieves errors comparable to those of global least-squares approximation, while being significantly faster, especially for larger bandwidths.

math.NA↗

Exact analytic solutions of classical time crystals

We present examples of classical time crystals (in a parabolic potential for the position) that admit simple exact analytic solutions. The kinetic energy in these crystals is piecewise parabolic as a function of speed and is minimised at nonzero values of the speed. The solutions are periodic, while the speed acquires definite periodic discontinuities. The solutions are stable along the branches where the action is lower. We also examine the classical time crystal with quartic kinetic energy. We show how the quartic kinetic energy can be emulated by a piecewise parabolic one, leading to accurate and transparent solutions.

cond-mat.other↗

3D Reconstruction from Arthroscopic Images using NeRF: a preliminary in-silico study

In knee arthroscopy surgery, accurate registration between preoperative and intraoperative anatomy is a critical step for patient-specific navigation. Achieving an accurate registration requires a reliable 3D reconstruction of the joint during surgery. Preoperative 3D models can be obtained from patient imaging through segmentation and reconstruction, but generating an intraoperative 3D representation remains particularly challenging. Arthroscopic imaging suffers from a limited field of view, low surface texture, and strong specular reflections, which make conventional feature-based 3D reconstruction methods unreliable. In this work, we investigate the application of MIS-NeRF (Minimally-Invasive Surgery Neural Radiance Fields) for reconstructing intraoperative knee 3D models from monocular arthroscopic images. The approach is evaluated on six simulated arthroscopic acquisitions representing six patient-specific knee 3D models. Both qualitative and quantitative results are presented to assess the reconstruction quality. The reconstructed knee 3D models were evaluated through their rendered images, achieving PSNR (Peak Signal-to-Noise Ratio) of 31.88 $\pm$ 2.82, SSIM (Structural Similarity Index) of 0.98 $\pm$ 0.004 and LPIPS (Learned Perceptual Image Patch Similarity) of 0.017 $\pm$ 0.006. These preliminary results suggest the feasibility of NeRF-based reconstruction in the challenging context of arthroscopy and may represent a promising step toward accurate in-silico preoperative-to-intraoperative 3D registration for computer-assisted orthopedic surgery. Further, validation on real arthroscopic data will be necessary to assess clinical applicability.

eess.IV↗

Benchmarking EMlog Calibration for Autonomous Surface Vehicles

Accurate velocity measurement is a fundamental requirement for autonomous surface and underwater vehicles. Commonly, velocity is provided by a Doppler velocity log (DVL) sensor, yet it becomes unavailable due to operational altitude constraints. In such situations, electromagnetic logs (EMLogs) provide a critically robust alternative for continuous velocity estimation. However, raw EMLog measurements are inherently corrupted by systematic errors, which need to be calibrated prior mission begins. Currently, a benchmarking comparative evaluation of how different calibration models perform under rapidly changing dynamic sea conditions is missing in the literature. To bridge this gap, this paper presents a comparative model-based calibration methodology that evaluates four distinct calibration models using two different estimation pipelines. The proposed framework is rigorously validated on a unique 221 minutes of continuous real-world telemetry collected from the MARVEL surface vehicle during dynamic sea trials. The dataset contains two different EMLogs and DVL recordings. Experimental results demonstrate that the bias and scale error model implemented with the Kalman filter improves the speed estimation by 71%. We also demonstrate that dynamical manoeuvres further improve the accuracy compared to standard straight-line paths, ultimately delivering a validated, real-time online calibration EMLog approach for autonomous surface vehicles.

cs.RO↗

Threshold Geometry, Bifurcation, and Data-Driven Analysis of a Vaccination-Treatment Model for Hepatitis B

Despite the availability of excellent vaccines and antiviral medicines, hepatitis B virus (HBV) remains a major public health problem. In this work, we develop a mathematical vaccination-treatment model for HBV transmission to investigate threshold conditions for illness persistence as well as the level of intervention. The analysis gives the characterisation of both disease-free and endemic equilibria and determines the basic reproduction number, $\mathcal{R}_0$, by the next-generation matrix approach. The model indicates that $\mathcal{R}_0=1$ is a critical threshold dividing disease-free from endemic states and characterises a forward transcritical bifurcation regarding the transmission parameter. Numerical equilibrium continuation and stability assessments confirm the analytical results. A two-parameter study of transmission and vaccination rates identifies threshold geometries and assesses impacts of important parameters using sensitivity analysis. The model is calibrated with India-specific data, such as HBsAg prevalence, HBV incidence and mortality rates and fits epidemiological targets closely through limited nonlinear calibration. Additionally, multi-start optimisation and practical identifiability profiling are used to investigate parameter uncertainty. Robustness analyses reveal that the calibrated $\mathcal{R}_0^*=1.1744$ is associated with a super-threshold regime; however, parameter flexibility permits both sub-threshold and super-threshold regimes. This combined analytical and data-driven framework connects epidemic thresholds, bifurcation structures, intervention efficacy, and parameter uncertainty. It provides a quantitative framework for analysing HBV persistence and evaluating control options, especially in the absence of epidemiological evidence.

math.DS↗

Excursion-Resolved Thermodynamic Inference without Observing Reverse Transitions

Experiments often detect only selected transitions, leaving the underlying state dynamics and dissipation hidden. We show that waiting-time distributions between directed visible events can reconstruct the boundary dynamics without observing the reverse of every detected transition. Under a simple source-target closure condition, they determine visible rates and the hidden excursions connecting observed states, yielding a rigorous lower bound on total entropy production that improves the standard waiting-time estimate when reverse events are available. Remarkably, observing only a spanning directed cycle is sufficient to identify the full Markov generator, even when all reverse cycle edges and other microscopic transitions remain unseen.

cond-mat.stat-mech↗

Plug-in Optimization Method for Penalty Parameters in Nonparametric GMANOVA model

In order to estimate the longitudinal trend, we often use the generalized multivariate analysis of variance (GMANOVA) model (Potthoff \& Roy, 1964), when the longitudinal data is balanced data. Usually, we use this GMANOVA model with some polynomial at the time of measurement. However, when the longitudinal trend has flexible curve, we cannot derive good fitting estimated curve when we use some polynomial curves. Then, Nagai (2011) proposed the nonparametric GMANOVA model which uses on several known basis functions instead of using the polynomial curves. If we use several basis functions, then overfitting problem is occurred. Nagai (2011) also proposed the estimation method for avoiding overfitting and unstable problems, and reducing computational iterative algorithm by extending the generalized ridge regression model (Yanagihara, Nagai \& Satoh, 2009). In the present paper, we extend one of the optimization methods in Nagai, Yanagihara and Satoh (2012) into the estimation method in Nagai (2011). Through numerical studies, we show some properties of each optimization method.

stat.AP↗

ASENA: Self-evolving Agents for Embodied Navigation

We present ASENA, an embodied agent system that connects general-purpose coding agents to robot sensing, computation, supervised execution, and persistent experience. Agents can write and execute programs, inspect recorded outcomes, repair failures, and reuse notes and executable skills while keeping their model weights fixed. We further introduce ASENA-VLN, a 4B monocular navigation policy that serves as an optional tool within this programmable system. ASENA-VLN predicts body-frame trajectories for both extended routes and short-horizon behaviors using a shared vision-language decoder trained on route instructions, visual question answering, and a newly curated dataset of geometry-derived atomic navigation tasks. As a standalone policy, ASENA-VLN achieves state-of-the-art success rates of 68.7% on R2R and 70.2% on RxR. When integrated with a coding agent, learned navigation improves ASENA's success rate by 11 percentage points on both agentic benchmarks while reducing execution time. Through persistent workspace evolution and simulator feedback, ten passes over recurring 100-task subsets further improve success from 72% to 98% on R2R and from 65% to 89% on RxR. On embodied question answering, ASENA achieves state-of-the-art accuracy with fewer interaction steps. Finally, real-world demonstrations on a Unitree G1 combine search, visual inspection, spatial reasoning, and synthesized gestures without a pre-built map, illustrating how online programming extends robot behavior beyond route following and predefined skills.

cs.RO↗

Probing Helical Primordial Magnetic Fields via Chiral Gravitational Waves in the LISA-TAIJI Network

Primordial magnetic fields (PMFs) are anticipated to be a cosmological candidate for generating a stochastic gravitational wave background (SGWB) through their anisotropic stress. In particular, if parity violation is present in the production of PMFs, it can induce a helical component in the PMFs, leading to circular polarization in the isotropic SGWB. In this study, we phenomenologically consider parity-violating inflationary magnetogenesis and compute the intensity and circular polarization of the SGWB for several assumed power spectra of helical PMFs. Using the planned sensitivity of the LISA-TAIJI network, we calculate the signal-to-noise ratio (SNR) and perform a Fisher forecast to assess the feasibility of constraining the parameters of helical PMFs. We conclude that it is possible to estimate the helical-to-non-helical power ratio~$r_H$ with $\mathrm{SNR}^{V}>2$ if the PMF strength is $\mathcal{B} \gtrsim 10~\mathrm{nG}$ ($50~\mathrm{nG}$) in the case of the delta-function-type (scale-invariant-type) PMF spectrum. Our results offer a useful criterion that indicates the future observational limit on helical PMFs at the small scale $k_\mathrm{LISA} \approx 10^{12}~\mathrm{Mpc}^{-1}$.

astro-ph.CO↗

O-Funnel: Lossless Structural Capture and Requirement-Driven Extraction from Drifting, Heterogeneous Documents

Pulling a fixed set of fields out of documents that arrive in many formats and under drifting schemas is usually done with hand-written byte patterns, which break whenever a key is renamed, a value is reformatted, or a lookalike value appears first. We argue the cause is structural: one pattern must both describe the value and locate it among its surroundings. O-Funnel separates the two. It transcribes any XML, JSON, CSV, HTML or key-value text document into one typed tree over five constructors, gated by an oracle that rejects any capture that does not reconstruct its source. Each needed field is declared in the tree's own terms and located by fusing independent evidence (key, path, value shape, synonym, key spelling, record neighborhood, value profile), so the best-supported node wins and a missing field is reported with a reason. Data no requirement claims becomes residue that a funnel traces back to the requirements to learn new key aliases. On 34,989 real PubMed records, O-Funnel matches a hand-written parser (F1 1.00). After a five-element schema rename, the parser's regular expressions fall to 0.20 while O-Funnel stays at 1.00, with every capture verified complete. On constructed suites that isolate regex failure modes it raises F1 from 0.43 to 1.00, and from 0.80 to 0.94 after self-improvement; on held-out schema-matching instances it is competitive with classical matchers without training. O-Funnel is a dependency-free Python library (pip install ofunnel).

cs.DB↗