arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 325 records · Page 18Linked to original sources

Implicit Behavior Coordination from Sub-Task Demonstrations by Exploiting Overlap-Induced Multimodality

Long-horizon robotic rearrangement is commonly formulated as a skill-sequencing problem, where distinct behaviors are explicitly represented and coordinated by a planner or high-level policy. We investigate whether such explicit behavior identities and sequencing interfaces are necessary at all. We introduce implicit behavior coordination from sub-task demonstrations, where separately collected behaviors are coordinated without behavior identity labels, complete-task demonstrations, or task-ordering supervision. Our key observation is that overlap between sub-task demonstrations induces multimodal action distributions that need not be resolved through explicit behavior partitioning. Instead, this overlap-induced multimodality can be exploited as a coordination resource. We instantiate this idea with a shared Flow Matching policy that preserves multiple action modes and critic-guided in-sample planning that propagates task value across demonstrations and selects task-relevant modes. Experiments in Habitat and on a real robot show that implicit behavior coordination remains effective under reduced cross-behavior overlap, larger behavior mixtures, longer horizons, and execution failures, supporting the idea that long-horizon coordination can emerge directly from sub-task demonstrations without explicitly recovering or sequencing behavior identities.

cs.RO↗

Macroscopic Zero-Mode Manifold Isolated by Quantum Chaos

Chaotic many-body spectra are expected to densely fill their energy window. We show that constrained spin chains with chiral symmetry evade this expectation by hosting an exponentially large manifold of symmetry-protected exact zero modes separated from the surrounding spectrum by a sharp gap at zero energy. The gap is generated by chaotic level repulsion, with width set by the number of zero modes times the mean level spacing. We verify this mechanism in an East-West kinetically constrained chain, develop a minimal random-matrix description, and show how the gap can be detected through linear-response spectroscopy.

cond-mat.str-el↗

Learning to Navigate with Minimal Parameters: Decomposing Visual Navigation Through Closed-Form Geometric Interfaces

Visual navigation policies have grown to hundreds of millions of parameters trained on billions of frames, with geometry, mapping, and control learned implicitly. We propose a decomposed point-goal navigation system in which operations with known closed-form structure, such as projective geometry, occupancy, and coordinate transforms, are computed analytically and serve as interfaces between three small learned modules: an egress predictor that grounds the episode goal as a local subgoal in the current view, a navigation predictor that estimates a goal-conditioned posterior over where trajectories travel, and an endpoint-pinned residual diffusion generator that samples trajectory shapes from this posterior. Only 0.58M out of 23M parameters are trained, on 44k frames, in under one GPU-hour. Across 6060 point-goal episodes in 60 environments, the system attains competitive success rates with the lowest collision rate among evaluated methods. We further show that under this decomposition, the frozen image encoder can be replaced by a 0.54M MobileNetV2 at a -2.0 SR cost, bringing the full system under 1.2M parameters. It also transfers to no-goal exploration by retraining only the 123k-parameter egress head, and its failure modes under sensor corruption are transparent and analytically correctable. We deploy and evaluate the system zero-shot on a low-cost UGV, running navigation and localization on a Jetson Orin Nano in real-time.

cs.RO↗

Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents

Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict. A critic reads each noisy observation and passes only the usable content to the ledger, discarding the rest, so the planner works from a compact context throughout the task. Each run records the state, action, and verdict at every step, and supervised fine-tuning and decision-level reinforcement learning on these trajectories further improve the planner. Omni-Decision achieves state-of-the-art accuracy of 81.4% on OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.

cs.AI↗

A hybrid analytical-PINN model for subsurface simulation of geothermal heat exchangers in heterogeneous underground

Accurate and efficient prediction of subsurface temperature fields is essential for the design and operation of borehole heat exchanger (BHE) systems. Here we develop a parametric hybrid analytical and physics-informed neural network (PINN) framework for long-term multi-BHE simulations in heterogeneous underground. The method analytically extracts the singular line source response and enables the effective training of neural correction associated with subsurface heterogeneity. An explicit parametrization of the thermal conductivity allows physics-informed learning of a single feedforward neural network to generalize across different subsurface conditions. By formulating the correction in borehole-centered relative coordinates, the learned correction can be reused as a universal corrector through spatial and temporal superposition principles. Numerical experiments based on the infinite line source (ILS), finite line source (FLS) and moving finite line source (MFLS) models show that the hybrid method outperforms analytical approximations with stable accuracy over long simulation horizons and achieves orders-of-magnitude speedups over traditional solvers. The proposed framework therefore combines the efficiency of analytical models with the ability of numerical methods to capture heterogeneous subsurface physics, providing a fast and accurate approach for repeated long-term simulation of multi-BHE systems.

cs.LG↗

A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation

The key-value (KV) cache is the dominant memory bottleneck in long-context language model inference. Existing compression methods apply low-rank factorization or quantization independently, without jointly allocating rank and precision under a shared storage budget. We introduce JoLT, a training-free compressor that treats grouped prefill caches as fourth-order tensors and applies partial Tucker decomposition along the token and feature modes, the two axes that carry low-rank structure, while leaving the head and layer modes intact. A rotated low-bit quantizer captures the truncation residual, and a single Lagrangian dual allocates per-group Tucker ranks and residual bit-widths under a global byte constraint. FlashJoLT replaces the exact token-mode SVD with a randomized approximation that matches JoLT within the free zone at a fraction of the compression cost, and a fused Triton decode kernel evaluates attention directly over the stored factors without materializing dense KV tensors. Across five models from four architecture families, covering multi-head attention, grouped-query attention, and mixture-of-experts architecture, JoLT achieves 2 - 3x compression with less than 0.2% perplexity degradation, without retraining. On RULER at 64K context with LLaMA-3.1-8B, retrieval accuracy remains near-lossless through 3x and declines by only 0.90 and 2.40pp at 4x and 5x, respectively. JoLT demonstrates that tensor-aware low-rank decomposition and quantized residuals, unified under a single storage budget, achieve near-lossless KV-cache compression across diverse model architectures without retraining.

cs.LG↗

Simulating the convection in red super-giant stars: wobbling jets in common envelope evolution

We use our newly constructed three-dimensional red supergiant (RSG) stellar model, which also mimics nuclear energy production and photospheric emission, to calculate the stochastic component of the angular momentum of the mass that a companion spiraling within the RSG's envelope accretes during common envelope evolution (CEE). The accreted mass has a fixed-direction angular-momentum component arising from the density gradient in the RSG envelope and orbital motion. The angular momentum component with a stochastically varying direction results from vigorous envelope convection. We do not include the companion's influence on the RSG envelope during the CEE and consider an undisturbed, non-rotating RSG stellar model. We find that the fluctuating angular momentum amplitude can be several times the fixed-axis angular momentum. The total specific angular momentum of the accreted mass easily forms intermittent accretion disks around neutron stars and black holes, but it is only marginally sufficient, or not at all, to form accretion disks around main-sequence stellar companions. The intermittent accretion disks we expect to form will launch wobbling jets with varying axes. We discuss aspects of wobbling jets in the CEE and the grazing envelope evolution (GEE), which might precede the CEE or replace it altogether. Studies have claimed that jets are a crucial ingredient in many cases of CEE, and the standard CEE should include jets that the companion launches, before (like the GEE), during, and/or at the exit from the CEE. Our study supports this claim and emphasizes the importance of wobbling jets.

astro-ph.SR↗

Right q-vector calculus at integral superdimension: localized decompositions and resonance

We specialize the intrinsic right $q$-vector derivative on radial algebras to integral superdimension. The formal dimension is encoded by an independent coefficient $Q$, and formal radial superspace of superdimension $M=m-2n$ is obtained by the coefficient specialization $Q\mapsto q^M$. This gives a rigorous universal calculus and, on finite blocks containing at most $m$ abstract vectors, a faithful coordinate realization on $\mathbb R^{m|2n}$. The localized exterior result is formulated as a Green decomposition by complementary projector images. Whenever the specialized finite determinant is nonzero, full left multiplication yields a determinant-localized right-monogenic Fischer decomposition. Beyond this base-change theory, we determine the exceptional one-vector calculus completely: for $M=-2\ell$ there is one additional singular monomial and one missing image monomial, whereas all other integral superdimensions give a surjective derivative with constants as its kernel. We then prove that the degree-zero Fischer operator is exactly diagonal on exterior blades, obtain its determinant explicitly, classify all support-resonance values in $0<q<1$, and give an exact kernel-rank formula as a sum of support multiplicities. On a block with $N$ auxiliary vectors, an even support rank $p$ has pure multiplicity $\binom Np$ on the support truncation. At an odd-support root, every lower odd factor is nonzero and any simultaneous lower resonance is unique, even, and characterized by one strictly monotone scalar equation. These results distinguish persistent nonpositive-even-superdimension defects from isolated support-dependent $q$-resonances. An appendix records that constant scalar projection of two independent orthogonal right $q$-vector derivatives does not descend to the Hermitian quotient.

math.CV↗

Lattice Boltzmann Methods for Navier-Stokes Equations in General Orthogonal Coordinates for Efficient Flow Simulations using Nonuniform Clustered Grids

Resolving multiscale fluid flows or boundary layers effectively requires the use of nonuniform meshes with local grid clustering. The standard lattice Boltzmann method (LBM), a kinetic theory-based approach for computational fluid dynamics, however, is restricted to the use of uniform Cartesian grids. We present new and improved formulations of the LBM that accommodate continuously varying spatial grids via coordinate transformations to simulate the Navier-Stokes equations (NSE) in the general orthogonal coordinates (GOC). They are constructed using a Chapman-Enskog analysis to specify the equilibrium moments of the distribution functions and the geometric force terms used in the collision step to be dependent on the local metric factors and their spatial derivatives, along with the density, momentum and their fluxes, and some correction terms related to the normal velocity gradients so as to accurately represent the NSE in the GOC. The resulting GOC-LBM importantly maintains the simplicity of the collide-and-stream approach and is Galilean invariant that is free of the cubic velocity artifacts. Our GOC-LBM is general and modular in that it can be used with any collision model with appropriate modifications to the equilibria and forcing terms. We present its implementation details for a variety of collision models while the central moments-based model using multiple relaxation times was found to be the most robust in practical implementations. We validate the GOC-LBM through numerical simulations for various benchmark flow problems. Moreover, we demonstrate significant computational advantages of our approach for a case study on simulating boundary layer flows efficiently that involves coupling the GOC-LBM for the NSE with a new GOC-LB scheme for solving the magnetic induction equation for magnetohydrodynamics (MHD), and for another case study involving orthogonal curvilinear grids.

physics.flu-dyn↗

Lieb-Thirring bounds for Melik-Adamyan canonical Hamiltonians

We study a class of positive matrix Hamiltonians arising from the canonical differential expressions of Melik--Adamyan and appearing in the appendix of Alpay--Gohberg. Let $J$ and $B$ be self-adjoint involutions on $\mathbb C^{2n}$ satisfying $JB=-BJ$, and let $H>0$ satisfy $HJH=J$. For $m>0$ we consider $$ \mathcal A_{m,H}=H^{-1}\left(-iJ\frac{d}{dt}+mB\right) $$ in the weighted space $L^2_H$. A locally absolutely continuous $J$-unitary gauge $Θ$ representing $H$ reduces this expression to the free massive Dirac operator plus the Hermitian coefficient $$ P_{m,Θ}=-iΘ^*JΘ'+m(Θ^*BΘ-B). $$ Whenever this coefficient belongs to $L^2$, the corresponding self-adjoint realization, including its operator domain, is independent of the chosen representing gauge. Minimizing $\int\mathrm{Tr}|P_{m,Θ}|^2$ over the gauge fibre defines an intrinsic energy. A two-sided Birman--Schwinger decoupling, combined with a truncated pseudo-relativistic estimate proved here, gives a $3/2$-moment bound for all eigenvalues in the gap $(-m,m)$ in terms of this energy. The Dirac estimate applies to arbitrary Hermitian matrix coefficients in $L^2$ and requires no sign condition. On the half-line we treat every self-adjoint Lagrangian boundary condition. Two reflection-compatible conditions require no endpoint correction, while an arbitrary condition contributes at most $2nm^{3/2}$. At zero mass, the optimal-gauge energy is computed explicitly in terms of $H^{-1/2}H'H^{-1/2}$. For a scalar hyperbolic-rotation family the massive gauge minimization reduces exactly to a one-dimensional phase functional. We prove existence of a minimizer in the principal phase sector and give an explicit trial phase that strictly and quantitatively improves the positive lift whenever the corresponding first variation is nonzero.

math.SP↗

SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions

In recent years KolmogorovArnold Networks KANs have attracted increasing attention due to their effectiveness in machine learning and scientific computing offering a new paradigm for neural network design In this paper we present SechKAN a novel KAN based on hyperbolic secant sech functions The hyperbolic secant basis is adopted for its smooth bellshaped form localized responses and wellbehaved gradients We employ a 1D linear projection to reduce the number of parameters allowing SechKAN to maintain a model size comparable to that of multilayer perceptrons MLPs Experimental results show the effectiveness of SechKAN on function fitting PDE surrogate modeling and image classification benchmarks including MNIST FashionMNIST CIFAR10 and CIFAR100 On function fitting SechKAN achieves performance comparable to both MLPs and representative KAN variants On PDE surrogate modeling it outperforms MLPs and achieves competitive or better performance than representative KAN variants On image classification benchmarks SechKAN achieves the best performance among the evaluated KAN variants while remaining competitive with MLPs using a comparable number of parameters However SechKAN still incurs higher computational cost than MLPs and some KAN variants Our source code is publicly available at https://github.com/hoangthangta/All-KAN.

cs.LG↗

The quantitative non-unique-product landscape at the global minimum: the Nielsen-Soelberg groups

Nielsen and Soelberg proved that a finite subset $A$ of a torsion-free group with $A\cdot A$ having no unique product satisfies $|A|\ge 8$, and exhibited two groups, here $G_1$ and $G_2$, attaining the bound. Nothing quantitative was known about these extremal configurations. We construct exact, independently verified models of both groups and compute the first quantitative invariants at the global minimum. In $G_1$ no $8$-element symmetric witness lies in the radius-$6$ ball ($933$ elements, certified infeasible), while the Nielsen-Soelberg witness lies in the radius-$7$ ball: the global minimum is spread out. In $G_2$, with its natural eight-generator metric, the witness and its inverse are the only two non-UP $8$-sets in the radius-$1$ ball, and the unique-product staircase takes the value $0$ at $n=8$ but $1$ at $n=9$ -- the first known minimizer whose square has exactly one uniquely represented element, so the simultaneous failure of t.u.p. and u.p. seen in the Promislow group is not universal. No $(7,9)$ two-sided witness exists in the searched balls, so the Nielsen-Soelberg profile bound may not be sharp. Finally we treat the universal group $G_3$. Its structure is known -- Soelberg's thesis identifies an index-$8$ Heisenberg subgroup of step $8$ and proves torsion-freeness, and Gardam, studying the same group as an amalgam of Klein bottle groups, shows it to be virtually nilpotent but not virtually abelian -- and what we add is a model in search coordinates in which balls can be enumerated. In it we reproduce the Nielsen-Soelberg two-sided pair and exhibit a symmetric $15$-element witness whose trivial-coset singleton generates the centre of that Heisenberg subgroup. It is rigid and rare: within $B(5)$ the size $15$ is exactly minimal, the coset profile is forced, and exactly four such witnesses exist in $B(4)$, one orbit. Hence $m_1(G_3)\in[8,15]$ against $m_2(G_3)=16$.

math.GR↗

Exceptional supersphere integration and logarithmic Pizzetti kernels

We study orthosymplectically invariant supersphere integration at the exceptional superdimensions $M=-2u$, where the harmonic Fischer structure becomes nonsemisimple and the Pizzetti pairing degenerates. For the meromorphically continued homogeneous inverse kernels we obtain the generating function $$ \mathscr G_μ(ρ;x,y) =\frac{Γ(μ/2)}{2π^{μ/2}} \bigl(1+ρ\{x,y\}+ρ^2x^2y^2\bigr)^{-μ/2}. $$ At $μ=-2u$, its Laurent expansion has a polynomial residue and a logarithmic finite part. We prove that these coefficients recover the complete degreewise duality structure on a fixed superspace with nonzero bosonic dimension. In degrees $k\le u$, the residue inverts a canonical renormalized pairing on $\mathcal P_k$. In the collision range $u<k\le2u$, the ordinary pairing has radical $(x^2)^{k-u}\mathcal P_{2u-k}$; the finite part reproduces the quotient, while the residue reproduces the radical after transport from the reflected degree. For $k\ge2u+1$, the finite part is the ordinary inverse kernel. We also establish the nondegenerate head--socle pairing on the generalized harmonic modules. As an application, we derive covariant right--left radial $q$-monogenic zonal symbols and identify precise degree-one obstructions to transferring scalar Pizzetti reproduction through a one-sided $q$-Fischer projection.

math.CV↗

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

Recent advances in automated audio captioning (AAC) are driving a shift from monolithic sentences toward structured formats that disentangle acoustic and semantic properties, such as timestamped captions for different sound events. Such representations can support faceted sound search for creators and richer access to auditory information for Deaf and Hard of Hearing people. Yet, it remains unclear how to meaningfully evaluate these hybrid, structured captions. We propose an evaluation framework for structured audio descriptions, spanning five complementary axes: tag sets, descriptions, reasoning, numeric measurements, and spectral profiles. The framework combines large language model (LLM) judges for semantic fields with deterministic metrics for temporal and acoustic attributes. To validate these metrics, we introduce controlled perturbations that apply typed, graded changes to ground-truth annotations. Results show that the proposed metrics remain robust to meaning-preserving paraphrases while responding to genuine semantic and acoustic corruptions, enabling more reliable evaluation of structured captions.

cs.CL↗

Disentangling mixed neutron fields: multi-source identification from few detected events

Identifying neutron-emitting materials is central to nuclear nonproliferation, safeguards, nuclear forensics, and emergency response, yet remains difficult when several sources contribute simultaneously: relevant fission, $(α,\text{n})$, and fusion sources emit broad, strongly overlapping energy distributions, and the associated spectral inversion is severely ill-conditioned. Here we demonstrate quantitative identification of mixed neutron fields directly from scatter-based (recoil) spectroscopy measurements, together with simultaneous estimation of the emission rate of each contributing source and a rigorous statistical confidence level for every candidate source combination. Using a compact $21.6\,\mathrm{cm}^3$ organic-glass scintillator spectrometer, we correctly identify Cf-252, a deuterium--deuterium (DD) neutron generator, and their mixture with decisive statistical support ($>\!4σ$), and further resolve a weak deuterium--tritium contaminant in the nominal DD generator field. High-fidelity Monte Carlo simulations spanning exhaustive single-, two-, and three-source mixtures show that identification requires remarkably little information: between $\mathcal{O}(10^1)$ and $\mathcal{O}(10^6)$ detected recoil events, set primarily by spectral similarity, mixture complexity, and emission-rate imbalance. For the compact spectrometer used here, this corresponds to acquisition times as short as a few minutes. These results substantially extend the operational reach of simple single-volume neutron spectrometers, enabling rapid, quantitative, and confidence-calibrated attribution of complex neutron fields in field-deployable instruments.

physics.ins-det↗

A Multi-level Information Integration Framework for Physically Verifiable Fault Diagnosis of Rotating Machinery

Integrating multi-level information, from physical models through data-driven diagnostics to natural language reasoning, into verifiable decision chains is a growing need in intelligent manufacturing. In bearing fault diagnosis, taken here as a representative testbed, the standard output is a class label and a confidence score derived from the classifier's own distribution, offering limited means of comparison against independent physical knowledge. Meanwhile, language models increasingly used for maintenance communication may introduce unsupported content. This work addresses both limitations from the output side. The proposed Diagnostic Evidence Network (DENet) is an encoder-agnostic multi-task framework that extends the output to a structured evidence record: the classification, a predicted characteristic frequency comparable against the theoretical value determined by bearing geometry and shaft speed, and a temporal localization of transient impulses inspectable on the raw waveform. Across four encoders and three public datasets, this evidence incurs no statistically significant accuracy cost, with a frequency error of about 6 Hz on 1,024-point segments. The deviation between predicted and theoretical frequency constitutes a label-free, inference-time validation signal. It detects misclassifications with AUROC of 0.970 and 0.871, and retains separation within the high-confidence subset. Finally, a QLoRA-adapted language model renders DENet's evidence into traceable maintenance reports without contributing diagnostic decisions, reducing unsupported-claim rates from 10-12% to 2% with no fabricated quantities observed.

cs.LG↗

No Free Lunch in Flow Surrogates under Time-Varying Boundary Conditions: A Two-Regime Study

We test whether an architecture that succeeds on a simple flow regime also succeeds on a richer one, with each trained separately on each regime. We explore two transient flows under time-varying boundary conditions: the three-dimensional slurry film in chemical-mechanical planarisation (CMP), central to semiconductor manufacturing, and the two-dimensional Kármán vortex street (KVS). Eight surrogate models on one shared pipeline differ in whether they learn the full field or a latent representation, and in whether they predict in one shot or step by step. No single architecture wins both regimes. On the film, a one-shot full-field model reconstructs the cumulative wall shear stress to 2.7% relative error. On the wake, a latent autoregressive DeepONet retains 90% of the shedding power that direct and one-shot models damp to almost zero. The treatment of time decides the outcome. The self-sustained wake calls for autoregressive feedback and the boundary-driven film for a direct map. Pointwise RMSE hides the damped oscillation on the wake, compresses the sixfold lead on the film's process target, and picks the damped model under wake extrapolation. The evaluation scores five physical questions. Trained surrogates answer queries 10^3 to 10^4 times faster than the finite-element solver and pay off from the first query beyond the training set on the film and from the third on the wake. Neither the winning architecture nor its validation holds across regimes. The choice of surrogate should follow the dynamical character of the target flow, and its validation should resolve the failure modes.

math.NA↗