arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

Spatio-Temporal Wireless-Optical Planning for Multi-UAV Networks

Multi-unmanned aerial vehicle (UAV) networks in urban low-altitude environments couple UAV mobility, wireless access, and optical backhaul resources. Existing path-planning methods optimize flight distance or wireless signal quality, but can still concentrate traffic on shared optical backhaul links. We present Spatio-Temporal Wireless-Optical (STWO) planner, a backhaul-aware path-planning algorithm that jointly considers flight distance, wireless link quality, and time-varying optical-link offered-load ratio. STWO updates backhaul occupancy during sequential multi-UAV planning, allowing later UAVs to avoid congested optical paths while maintaining wireless connectivity. Experiments show that STWO reduces peak optical-link offered-load ratio by up to 56.4\% and congestion ratio by up to 72.6\% under dense UAV deployment, demonstrating the importance of wireless-optical awareness for reliable multi-UAV transmission.

cs.NI↗

TripleFlow: Training-Free Video Object Removal by Bridging Residual Editing and Native Generation

Video object removal presents a uniquely difficult editing challenge. Because a removal prompt specifies only what to erase rather than what to generate, the model must infer and reconstruct a highly specific occluded background entirely from the surrounding context. Existing training-free methods struggle with this because their editing mechanisms act primarily as localized erasers. They fail to actively synthesize the missing background details and often leave behind ghosting artifacts. To solve this, we propose TripleFlow, a training-free framework that tightly couples erasure and generation. It coordinates a source flow, a residual flow, and a synthesis flow throughout the entire process. By reusing a single target prediction, the residual flow isolates and suppresses the object, while the synthesis flow independently reconstructs the occluded background. Crucially, TripleFlow injects this newly synthesized background back into the editing trajectory at every step. This continuous feedback loop ensures that the generated structures actively guide the removal process, achieving seamless completion that is spatiotemporally consistent with the unedited scene. Extensive evaluations across five challenging benchmarks demonstrate that TripleFlow establishes a new state-of-the-art, significantly outperforming existing baselines in both reconstruction fidelity and temporal consistency.

cs.CV↗

Heavy Dark Baryons at Large $N$: Self-Interactions across All Scales

We show that heavy dark baryons in a confining $SU(N)$ sector with a moderately large number of colors, $N\sim10$--$100$, can generate self-interactions spanning the full range of astrophysical velocities. Their chromo-electric polarizability induces van der Waals interactions that yield $σ/m_{\rm DM}\sim30$-$100~\mathrm{cm^2/g}$ at dwarf-galaxy velocities, decrease to the values required by galaxy clusters, and rise to $100$-$1000~\mathrm{cm^2/g}$ at velocities of a few $\mathrm{km/s}$, as suggested by the second perturber of JVAS B1938+666. Remarkably, for a given ratio $m_Q/Λ_d$, these observations can be used to infer the value of $N$. The resulting self-interaction phenomenology favors GeV-scale dark baryons, independently of the relic-abundance requirement. At these masses, a dark baryon asymmetry comparable to the visible one can account for the observed dark-matter abundance.

hep-ph↗

The Exact Comass Criterion for the Tsai--Wang Form

For a map $F:Ω\subset\mathbb{R}^n\to\mathbb{R}^m$, let $Θ(F)$ be the $n$-form on $Ω\times\mathbb{R}^m$ introduced by Tsai and Wang [arXiv:2604.04336]. If $λ_1,\ldots,λ_{r}$ are the nonzero singular values of $d F$ at any fixed point, we prove that $$ Θ(F) \text{ has comass one if and only if }~ \mathcal{S}(λ):=\sum_{i=1}^{r}\frac{λ_i^2}{1+λ_i^2}\leq1. $$ Note that if $\text{rank} d F\le1$, $\mathcal{S}$ is always less than $1$. We also prove a dichotomy that if $\mathcal{S}\leq 1$, then either $\mathcal{S} < 1$ or $\mathcal{S}\equiv1$. When $\mathcal{S}<1$, the graph tangent plane is the only calibrated plane, generalizing the corresponding result for hypersurfaces. When $\mathcal{S}\equiv1$, we prove that $F$ is affine or of rank $2$.

math.DG↗

Quotient Morita Theory with Applications

We study equivalences between quotient categories of module categories associated with Gabriel topologies. We establish necessary and sufficient conditions for a Morita context to induce an equivalence between quotient categories and characterize such equivalences in terms of Morita contexts after passing to rings of quotients. Motivated by Artin--Zhang's noncommutative Serre theorem, we introduce a notion of ampleness relative to a Gabriel topology and use it to characterize quotient categories of finitely generated modules under suitable noetherian hypotheses. As an application, we establish a sufficient condition under which the noncommutative Auslander theorem holds for actions of finite-dimensional Hopf algebras on AS-regular algebras without assuming that the Hopf algebras are semisimple.

math.RA↗

Spectroscopic Evidence for Nontrivial Band Topology in Superconducting FeTe/MnTe Heterostructure

FeTe has long been regarded as a nonsuperconducting antiferromagnetic metal with trivial band topology, but recent advances in stoichiometry control have begun to challenge this picture. Here we use higher-order epitaxy on zinc-blende MnTe to stabilize near-stoichiometric FeTe with strongly suppressed interstitial Fe and a superconducting transition onset near 13 K. Angle-resolved photoemission spectroscopy reveals markedly enhanced quasiparticle coherence, well-defined Fe-derived hole bands, and a nearly two-dimensional Dirac-cone-like state near the Fermi level. First-principles calculations identify an inversion between odd- and even-parity bands, yielding nontrivial $Z_2$ topology and a Dirac surface state consistent with experiment. These results elucidate the intrinsic electronic structure of stoichiometric superconducting FeTe and provide evidence for nontrivial band topology, positioning FeTe/MnTe as a promising platform for exploring topological superconductivity.

cond-mat.supr-con↗

Training-Free Affinity Fusion of Neural and Embedding-Based Speaker Diarization

Speaker diarization systems based on speaker embeddings and neural diarization exploit complementary forms of speaker information, but their intermediate representations are not directly compatible. We introduce Training-Free Affinity Fusion (TFAF), which integrates the speaker structure inferred by a neural diarizer into an embedding-based diarization system. The neural speaker partition is used to condition local speaker representations, from which we construct a continuous affinity matrix and combine it with the embedding-based acoustic affinity before a single global clustering step. The method requires no additional training, shared embedding space, speaker-label alignment, or hard transfer of the neural diarizer's speaker count. Experiments on AMI and CALLHOME show consistent DER improvements over both constituent systems; on AMI, fusion also improves speaker-attributed transcription. Ablations show that the neural speaker partition accounts for most of the gain, while retaining the continuous embedding-based affinities provides additional benefit over hard partition fusion.

cs.SD↗

Pair rationality and top trading cycles on single-peaked and single-dipped domains

In a recent study, Ekici and Yenmez (2026) characterize top trading cycles (TTC) on the unrestricted domain by strategy-proofness, individual rationality, and pair rationality, uncovering a correspondence between the axiomatic foundations of TTC and deferred acceptance, matching theory's two canonical rules. We show that their characterization fails on the single-peaked domain, but it holds on the single-dipped domain even without strategy-proofness. They also show that pair rationality can be replaced in their characterization by two weaker conditions concerning agents' first and second choices. In contrast, this replacement fails on the single-dipped domain, even when strategy-proofness is imposed.

econ.TH↗

QuanVI: Score-based Variational Inference via Quantum Maximally Mixed States

Score-based variational inference (VI) provides an alternative to Kullback--Leibler (KL)-based VI by minimizing the Fisher divergence between the variational distribution and the target. A prior score-VI approach formulates this optimization as an eigenvalue problem, with the variational distribution constructed from low-energy eigenstates. However, this eigenvalue-based formulation faces two high-dimensional obstacles: an intractably large parameter count due to exponential scaling and non-uniqueness of individual eigenvectors in degenerate or nearly degenerate low-energy subspaces. We propose QuanVI, a scalable quantum-inspired algorithm that combines a mixed-state density-operator formulation with a quantum tensor network (QTN) parameterization using the matrix product operator (MPO) structure. In degenerate low-energy subspaces, the density-operator formulation represents the subspace by its maximally mixed state rather than relying on a non-unique individual eigenvector, while the QTN parameterization compresses the density operator to avoid exponential parameter growth. Experiments and ablations show that QuanVI agrees with exact solutions in low dimensions and scales to high-dimensional synthetic and Bayesian posterior-approximation benchmarks, including challenging non-Gaussian targets.

cs.LG↗

Componentwise linearity, Fröberg's analogue, and classification of linear support-two monomial ideals

In this paper, we investigate the componentwise linear property of support-two monomial ideals. Our first main result shows that if $I$ is a support-two monomial ideal, then its underlying simple graph $G_I$ is co-chordal; equivalently, by Fröberg's theorem, $\sqrt{I}$ admits a linear resolution. This phenomenon is quite rare for general monomial ideals. In fact, there exist monomial ideals with linear resolutions whose radicals fail to have linear resolutions, even when the radical is the edge ideal of a graph. For any monomial ideal $I$, one always has $μ(I)\geq μ(\sqrt{I})$. In the literature, support-two monomial ideals satisfying $μ(I)=μ(\sqrt{I})$ are of special interest, as they include edge ideals of simple graphs, weighted oriented graphs, edge-weighted graphs, and vertex-weighted graphs. We refer to such ideals as minimal support-two monomial ideals. We explicitly characterize all minimal support-two monomial ideals, as well as their powers, that admit linear resolutions. Next, we classify the linearity of non-minimal ones. Consequently, we obtain a complete classification of linear support-two monomial ideals, which shows that the property of being linear does not depend on the characteristics of the base field for support-two monomial ideals.

math.AC↗

Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds

Persistent spatial memory enables embodied agents to navigate familiar environments across repeated visits. However, targets may move while unobserved, including during navigation, making remembered locations unreliable by the time an agent arrives. Despite advances in memory retrieval and state prediction, accounting for continued hidden world evolution and revising beliefs under limited visibility remain challenging. We study Evolving-World Navigation, where agents infer target locations from intermittent observations, predict their states at inspection time, and revise beliefs using visual evidence. We propose EvolvingNav, which constructs a time-indexed belief from timestamped 3D object histories through a structured persistence-relocation model. The belief distinguishes persistence at the last observed location from relocation to alternative locations and retains probability mass outside the known candidate set. An event-driven filter propagates the current belief as time elapses, forecasts target occupancy at candidate inspection times, and incorporates new RGB-D evidence. Negative observations downweight location hypotheses according to calibrated, visibility-conditioned detection probabilities, while evidence tracking prevents repeated use of the same observations. A frozen, zero-shot vision-language controller uses the updated belief to choose actions and replan. We further introduce EvoWorld-Bench, a benchmark grounded in human activity traces, comprising 54 scenes and 803,680 tasks with controlled changes before and during navigation. In simulation and real-robot experiments, EvolvingNav improves navigation success and search efficiency over the evaluated baselines. Paired experiments show the clearest gains under learnable temporal patterns, while ablations demonstrate the value of preserving uncertainty and incorporating visibility-aware evidence.

cs.AI↗

Deep Learning-Based Tri-Hybrid Multi-User MIMO Precoding: The Blessing of EM-Reconfigurable Antennas

Electromagnetic (EM)-reconfigurable antennas provide multiple candidate radiation patterns per element, thereby introducing an additional EM-domain degree of freedom. Integrating radiation-pattern reconfigurability, realized as EM-domain precoding, with conventional hybrid analog-digital precoding yields tri-hybrid multiple-input multiple-output (MIMO) precoding, which can substantially improve the spectral efficiency of wideband multi-user MIMO orthogonal frequency-division multiplexing (OFDM) systems. However, the joint design of EM, analog, and digital precoding remains challenging. To address this challenge, we propose a tri-hybrid precoding network (Tri-PNet) based on Conformer, an emerging neural architecture that combines the local modeling strength of convolutional neural networks with the global dependency modeling of Transformers. Furthermore, two representative radiation-pattern modes, i.e., the non-regular mode and the 3rd Generation Partnership Project (3GPP) Technical Report (TR) 38.901 mode, are investigated. Tri-PNet is trained in an unsupervised manner to jointly learn EM, analog, and digital precoding by maximizing the average sum spectral efficiency. Its radiation-pattern selection network (RPSNet) employs a Conformer encoder to capture both local and global frequency-domain correlations, whereas its hybrid analog-digital precoding network (HPNet) combines cross-attention and dual-path processing with singular-value-decomposition (SVD) and zero-forcing (ZF) priors. Simulation results under both radiation-pattern modes demonstrate that Tri-PNet outperforms random EM precoding and conventional hybrid MIMO without EM precoding, approaches the greedy EM precoding search scheme with substantially lower online complexity, and remains robust to imperfect channel state information (CSI).

eess.SP↗

Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at https://github.com/Zhihan72/TPAE.

cs.AI↗

Edgeworth expansions of extreme eigenvalue distributions for random unitary ensembles

In this paper, we establish Edgeworth expansions of extreme eigenvalue distributions for two types random unitary ensembles at the spectral edges and reveal certain universal structures for the correction terms. More precisely, we show the correlation kernels for the Gaussian-type unitary ensembles and Laguerre-type unitary ensembles admit full expansions at the soft edge and the hard edge, respectively. The coefficients of the correction terms are given by finite sums of Airy function or the Bessel functions of the first kind and their derivatives with polynomial coefficients. By lifting the kernel expansion to the associated Fredholm determinant, we further obtain full expansions for the largest eigenvalue distribu- tion of the Gaussian-type unitary ensembles and the smallest eigenvalue distribution of the Laguerre-type unitary ensembles. In addition, the first correction term therein involves the derivatives of the leading term.

math.PR↗

Quantum Frozen--Oseen homotopy analysis method with LCHS for solving nonlinear partial differential equations

Nonlinear partial differential equations (PDEs) underpin computational fluid dynamics, yet resolving nonlinear transport on fine grids remains computationally demanding. Quantum linear-evolution algorithms offer a possible route to large-scale simulation but cannot directly propagate nonlinear coupling. Here we develop a Frozen--Oseen quantum homotopy method (FOQHAM) framework that links transport-aware auxiliary-operator selection to the size of the resulting linear representation. We freeze the full Fréchet derivative at a prescribed flow profile, retaining transport and profile-gradient coupling and close each prescribed finite-order homotopy hierarchy as a fixed affine linear system by product lifting. We formulate its propagation through Linear Combinations of Hamiltonian Simulations (LCHS), without outer homotopy iterations or profile updates. For spatially semidiscrete equations, we establish local approximation bounds and exact closure of the finite-order hierarchy. Classical tests on Burgers, Korteweg--de Vries, Zakharov--Kuznetsov, and one- and two-dimensional isothermal compressible Navier--Stokes equations provide numerical evidence of accurate non-iterative approximations in the tested regimes. The framework thus connects nonlinear PDE approximation to quantum linear evolution and offers a conditional path toward quantum acceleration.

math.NA↗

Marcinkiewicz-Zygmund inequalities for hemispherical $t$-designs

In this paper, we introduce a new set of $\mathbb{L}_2$-orthogonal polynomials on the hemisphere $\mathbb{S}^d_+:=\{\mathbf{x}\in\mathbb{R}^{d+1}:\|\mathbf{x}\|=1, \mathbf{x}\cdot\mathbf{e}_{d+1}\geq0\}$ and establish $\mathbb{L}_1$ Marcinkiewicz-Zygmund inequalities on the hemisphere for even and odd spherical polynomials. Hemispherical $t$-designs provide equal weight cubature rules on the hemisphere which are exact for polynomials up to degree $t$. We propose variational characterizations of hemispherical $t$-designs by the $\mathbb{L}_2$-orthogonal polynomials. Moreover, based on the Marcinkiewicz-Zygmund inequalities and the $\mathbb{L}_2$-orthogonal polynomials, we prove that for each $n\geq C_d t^d$, there exists a hemispherical $t$-design on the hemisphere with $n$ points, where $C_d$ is a constant depending only on $d$.

math.NA↗

A Width-Matched Comparison of Hybrid Quantum-Classical Self-Supervised Learning for Fingerprint Recognition

Fingerprint recognition is a widely deployed biometric, but supervised training requires large labeled enrollment sets. Self-supervised learning (SSL) removes this requirement, and hybrid quantum-classical models have been proposed to enrich the learned representations. Prior quantum SSL studies consider a single contrastive objective, so it is unclear whether reported benefits depend on the objective or can be attributed to the quantum circuit. We insert the QuFeX quantum feature-extraction module into three SSL frameworks, the contrastive SimCLR and MoCo v2 and the non-contrastive BYOL, and compare each hybrid with its classical counterpart at matched representation width (8 features, equal to 8 qubits) on the SOCOFing fingerprint dataset, with a CIFAR-10 control, using k-nearest-neighbor identification on encoder features. In single-run experiments the hybrid scores clearly higher for both contrastive objectives, whereas for BYOL a multi-seed analysis shows no reliable difference, suggesting that any benefit depends on the SSL objective. A hardware-efficient circuit (QNet) does not show the same gain. We examine whether the gains can be attributed to the quantum circuit, considering circuit architecture, trainable parameter count, nonlinearity, and the classical simulability of 8-qubit circuits.

quant-ph↗

Asymptotic Properties of Support Vector Machines in High-Dimension, Low-Sample-Size Settings under a Spiked Model

In this paper, we consider asymptotic properties of the support vector machine (SVM) in high-dimension, low-sample-size (HDLSS) settings under a spiked model. The existing theory of the SVM in the HDLSS context relies on the geometric representation of HDLSS data, which requires that the eigenvalues of the covariance matrices are not dominant. We first show that the geometric representation does not hold under the spiked model. We show that the Gram matrix of HDLSS data converges in distribution to a random matrix, namely, the HDLSS data converge to a random configuration in a finite-dimensional space whose dimension is given by the number of the spikes. We show that the misclassification rates of the SVM do not tend to zero, that is, the SVM does not hold the consistency property. We also show that the bias-corrected SVM (BC-SVM) does not give preferable performance in this setting because the bias term itself should be modified. In order to overcome such difficulties, we propose a spike-corrected SVM (SC-SVM). We show that the SC-SVM holds the consistency property when the sample size goes to infinity, and that the growth of the sample size is essential in the sense that any projection-based procedure fails when the sample size is fixed. Finally, we check the performance of the classifiers by numerical simulations.

stat.ML↗