arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Generative Atmospheric Super-Resolution from Heterogeneous In Situ Observations through Composable Interfaces

Atmospheric observations are sparse, heterogeneous, and unevenly distributed, whereas many generative atmospheric models learn distributions over regularly gridded multivariate states. Once pretrained, diffusion models can supply atmospheric priors that can be combined with observation-derived likelihood factors in a Bayesian formulation. However, these observation sources differ substantially in geometry and sampling density, complicating the consistent use of their observations within a common inference framework. Here, we formulate this reconstruction problem as generative atmospheric super-resolution and introduce composable observation interfaces for conditioning a single pretrained 13-variable atmospheric diffusion model. The interfaces convert sparse radiosonde (R), clustered aircraft (A), and dense irregular surface-station (S) observations into source-specific likelihood factors that specify where observations constrain the gridded state, how residuals are counted under uneven sampling, and how strongly each source guides posterior sampling. We developed the aircraft and surface observation interfaces using 2019 observations and evaluated the selected interfaces throughout 2020 without further tuning. Compared with reconstructions conditioned only on radiosonde observations, the composed R+A+S interface reduces RMSE evaluated against ERA5 by $9.24\%$ across all 13 state variables over the CONUS domain. The aircraft and surface factors provide complementary improvements in upper-air and surface variables. The R+A+S combination also lowers the Continuous Ranked Probability Score (CRPS), while evaluations at held-out aircraft and surface-station observations show reduced prediction errors. Together, these results demonstrate a modular route for conditioning a pretrained atmospheric generative prior on heterogeneous in situ observations without retraining the underlying model.

cs.LG↗

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-grained categories, largely surpassing the category diversity of existing popular RGB-D benchmarks (e.g., NYUv2 with 40 classes and SUN RGB-D with 37 classes). With such enriched semantic coverage, we expect to promote the learning of more generalizable segmentation models. (2) Larger Scale. Compared with current benchmarks, RGBD20K offers 20,000 RGB-D image pairs, providing a substantially larger training resource that benefits the development of more powerful deep models. (3) High-Fidelity Annotation. We perform rigorous re-evaluation and correction of existing labels to resolve long-standing annotation noise, resulting in a clean and reliable ground-truth foundation. Furthermore, we propose a novel score-purified fusion (SPF) method, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of our approach in leveraging high-quality multimodal information for RGB-D semantic segmentation. The dataset is here: https://github.com/ShaohuaDong2021/RGBD20K/.

cs.CV↗

Exploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge

Onboard vision-language models could enable satellites to answer queries directly, but exhaustive tiled inference over high-resolution imagery is slow and energy-intensive. We identify answer-invariant token redundancy (AITR): image tiles and vision tokens that can be removed without changing the final answer. We present Rift, a two-stage system that performs query-conditioned tile pruning followed by elastic prefill to reduce token budget. We evaluate it on LLaVA-1.5 7B running on Jetson AGX Orin. Compared with exhaustive tiled inference, Rift reduces energy by 78% and latency by 69%, while increasing accuracy from 45% to 73%.

cs.CV↗

Frustrated junctions in interfacial networks

Networks of interfaces in materials and living systems evolve through motion and rearrangement of interfaces and junctions. Local equilibrium at a junction requires interfacial force balance. We show that in three dimensions the classical Herring condition for triple lines is necessary but insufficient: four triple lines meeting at a quadruple node may each satisfy Herring equilibrium, while the four conditions remain mutually incompatible. Such a frustrated node has no admissible local equilibrium geometry. In the semi-isotropic limit, the six interfacial energies must form the edge lengths of a non-degenerate tetrahedron, with constructibility determined by the Cayley-Menger determinant. This hidden compatibility constraint arises from three-dimensional geometry and applies broadly to foams, polycrystals, tissues and other interfacial networks.

cond-mat.mtrl-sci↗

Simple Torque-Observation Alignment for Zero-Shot Sim-to-Real Grasping with a Direct-Drive Gripper

Torque observations in reinforcement learning remain challenging because simulated and measured torque differ in scale, offset, and noise. In this paper, we propose a simple torque observation alignment method for robots with direct-drive (DD) actuators, in which motor current maps linearly to joint torque through a motor-type-specific torque constant K_tau. First, dynamometer calibration identifies K_tau* and corrects the scale mismatch between simulated and real torque. Second, the method uses delta_tau(t) = tau(t) - tau(t-1) as the observation in both domains to eliminate the constant offset instead of using the direct torque tau(t), which carries a domain-dependent bias. Third, Gaussian noise obtained from the dynamometer measurement data is injected during the learning process. To validate the proposed method, we train a teacher-student grasping policy entirely in simulation and deploy the distilled student on a multifingered DD gripper. The deployed policy performs proprioceptive grasping using only joint positions and torque differences. We conduct an ablation study comparing the proposed method with alternative alignment variants on nine in-distribution (ID) objects. The proposed method achieves 100% grasp success. These results demonstrate that the proposed alignment method improves the robustness of zero-shot policy transfer on the DD gripper against real-world torque-observation mismatches.

cs.RO↗

Paging the Experts: A Reproducible Characterization of Flash-Backed MoE Inference on iPhone

Sparse activation reduces mixture-of-experts computation without eliminating the need to store all experts. We present Routide, a Swift/MLX runtime that executes the text path of a pinned public Qwen3.6-35B-A3B quantized checkpoint while keeping expert weights in iPhone storage and a byte-budgeted subset in memory. We characterize cache-policy sensitivity, numerical comparison boundaries, and measurement limits. Across five recorded 128-token workloads, fixed-route replay gives 0.00% demand hits with a 512 MiB LRU cache, 18.80% with seeded random eviction at the same budget, and 38.58% with 576 MiB LRU. The apparent capacity cliff is therefore a policy/workload interaction, not a universal memory requirement. Same-runtime Mac controls preserve generated sequences across eviction and asynchronous prefetch, including 2,560 exact token comparisons and 10,334 speculative loads. In contrast, complete resident-Python versus recorded-phone sequences disagree on all five tested cases, precluding a general numerical equivalence claim. Two separately scoped iOS 27 memory protocols observe sampled process-footprint peaks of 1.87-2.32 GiB on short prompts and 2.39-2.73 GiB on one longer prompt. We retain a thermal stopping event, negative timing comparisons, and a single qualified whole-device power estimate. These results establish bounded feasibility and identify limitations that a deployment claim must not hide.

cs.PF↗

Proof of the positive trace gap conjecture

We prove that a lattice $Γ$ in $\mathrm{PSL}_2(\mathbb{R})$ or $\mathrm{PSL}_2(\mathbb C)$ has positive trace gap, meaning that its traces are uniformly separated, if and only if it is derived from an admissible quaternion algebra. For cocompact Fuchsian groups, this proves the positive trace gap conjecture attributed to Sarnak by Geninska and Leuzinger in 2008. The same characterization by quaternion algebras holds if the difference set of traces is not dense. Our method also gives a similar result for lattices in $\mathrm{SL}_d(\mathbb R)$, for every $d\ge3$: a lattice $Γ<\mathrm{SL}_d(\mathbb R)$ has positive trace gap if and only if all its traces are integers. Equivalently, after conjugation, it has finite index in the norm-one group of an order in a central simple algebra of degree $d$ over $\mathbb Q$ that splits over $\mathbb R$. We also discuss spectral consequences and other applications of the results and techniques.

math.GR↗

Constraining Energy Density Functionals via Bayesian Analysis of Nuclear Densities

In nuclear many-body physics, energy density functional (EDF) theory is one of the most powerful approaches for describing finite nuclei and nuclear matter. However, its predictive capability depends on calibrating model parameters to experimental and observational data. In this work, we investigate an alternative approach: Constraining the parameters with the continuous density profiles of finite nuclei obtained from ab initio calculations. We apply Bayesian analysis to infer the parameters of Skyrme EDF from the density profiles and binding energies of 16O, 40Ca, and 48Ca. We show that the data effectively constrain the parameters associated with the properties of uniform nuclear matter, whereas those governing non-uniform nuclear matter remain partially constrained and require additional input. Furthermore, using the inferred parameter distributions, we successfully predict the density profiles and binding energy of 208Pb, which is excluded from the training data. This demonstrates the predictive capability of the framework. In conclusion, these results establish Bayesian analysis of density profiles as a promising route for incorporating accurate ab initio results of light nuclei into EDF development and strengthening the connection between both approaches.

nucl-th↗

Phylodynamic inference with the bounded coalescent: a point process perspective

The coalescent is a central framework in population genetics for modelling the ancestral relationships among sampled individuals through a genealogy, represented as a rooted and ranked binary tree. In this model, lineages coalesce at a rate inversely proportional to the effective population size, a time-varying quantity of primary interest. The bounded coalescent conditions genealogies on the time to the most recent common ancestor being bounded above by a fixed time. This model is useful in various contexts, such as phylodynamics of infectious diseases with known introduction times and single-cell lineage tracing in synthetic barcoding experiments. To our knowledge, there is no existing tool that infers variable effective population size trajectories under the bounded coalescent. We view estimation under the bounded coalescent as equivalent to estimation of the intensity function of an inhomogeneous point process. We provide an efficient algorithm for coalescent simulation under the bounded coalescent using point process methods, retaining the exactness of naive rejection sampling while substantially reducing computational cost and avoiding repeated numerical inversion of the bounded cumulative hazard. We then develop a Markov chain Monte Carlo procedure for posterior inference of effective population size trajectories that avoids discretization of the likelihood integrals. In simulations, conditioning on the bound reduces the median sum of squared errors in two of three settings, with less favourable results in the most rapidly varying setting. We illustrate the method using severe acute respiratory syndrome coronavirus 2 sequence data from Washington State.

stat.ME↗

Reconstruction of Black Hole Metric with Gravitational Shadows

Based on a two-parameter perturbative framework, we derive perturbative formulae for the radii of massive particle spheres and massive shadow radii, where an extremal black hole is employed as the perturbative background. Adopting extremal black holes as the perturbative background enables energy-dependent scans to probe stronger gravitational regimes. Using these formulae, we perform perturbative analyses of the massive shadows for several non-standard black holes, and obtain the expansion expressions near the photon sphere of the extremal black hole. It is found that, although the dependence on the energy parameter $ε$ is the same across models, the coefficients associated with the deviation parameter $δ$ differ significantly, which can serve to distinguish among different black hole models. Furthermore, we construct a two-point Padé approximant in the form of a continued fraction which achieves a globally accurate approximation for the massive shadow radius of the black hole over the entire parameter range. Numerical tests show that the relative errors based on this approximant are small enough to provide a reliable foundation for model-independent reconstruction of metric parameters.

gr-qc↗

Personalised federated learning for Riemannian and Euclidean EEG decoding

Federated learning (FL) lets EEG decoders learn from recordings of several subjects without pooling them. We consider two light EEG decoders, the Riemannian SPDNet and the Euclidean EEGNet. Both split into a trunk, which builds a latent representation, and a head, which classifies it. Inter-subject variability, however, makes a single shared FL model a poor fit for each subject. Personalised FL addresses this: all subjects learn a common trunk, and each subject keeps its own head. We adapt it for SPDNet and study its effects against standard FL and centralised training, with EEGNet as a Euclidean baseline. Experiments cover three motor-imagery datasets that span diverse regimes in channels, subjects and classes. We observe that personalised SPDNet reaches higher accuracy than both standard FL and centralised training, while converging in fewer rounds and communicating fewer parameters than standard FL. It also outperforms every EEGNet configuration on two of the three datasets, although centralised EEGNet outperforms centralised SPDNet.

stat.ML↗

Spectral eigenvalue problem of Cantor measures and Artin's primitive root conjecture

The eigenvalue problem for a probability measure $μ$ with compact support in $\R$ is whether there exist a countable set $Λ$ and a nonzero real $t\ne 1$ such that both $Λ$ and $tΛ$ are spectra of $μ$, that is, the family $$E_{aΛ}=\{e^{-2πi aλx}:λ\inΛ\}$$ is an orthonormal base for $L^2(μ)$ for $a=1, t$. The eigenvalue problem was discovered independently by Strichartz \cite{Str00}, Łaba and Wang \cite{LW02} for the Cantor measures $μ_{4,\{0,1\}}$ and $μ_{6,\{0,1,2\}}$, respectively. In this paper, we investigate the spectral eigenvalue problem for the general spectral Cantor measure $μ_{b,\mathcal{D}}$. This topic is naturally related to elementary number theory. Unexpectedly, however, our main results depend on the theory of integers, especially Artin's primitive root conjecture. To some extent, our results suggest that Artin's primitive root conjecture may hold and confirms some viewpoints implied by Minkowski in \cite{Min57}.

math.CA↗

Structure-preserving diffuse-domain accelerated saddle dynamics for wetting transitions

Wetting transitions on textured substrates play a central role in the design of functional surfaces, but resolving their transition mechanisms requires efficient exploration of complex energy landscapes. In particular, saddle points are essential for revealing the connectivity among metastable states and transition pathways. In this work, we develop a diffuse-domain accelerated saddle dynamics (DD-ASD) framework for mass-constrained wetting transitions on complex textured substrates. The diffuse-domain formulation embeds complex solid geometries into a Cartesian grid, avoiding body-fitted mesh construction in repeated saddle-point searches. An orthogonal projection is applied consistently to the phase-field state and unstable-direction dynamics, preserving the prescribed droplet mass at the discrete level. Momentum acceleration is incorporated into the saddle dynamics to improve the efficiency of high-index saddle searches. Under the stated local spectral and exact-eigenspace assumptions, we establish a local convergence rate of $1-\mathcal{O}(1/\sqrtκ)$ on the effective mass-conserving subspace. Numerical experiments demonstrate the effectiveness of the proposed method in resolving wetting solution landscapes, transition pathways, and energy barriers on textured substrates. We further extend the framework to fully three-dimensional wetting-landscape computations.

math-ph↗

The Vulnerability of Neural Audio Watermarks under Speech Enhancement

Neural audio watermarks are increasingly deployed in commercial speech generation systems to make AI-generated speech traceable, yet their robustness has been studied mainly under conventional signal distortions. Since a watermark can be regarded as imperceptible noise added to the speech signal, a natural question is whether speech enhancement (SE), as a denoising model, can remove it. In this paper, we cascade Gaussian noise with SE models as a black-box watermark removal attack, covering both discriminative and generative SE paradigms, against six neural watermarks: AudioSeal, WavMark, SilentCipher, Timbre, Perth, and AlignMark. Experimental results show that the proposed attack significantly outperforms existing neural re-synthesis methods in watermark removal. In particular, we find that generative SE, which reconstructs the harmonic regions of speech while denoising, is highly destructive to watermarks. These findings show that SE poses a serious threat to current audio watermarking methods, and we call for SE-aware robustness evaluation in watermark design.

cs.SD↗

From Scattered Gaussians to Structured Maps: Efficient Gaussian Splatting Coding via Dual-phase Morton Sorting

3D Gaussian Splatting (3DGS) enables high fidelity novel view synthesis but suffers from excessive storage and bandwidth requirements due to its unstructured representation. To address this, a projection based video coding framework has emerged as a leading approach, supported by MPEG's ongoing standardization, where 3DGS attributes are converted into 2D maps to take advantage of efficient compression using established video codecs such as HEVC and VVC. However, the effectiveness of this approach depends heavily on the spatial coherence of the projected video, which current sorting strategies such as PLAS and Morton ordering fail to preserve adequately, either incurring high computational cost or achieving limited correlation retention. To overcome these limitations, we propose a dual phase Morton spatial sorting algorithm that improves both coding efficiency and processing speed. In the first phase, Morton based 1D indexing is applied to high dimensional attributes to enhance spatial locality. The second phase further refines layout continuity through a structured 2D Morton mapping table that enforces spatial adjacency. This hierarchical strategy generates highly regular, block wise feature maps with strong local correlation, making them well suited for compression via conventional block based coding tools. Experimental results show that our method significantly outperforms existing approaches in both compression performance and runtime efficiency, providing a practical and standard compatible solution for 3DGS data coding.

cs.MM↗

Quantitative QSD convergence in 1-Wasserstein distance via the Föllmer drift

We develop a novel pathwise approach to study the convergence of the law of killed diffusion processes conditioned on non-absorption, towards a quasi-stationary distribution (QSD) as time goes to infinity. We start from the general observation that the dynamics of an absorbed Markov process conditioned upon survival up to time $T>0$ is the minimizer of the pathwise relative entropy with respect to its unconditioned dynamics, under a simple distributional constraint at that time; in other words, a Föllmer process. We then show how this result applies to a Brownian diffusion process softly-killed at a state-dependent regular rate, and characterize the associated drift change. In the case when the diffusion process is moreover reversible, we leverage this idea and recent results on the propagation of weak log-concavity of HJB semigroups to prove that, under strict asymptotic convexity of the potential, the conditioned dynamics satisfy a contractivity property in $1$-Wasserstein distance, uniformly in $T>0$. Under a general ergodicity condition on the associated Feynman-Kac semigroup, we then establish the existence of a QSD with a large domain of attraction, and the exponentially fast convergence to it of the conditioned semigroup in the $1$-Wasserstein distance as $T$ goes to infinity. Finally, we deduce the exponentially fast convergence, also in $1$-Wasserstein distance, of the law of the corresponding Q-process towards its equilibrium.

math.PR↗

Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots

General Purpose Service Robot (GPSR) tasks, as defined in the RoboCup@Home benchmark, require robots to interpret diverse natural language commands and generate multi-step action sequences in real home environments. Conventional Single Prompt (SP) approaches suffer from context bloat and the "Lost in the Middle" phenomenon, leading to unreliable task planning. We propose an LLM chaining architecture that separates instruction classification and action generation into two specialized stages, reducing per-inference prompt length by approximately 45% while improving planning consistency. We evaluate our method using 100 randomly generated GPSR commands across three language models spanning local open-source and frontier cloud deployment contexts. Results show consistent planning improvements over SP across all models, with gains of up to +37 percentage points on local models. Further, real-robot execution experiments on the Toyota Human Support Robot (HSR) reveal that planning success alone does not guarantee task completion, with 6 of 10 tasks completing successfully and execution-layer failures identified as the primary remaining bottleneck.

cs.RO↗

Multi-Agent Orchestration of 3GPP Channel Estimators

Pilot-aided channel estimation is a decisive block in orthogonal frequency-division multiplexing (OFDM) receivers for both 5G New Radio (5G-NR) and Long-Term Evolution (LTE). A large body of estimators exists, from simple least-squares (LS) interpolation to statistically optimal linear minimum-mean-square-error (LMMSE) variants and, more recently, deep convolutional denoisers, yet no single estimator is uniformly best: the winner depends on the propagation scenario, the numerology, the operating signal-to-noise ratio (SNR), the mobility (Doppler), and the antenna configuration. In this paper, we quantify this fact through a unified study of eight literature estimators evaluated over the 3GPP TR~38.901 Urban-Macro (UMa), Urban-Micro (UMi), and Rural-Macro (RMa) channels generated with NVIDIA Sionna, for both 5G-NR and LTE numerologies, in single-input single-output (SISO) and $8\times2$ multiple-input multiple-output (MIMO) settings. We then propose a \emph{condition-adaptive multi-agent orchestrator} that treats each estimator as an independent agent and dispatches, per operating condition, to the agent that is best on a validation split without any genie knowledge. The orchestrator tracks the per-realization oracle to within $1.07$~dB and improves the normalized mean-square error (NMSE) over the best \emph{fixed} strategy by up to $3.6$~dB at high SNR, where the low-SNR champion is no longer optimal. Because the agents are independent, running them concurrently delivers this best-of-eight accuracy at essentially single-estimator latency: a data-parallel partition scales the wall-clock nearly as $1/K$ with $K$ workers (up to $6.9\times$), whereas naive by-algorithm partitioning is Amdahl-limited by the heaviest agent. The results substantiate multi-agent orchestration as a practical route to robust channel estimation across heterogeneous 5G-NR/LTE deployments.

cs.IT↗