arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair

Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether verifier-admitted outputs can supervise a bounded model by consolidating occasional or expensive capability into reliable one-shot execution. On fresh, structure-disjoint PlanBench replanning cases, Qwen3-14B thinking produced 24 plans admitted by independently authored VAL. They trained the same checkpoint for non-thinking execution, without oracle targets or a stronger teacher. VAL acceptance rose from 1/80 to 57/80. A prospective replication held targets, model revision, recipe, and evaluation corpus fixed across eight LoRA seeds and three inference realizations per seed. Every seed produced a clear lift: adapters reached 45/80 to 70/80 against 1/80 for every matched base report; the exact seed-level sign-flip test gave p=0.0078125. Target-selection performance was less stable. An initial matched seed gave 102/160 accepted plans after VAL selection versus 69/160 after blinded model self-selection. Across eight prospective seeds, the contrast was seed-dependent, included one clear reverse seed, and did not replicate (p=0.3672). VAL also had a positive descriptive aggregate over schema-only selection but failed its preregistered seed-level reliability gate (p=0.0703). The verifier remains the admission authority; no reliable downstream capability advantage of semantic selection is established. A complementary Phi arm supports stronger-teacher distillation. Earlier synthetic studies bound teachability, cumulative learning, transfer, and stopping. The evidence supports robust consolidation of one fixed, machine-checkable capability, not arbitrary planning, enterprise validity, or unrestricted self-improvement.

cs.AI↗

J0011+3443: a GPS compact symmetric object, gravitational lens, or dual AGN?

We present new multi-frequency VLBA observations of J0011+3443 (TXS 0008+344, z = 0.89) at 2.3, 4.9, 8.5, and, for the first time, 23.6 GHz. The source consists of two compact components A and B at a projected separation of 314 +/- 2 pc, plus a third feature C detected at 23.6 GHz at 0.6 mas from A. Archival low-resolution radio measurements confirm an integrated gigahertz-peaked spectrum, with an observed-frame peak frequency of 0.73 +/- 0.08 GHz and a peak flux density of 879 +/- 125 mJy. Comparison with nearly frequency-matched low-resolution measurements shows that the VLBA recovers 0.85 +/- 0.11 of the 4.85 GHz flux density and 0.50 +/- 0.06 of the 8.46 GHz flux density. The lower recovered fraction at 8.5 GHz suggests that low-surface-brightness emission is resolved out or falls below the VLBA surface-brightness sensitivity. We therefore interpret the VLBA component spectra as spectra of the compact recovered emission only. The 23.6 GHz morphology, the absence of a detected flat-spectrum core, the similar compact spectra of A and B, and the steep integrated GHz spectrum favor an interpretation of J0011+3443 as a GPS-class compact symmetric object, possibly in a short-lived or relic phase, although a dual-AGN origin cannot be excluded without multi-epoch astrometry.

astro-ph.HE↗

DecoVAE: a Lightweight Interpretable Trend-Seasonal VAE Framework for Efficient Probabilistic Time Series Forecasting

Probabilistic time series forecasting remains challenging, largely because modeling distinct trend and seasonal dynamics requires specialized approaches. Existing methods often fail to capture the unique inner properties of these components, lack interpretability, or suffer from heavy memory and runtime overhead. To address these limitations, we propose DecoVAE, a lightweight interpretable trend-seasonal VAE framework that explicitly decomposes time series into trend and seasonal components by applying domain-specific inductive biases. The trend stream enforces structural smoothness using a differential regularizer on the latent trajectory, analogous to the Hodrick-Prescott filter. Concurrently, the seasonal stream operates in the frequency domain via a complex Gaussian VAE, natively capturing the amplitude and phase of periodic patterns. Extensive evaluations across seven real-world benchmarks show that DecoVAE consistently outperforms strong baselines. It achieves reductions of up to 14.96\% in CRPS and 23.30\% in NMAE for short-term forecasting, and up to 52.68\% and 26.51\% for long-term horizons. Crucially, DecoVAE yields these accuracy gains while remaining highly efficient, reducing model weight by up to 93\% and accelerating speed by up to 74\% compared to the second-best method.

cs.LG↗

Algorithms, Complexity, and Entropy of the Bernard-Letac Fair-Sampling Construction

Bernard and Letac (1971) introduced a method for uniform random sampling among $m$ outcomes from an unknown biased source of independent and identically distributed symbols. The process terminates when the multinomial coefficient of the cumulative symbol counts equals zero modulo $m$. This study extends the computational and information-theoretic analysis of their construction by presenting five algorithms with formal correctness guarantees and comprehensive complexity analyses. For prime $m=p$, the Bernard-Letac framework is analyzed in greater detail. The Rényi entropies of the source yield an exact product formula for the expected number of draws. A first-order approximation consistently overestimates this value, and the entropy lower bound is never attained. As $p$, treated as a continuous parameter, approaches 1, the expected cost converges to a constant greater than 1, determined by the entire source distribution. Furthermore, for every prime modulus $p$ and every finite alphabet $I$, an explicit automaton with $p + |I| + 3$ states computes the mod-$p$ first-passage kernel of the walk from the base-$p$ digits of its arguments, reducing the fair assignment cost from quadratic to nearly linear.

cs.IT↗

The Communication Map of a Transformer

The components of a transformer communicate by writing to and reading from a shared residual stream, and the mechanistic interpretability literature has mapped these connections by hand, one circuit at a time. We present the communication map, which charts every potential communication channel from the geometry of the model's weights alone, generalizing the composition score of Elhage et al. (2021) into a single coupling coefficient covering all 18 connection classes, from head-to-head to neuron-to-neuron and everything in between. We provide an account of the properties of the coupling coefficient, including its geometric interpretation and its exact chance level. The census finds that 70-89% of head pairs are oriented far from chance, some coupled strongly and others actively avoiding each other. We demonstrate the communication map in two novel applications. In Application 1, we recover the known induction circuits blind from the strongest head-to-head couplings and group the heads into communities, and ablating one such community destroys the model's in-context copying. In Application 2, we pool the coupling coefficients of every head to identify a distinct two-dimensional residual stream subspace, whose deletion abolishes the induction capability in six models up to Pythia-6.9B. We show that this subspace is different from those identified by either activation PCA or outlier dimensions. We release the map, the statistical machinery, and the intervention suite.

cs.LG↗

On the formation of subdwarf B stars via engulfment of substellar companions I. Conditions for and occurrence of envelope ejection events

The canonical scenario for the formation of subdwarf B (sdB) stars involves the ejection of the progenitor's envelope near the tip of the Red Giant Branch (RGB), concurrent with the onset of core He-burning. While binary interactions are known to dominate sdB formation, the origin of apparently single sdB stars remains uncertain. We investigate the conditions under which an sdB progenitor can eject its envelope through the engulfment of a substellar companion, using an energy balance approach. We simulate the orbital evolution of substellar companions during the subgiant and RGB phases of their host stars (1.2-2.0 M_Sun) to determine the onset of engulfment and calculate the energy released during inspiral within the stellar envelope. By comparing this energy with the envelope binding energy and exploring different efficiencies for its deposition, we identify the conditions required for envelope ejection. We then estimate the occurrence of such events using observed population of substellar companions. The engulfment of substellar objects can lead to envelope ejection within a limited region of the mass-semi-major axis parameter space, whose extent depends on the host star's properties and energy transfer efficiency. The minimum companion mass required for ejection increases with decreasing efficiency and increasing stellar mass. As a result, the frequency of envelope ejection is strongly influenced by the distribution of substellar companions, particularly by the paucity of objects in the brown-dwarf desert. The engulfment of low-mass brown dwarfs and massive planets during the late RGB can provide sufficient energy to eject the stellar envelope, which could ultimately lead to the formation of sdB stars. Our results define the region of parameter space where this mechanism is energetically possible and provide a guide for future multidimensional hydrodynamical simulations.

astro-ph.SR↗

PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors

Recent progress in the field of TinyML has demonstrated that low-power hardware based on microcontrollers can achieve bird species monitoring in real time based on acoustic sensor data for an entire breeding period on a single battery charge. However, the state of the art on low-power microcontrollers was so far limited to binary classification of a single species. In contrast, real fauna monitoring deployments often target multiple species simultaneously. To address this challenge we develop PolyChirp, an approach combining biological domain expertise, automated dataset curation, neural architecture optimization and novel hardware to achieve multiclass bird species detection in the wild. PolyChirp is based on newly designed tiny multiclass models that leverage recent microcontrollers and hardware acceleration with a neural processing unit (NPU). We evaluate the predictive performance of these models, and we measure their computational performance -- flash footprint, latency, energy consumption -- on common microcontroller hardware. Our results demonstrate that PolyChirp matches or exceeds the TinyChirp architectures retrained under our protocol on single-species detection, and further achieves robust classification of up to 10 species simultaneously (macro F2 up to 0.97), while still fitting the flash, latency and energy budget of a low-power microcontroller sensor. A data-driven front-end redesign additionally makes on-device mel feature extraction 7x to 11x cheaper.

cs.LG↗

The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge: A Benchmark with Natural Mixtures and Degraded Video

Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge addresses both gaps. Track~1 combines naturally recorded two-talker mixtures, which lack a clean reference, with reference-available remixes of the same speakers; Track~2 additionally degrades the target video in five ways and adds 3-m far-field recordings. Sixteen and twelve teams were ranked on speaker-disjoint test data by rank averaging over waveform fidelity, predicted quality, transcription accuracy, and speaker similarity. On Track~1 remixes, the best system reaches 12.7~dB SI-SDR and 0.85 STOI, but natural recordings remain harder: even the lowest CER rises from 9.0\% to 14.7\%. The leading systems use video mainly for speaker attribution rather than signal reconstruction, and the top two lose under 0.5~dB SI-SDR on Track~2. UTMOS and DNSMOS rank systems differently from the other metrics, so no single metric captures target-speech recovery. We release the baselines, evaluator, and official results.

eess.AS↗

Provable Quantum-Classical Separation for Continuous Gibbs Sampling

We prove the first quantum-classical separation for a sampling problem over a continuous domain. For a class of Gibbs states $p\propto e^{-βE}$ on the torus $\mathbb{T}^d$ with smooth ($s$-Gevrey) potential and barrier amplitude $α=e^{βΔ}$, where $Δ= \max E-\min E$, every classical algorithm querying the value, gradient, or any higher-order derivatives of the log-density requires $Ω(α)$ queries to sample at constant accuracy in total variation distance, while a quantum algorithm based on quantum singular value thresholding and temperature annealing samples with $\tilde{O}\left(\sqrtα\right)$ queries to an oracle for the gradient. The advantage is quadratic in the barrier amplitude, which becomes exponential in the dimension, $e^{Ω(d)}$, at low temperature. The classical bound is information-theoretic, holding for every classical algorithm with query access to the Gibbs potential and its derivatives at any order.

quant-ph↗

PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their ability to solve complex physics problems. To address these issues, we present PhysElite, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning. PhysElite contains 11,586 Olympiad-tier problems. For each problem, we provide corresponding visual diagrams, step-by-step bilingual Chinese-English solution derivations, and the final answer. We benchmark 18 open-source and closed-source MLLMs, and find that even the strongest model reaches only 33.7% answer accuracy. We additionally conduct step-level process evaluation to diagnose where models fail in the reasoning chain. Our datasets are released at https://huggingface.co/datasets/physelite/PhysElite.

cs.AI↗

LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents

Modern GUI-agent frameworks achieve strong desktop task performance with frontier API models, yet persistent control information often remains implicit in growing interaction trajectories. At each step, the planner reconstructs the active task stage, accumulated evidence, and runtime feedback before deciding the next action. This dependence becomes more pronounced under weaker local reasoning backbones. Across four representative state-of-the-art frameworks, replacing GPT-5 with Qwen3.5-9B reduces average OSWorld SR-100 from 60.9\% to 35.2\%. Trajectory annotation further identifies at least one control failure in 94.7\% of failed trajectories. To address this problem, we introduce LocalLSTC, a training-free architecture that organizes control by temporal scope, maintaining persistent cross-step state to guide short-term execution commitments. The framework comprises two complementary mechanisms. Long-to-Short Planning forms each commitment from persistent state, while Short-to-Long Control integrates execution outcomes back into that state for progress assessment, recovery, and termination. With Qwen3.8-27B, LocalLSTC reaches 73.7\% SR-100 on OSWorld and 68.5\% on WindowsAgentArena, achieving the strongest local OSWorld result and a new state of the art on WindowsAgentArena. Ablations further support contributions from both mechanisms. Together, these results show that temporal organization of control information can reduce cross-step control failures and complement backbone scaling in locally deployed GUI agents.

cs.AI↗

SimCast-S2S: A Computationally Efficient Diffusion Model for Subseasonal Precipitation Forecasting

Subseasonal-to-seasonal (S2S) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains simulation fidelity. We introduce SimCast-S2S, a generative latent-diffusion framework for probabilistic S2S precipitation forecasting that addresses three major bottlenecks in data-driven prediction. First, because S2S prediction requires uncertainty quantification rather than only deterministic point forecasts, SimCast-S2S is the first data-driven system that uses a diffusion-based generative pipeline for S2S prediction, enabling effective sampling from the underlying conditional distribution. Second, since generating large probabilistic ensembles is computationally costly in physical space, SimCast-S2S instead operates in a compact latent space learned by variational autoencoders (VAEs), enabling efficient large-ensemble generation. Third, diffusion models typically require large training datasets; SimCast-S2S overcomes this via transfer learning with low-rank adaptation (LoRA), pretraining on large ensembles of climate simulations before fine-tuning on limited reanalysis data. On reanalysis data, SimCast-S2S outperforms deep learning baselines, including convolutional neural networks and U-Net architectures. Notably, despite using only a subset of atmospheric input variables and no post-processing, bias correction, or calibration, SimCast-S2S remains competitive with, and in many aspects outperforms, state-of-the-art operational systems such as the ECMWF-S2S baseline. These results indicate that latent generative modeling combined with simulation-to-reanalysis transfer learning offers an efficient and scalable path toward data-driven probabilistic S2S precipitation forecasting.

cs.LG↗

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.

cs.AI↗

On Eigenvalue Bounds for Bounded Genus Graphs and Minor-Free Graphs

In this paper, we resolve a 30-year-old conjecture of Spielman and Teng concerning the performance of the spectral partitioning method on graphs embeddable on an orientable surface of genus $g\ge 1$. In particular, for such a graph $G$ with $n$ vertices and maximum degree $Δ$, we show that the second-smallest eigenvalue of its Laplacian matrix satisfies $λ_2(L_G)\lesssimΔ\frac g n$. We also obtain an improved eigenvalue bound for $K_h$-minor-free graphs of $λ_2(L_G)\lesssimΔ\frac{h^2(\log h)^2}n$. In fact, our results directly prove much stronger results for reweighted eigenvalues, including higher reweighted eigenvalues. As a consequence, we obtain bounds not just on Laplacian eigenvalues, but also on normalized Laplacian eigenvalues and Steklov eigenvalues. Our results for genus-$g$ graphs are optimal for all of these kinds of eigenvalues, while our results for $K_h$-minor-free graphs are optimal up to $\log(h)$ factors. Our techniques for genus-$g$ graphs bootstrap bounded-degree bounds of normalized eigenvalues for entire classes to bounds for reweighted eigenvalues for the same classes without the bounded-degree limitation, while our techniques for $K_h$-minor-free graphs generalize an argument of Korhonen and Lokshtanov, making use of the Lovász local lemma.

math.CO↗

Tensegrity Continuum Robots Enable Task-Adaptive Morphologies for Cooperative Behaviors

Robots that can change their morphologies and behaviors for different tasks and environments hold great promise for adaptable, multifunctional systems. Modular reconfigurable robots (MRRs) can achieve such functionalities by docking and rearranging individual units, but most rely on rigid modules that lack structural compliance, resulting in limited capabilities. Continuum robots offer compliance through flexible backbones, yet they cannot self-reconfigure into task-adaptive multi-robot configurations. Here, we introduce an MRR that unifies the advantages of both architectures by combining a tensegrity-based compliant body with claw-based connection mechanisms. Each robot can manipulate and locomote independently, and multiple robots can self-reconfigure into different morphologies (e.g., chains, loops, branches) for cooperative manipulation and locomotion. We demonstrate the robots' capability across diverse tasks and environments, including coordinated object manipulation and transport, multimodal locomotion, and loco-manipulation in real-world scenarios. These results lay a foundation for adaptable and multifunctional robotic collectives, with broad potential applications in manufacturing, space exploration, and search-and-rescue operations.

cs.RO↗

Exploring continuous beta-ensembles: A Python implementation for random matrix spectral statistics

We present an open-source Python package for sampling the Gaussian, Laguerre, Jacobi, and Circular $β$-ensembles of random matrix theory. The package implements the Dumitriu--Edelman and Killip--Nenciu constructions, allowing efficient generation of random spectra for general $β> 0$. In addition to spectrum generation, it includes tools for the analysis of spectral statistics, from standard nearest-neighbor spacings and spacing ratios to non-adjacent $k$-spacings and the spectral form factor. These tools can be applied to generic spectral data, allowing users to compare them with and fit them to $β$-ensemble predictions. In this note, we review the $β$-ensembles, describe the package interface, and illustrate its use through several numerical experiments motivated by applications to quantum chaos. Our numerical results include an analysis of $β$ as a continuous fitting parameter in spacing ratio statistics, an examination of the numerical evidence for the conjectured $k$-spacing ratio distributions, and a study of the spectral form factor for general values of $β$.

nlin.CD↗

A Sharp Spectral Mantel Theorem for Quantum Transport

We study the maximum average quantum transition probability between distinct vertices of a triangle-free graph, equivalently the minimum average return probability, at times inversely proportional to the order. For every compact interval of positive scaled times below $τ_{\mathrm c}\approx3.79099$, the balanced complete bipartite graph is the unique maximizer for all sufficiently large orders. The threshold is sharp: it solves $τ_{\mathrm c}=4\sin(τ_{\mathrm c}/2)$, and the balanced complete bipartite graph is eventually suboptimal at every larger scaled time. The corresponding graphon functional combines the edge density with alternating even cycle densities. We prove that the balanced complete bipartite graphon is its unique maximizer through the critical time. On a nonempty interval immediately afterward, the unique maximizer is a balanced bipartite graphon with constant cross-edge weight below one. After normalization by the squared time, we obtain uniform square-root $L^1$ stability through the critical endpoint and at zero time. We also solve the bipartite problem at every positive time and establish sharp cut distance stability. The proofs combine spectral interpolation, an exactly certified six-vertex moment inequality, vertex-measure variation, and a minimum-degree argument. A uniform finite approximation bound and vertex cloning connect the graphon and exact finite results.

math.CO↗

Efficient Quantum Simulations of Yang-Mills theory with Maximal-tree Gauge

We develop a quantum algorithmic framework for the efficient simulation of Yang--Mills theories, including the $\mathrm{SU}(3)$ gauge theory in Quantum Chromodynamics (QCD). The framework uses maximal-tree gauge in terms of gauge field variables that removes all local gauge redundancies. In the resulting gauge-fixed formulation and digitization in the field-amplitude basis, we show that Hamiltonian time evolution admits an efficient implementation based on quantum singular value transformation (QSVT). We derive upper bounds on the total number of qubits and gate complexity, finding polynomial scaling with the inverse simulation precision $1/\varepsilon_s$, lattice volume $\mathcal{V}$, gauge coupling $g$, and target energy scale $E$. Our results provide a rigorous complexity-theoretic demonstration that non-Abelian Yang--Mills theories can be simulated efficiently on quantum computers, paving the way toward first-principles quantum simulations of non-perturbative QCD dynamics.

hep-lat↗