arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Nahm-Kirchhoff Moduli Spaces, Hyperkähler Quotient and 2D Field theories

To each oriented quiver with boundary $Γ$ and compact connected Lie group $G$, we associate a Nahm--Kirchhoff moduli space $\mathcal{M}_{\mathbb{H}}(Γ)$ of solutions to Nahm's equations along the edges, subject to Kirchhoff matching at the interior vertices, modulo gauge transformations trivial on the boundary. For connected $Γ$, we prove that $\mathcal{M}_{\mathbb{H}}(Γ)$ is a smooth manifold of dimension $4(|E|-|Γ_{\mathrm{int}}|)\dim G$ carrying a hyperkähler structure obtained as an infinite-dimensional hyperkähler quotient. Using a Kempf--Ness argument on $G_{\mathbb{C}}/G$, we establish a quiver analogue of Donaldson's theorem, identifying $\mathcal{M}_{\mathbb{H}}(Γ)$ with the complex-symplectic quotient $\mathcal{M}_{\mathbb{C}}(Γ)$ of solutions to the complex Nahm equations modulo the complexified gauge group. The space $\mathcal{M}_{\mathbb{C}}(Γ)$ is independent of the edge lengths and isomorphic to a finite-dimensional complex-symplectic quotient of $(T^*G_{\mathbb{C}})^E$, whereas the hyperkähler metric depends on them. Through cutting and gluing laws, $Γ\mapsto \mathcal{M}_{\mathbb{C}}(Γ)$ defines a 2D TQFT valued in complex Hamiltonian manifolds, in the spirit of Moore--Tachikawa, while its hyperkähler refinement fails to be functorial, since subdividing an edge into sub-edges of arbitrary lengths alters the metric. Retaining the edge lengths restores functoriality: the spaces $\mathcal{M}_{\mathbb{H}}(Γ)$ assemble into a metric field theory on metric cobordisms, valued in hyperkähler manifolds composed by hyperkähler reduction, which fibres over the tropical moduli space $M_{g,n}^{\mathrm{trop}}$ with fixed complex-symplectic fibre. We also describe the degenerations of the metric as an internal edge collapses or becomes infinitely long.

math.DG↗

Active Phase and Amplitude Modulation in Perovskite Polaritonic Waveguides

Integrated photonic technologies require active materials capable of providing large optical modulation within compact footprints and low losses. The large optical nonlinearities of exciton-polaritons offer a promising route to meet these requirements, yet their integration into photonic circuits remains challenging because strong light-matter interactions must be combined with low-loss transport and spatially localized control. Here, we demonstrate active phase and amplitude modulation in halide perovskite polaritonic waveguides integrated with electrical microheaters. Local electrothermal actuation drives a structural phase transition in the perovskite, modifying the complex refractive index dispersion near the excitonic resonance. By tuning the excitonic fraction of the propagating mode, the waveguides can therefore operate either as a phase shifter with a tuning efficiency of 0.2 mW mm, approximately one order of magnitude lower than silicon modulators with similar geometries, or as an intensity switch , providing 1 dB/um modulation depth at 20 mW. Our work establishes phase-change excitonic reconfiguration as a route towards compact polariton circuits for nonlinear photonic applications.

physics.optics↗

Mathematical study of a new Navier-Stokes-alpha model with nonlinear filter equation - Regularity theory and refined time asymptotics

This article is devoted to the mathematical analysis of a new Navier--Stokes-$α$ model involving a nonlinear filter equation. The resulting model is governed by a doubly nonlinear parabolic--elliptic coupled system. In our previous work [M.~F.~Cortez and O.~Jarrín, \emph{Mathematical study of a new Navier--Stokes-alpha model with nonlinear filter equation -- Part I}, J. Math. Fluid Mech. 28, no.~12 (2026)], we established the global well-posedness of weak Leray-type solutions and the existence of a global attractor. In the present work, under natural assumptions on $A(\cdot)$, we investigate several additional properties of this model, with particular emphasis on higher-order regularity and its consequences for the long-time dynamics. The regularity analysis constitutes a central and delicate issue, since the nonlinear structure of the elliptic filter and its coupling with the evolution equation preclude a direct application of the standard regularity arguments available for linearly filtered Navier--Stokes-$α$ models. Overcoming this difficulty requires new higher-order estimates for the nonlinear elliptic filter equation, which may also be of independent interest. Among the results obtained, two constitute the main contributions of the article. First, we establish uniform higher-order regularity for the global attractor. This result is then used as a fundamental ingredient in proving an exact determining-modes property for complete trajectories on the attractor: if the projections of two complete trajectories onto a sufficiently large finite-dimensional space of Stokes modes coincide at every time, then the two trajectories coincide identically.

math.AP↗

Benchmarking Candidate Coverage in Typed Decision Models

Typed decision models return choices or distributions over answer options supplied at request time. Accuracy with complete options does not establish whether a model recognizes that a reference answer is missing or avoids rejecting valid candidates. We present a paired candidate-coverage benchmark protocol and an initial evaluation of Laya and Jev across AG News, DBpedia, Emotion, and TREC. The models receive identical frozen texts and requests: 300 calibration and 589 test texts yield 23,932 predictions per model. Present/absent pairs match ordinary candidate count, and name variants preserve descriptions, members, and order. Native rejection behavior differs sharply: at five TREC candidates with natural names, Laya detects 97.2% of missing-answer cases but falsely rejects 69.7% of present controls; Jev's rates are 24.8% and 0.0%. Calibration-only none-score thresholds change these rates to 33.9%/3.7% and 45.0%/1.8%, respectively. On DBpedia, Jev's high coverage-score AUROC supports a stronger operating point, whereas both models have weak complete-set accuracy on Emotion. Competence-conditioned analysis, probability-precision sensitivity, and interface audits show why classification, score ranking, and rejection policies need separate measurement. This initial benchmark is descriptive and limited to reference-label omission; it does not establish natural out-of-scope generalization, causal mechanisms, or a new rejection method.

cs.AI↗

KungfuAthleteBot: learning high-dynamic humanoid motion from video with unified robust recovery

Video is an abundant, inexpensive source of human motion data that is rich in extreme athletic behaviors. Making it usable for humanoid robots, however, is not a matter of simply retargeting a reconstructed trajectory: video-derived motion is physically inconsistent, devoid of actuation information, and says nothing about failure or recovery. We present KungfuAthleteBot (KAB), a framework that treats learning high-dynamic motion from video as the central problem and resolves each of these three failure modes in turn. (C1) We build the KungfuAthlete dataset from videos of national-level martial artists and introduce a physics-guided parabolic trajectory correction that removes height floating, ground penetration, and high-frequency jitter from reconstructed aerial and landing phases. (C2) Because video carries no force information, strict tracking of a reconstructed trajectory is dynamically infeasible, and error-driven initialization keeps re-launching the policy from infeasible aerial poses. We introduce physics-driven pseudo-low-kinetic-energy (LKE) sampling, our central mechanism for making such references learnable: it biases initialization towards dynamically feasible states, letting the policy discover feasible actuation patterns instead of imitating infeasible ones. (C3) Finally, we introduce a direct training paradigm in which disturbance rejection and fall recovery are learned inside the same policy that tracks the video motion, requiring no recovery reference data and no manual mode switching. On a humanoid robot, KAB learns dynamic skills from video and recovers from arbitrary falls in about 0.7 s, the fastest reported recovery for a unified policy. Ablations on the unified policy confirm the necessity of its components, supporting the view that repairing and compensating video data, rather than only collecting more of it, is what unlocks high-dynamic humanoid skills.

cs.RO↗

From Patching to Pruning Visual Computation in Vision Language Models

Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P performs validation-guided forward and backward layer sweeps to identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance. Unlike conventional token-pruning methods, P2P preserves the sequence length, token order, positional information, attention mask, and residual pathways, thereby pruning computation without removing tokens or modifying the pretrained model weights. We evaluate P2P on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks using mutually disjoint calibration, validation, and test partitions. P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%. Beyond these efficiency gains, our layer-wise analysis suggests that visual processing in VLMs is non-uniformly distributed across decoder depth: early and late layers often require little token-specific visual computation, whereas intermediate layers appear to perform most task-relevant visual integration, enabling later reasoning to rely largely on visual information already embedded in shared residual and textual representations. This makes P2P both an efficient inference framework and a causal lens into visual information processing in VLMs.

cs.CV↗

DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift

Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git

cs.SD↗

Native Action-Prior Learning from Videos for World Action Models

World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual representations that must later be adapted for control, or to infer latent actions that are subsequently grounded to robot commands. We present NAVA-WAM, which introduces native action-prior learning by directly pretraining the action policy from observation-only videos, avoiding indirect representation-to-control transfer or a separate latent-action model. Our training consists of two stages. First, we pretrain on observation-only videos, where future-video flow-matching supervision over visual transitions is propagated through transition-structured joint attention to optimize the Action-DiT and learn action-relevant priors. Second, we use action-labeled demonstrations to post-train the Action-DiT for robot control through joint video--action flow matching, while asymmetric attention decouples the visual branch from iterative action denoising and enables efficient action-only inference. Extensive experiments show that NAVA-WAM consistently outperforms prior approaches under both in-distribution and out-of-distribution settings, while demonstrating strong action-label efficiency and effective real-robot generalization. These results establish native action-prior learning as an effective approach to directly pretrain action policies from observation-only videos, providing a scalable path beyond action-labeled robot data.

cs.CV↗

The Plan Language of a Curriculum: A Formal Model and the Complexity of Degree Planning

We model an academic curriculum as a generator of a language of feasible study plans: prerequisites are monotone Boolean formulas in conjunctive normal form, degree requirements are credit-threshold covering constraints, and a study plan is a sequence of terms bounded by a per-term credit capacity. Within this model, we settle the complexity of the two natural planning objectives, the number of terms to a degree and the total credit load, and we isolate the structural commitment responsible for each source of hardness. Time to degree is polynomial whenever the per-term capacity is unbounded, for arbitrary disjunctive prerequisites and arbitrary electives, so disjunction never contributes to its hardness, yet it becomes strongly NP-hard as soon as capacity binds, even without any prerequisite. Load is complementary: disjunction and overlapping electives are each strongly NP-hard in isolation and their complexity does not depend on capacity, while load is polynomial on the conjunctive, mandatory fragment. The two objectives therefore have disjoint sources of hardness. We show that the delay-factor component of the standard curricular-complexity metric is a polynomially computable upper bound on time to degree, exact on the conjunctive fragment and loose elsewhere by a quantity we name the disjunctive slack, and we prove that program subsumption is coNP-complete and consensus prerequisite recovery is NP-complete. Instantiating the model on a corpus of twenty-two universities, we find that 88 percent of prerequisite-bearing courses are purely conjunctive and that capacity, not prerequisite logic, is the operative constraint on time to degree. The curriculum corpus is openly available (https://doi.org/10.5281/zenodo.22334674), and the analysis and figure-generation code accompany the paper.

cs.DS↗

Multivariate quantum signal processing with optimal query complexity

Polynomial transformations of data encoded in quantum operations are a building block of quantum algorithms, and their query complexity measures how often those operations are used. For multiple variables, existing methods either realize restricted polynomial families or implement monomials separately at a query cost exponential in the number of variables. We present a quantum circuit that implements every multivariate trigonometric polynomial up to an explicit normalization. The circuit processes one variable coherently over the frequency components of the remaining variables, allowing all components to share the same signal queries. Each variable is queried exactly as many times as its degree in the polynomial, and we prove that this query complexity is optimal. We further extend the construction to commuting unitaries. The resulting circuit applies the multivariate polynomial simultaneously across their shared eigenbasis, using the minimum numbers of forward and inverse queries to each unitary. We also study trainable rotations with fixed preparation and a local readout of the last input. At query depth at least two, independent uniform initialization gives every active readout gradient a variance lower bound set by its squared branch weight and the inverse square root of the depth, uniformly over data states and dimensions. For supervised squared loss, retaining sample correlations gives an initialization gradient-variance bound in terms of targets and input spectra. This also bounds the expected loss decrease after the first exact gradient step with a specified step size. These results provide a scalable and practical solution for developing multivariate quantum algorithms and quantum learning models.

quant-ph↗

Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so the full distance to the target must be covered from that side. We propose the Bidirectional Voronoi-biased Exploration curriculum for Reinforcement learning (BVER), which expands from both ends at once. Inspired by bidirectional RRT planning, BVER grows start states outward from the goal and goals outward from the initial state distribution, biases both toward unexplored task space, and steers them toward each other, training one goal-conditioned policy on both. On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than all compared reference-free curricula. On box climbing, it reaches 95% success on a 0.4 m box in roughly 65% fewer iterations than the best of them, is the only one of them to learn to climb a 0.7 m box, and yields a policy robust to start, goal, and yaw variation. Without a demonstration, it approaches the sample efficiency of reference-based curricula on the 0.4 m box and on ring-on-peg transfer. Ablations show that expanding from both ends outperforms either direction alone.

cs.LG↗

A Unified Framework for Bayesian Data Assimilation with Generative Models and Observation Interpolants

Bayesian data assimilation combines model forecasts with noisy observations, but sampling high-dimensional, non-Gaussian posteriors remains challenging. We introduce an observation-interpolant framework that turns pretrained stochastic interpolant, flow matching, and diffusion models into posterior samplers without retraining. Conditioning the interpolant path on observations yields a shared likelihood-score correction to the drift or velocity, unifying stochastic and deterministic posterior sampling. The resulting SDEs and ODEs sample the exact posterior when the intermediate likelihood score is known. For practical computation, we approximate this score using a closed-form Gaussian surrogate with a bias-corrected mean and covariance inflated by the model's source covariance. Jacobian-free and ensemble-shared approximations make the method tractable in high dimensions. We evaluate the framework on linear-Gaussian dynamics, stochastic two-dimensional Navier-Stokes, and urban airflow with up to $O(10^4)$ degrees of freedom.

cs.LG↗

Anisotropic Domain Walls and Cosmologies

We investigate anisotropic domain wall and cosmological solutions in $D$-dimensional gravitational theories of arbitrary spacetime signature coupled to scalar fields with a scalar potential. We identify a class of scalar potentials for which the second-order equations of motion admit a first-order formulation. The associated bosonic flows allow non-trivial Ricci-flat worldvolumes, while their Killing-spinor realization further requires the worldvolume to admit a non-vanishing parallel spinor. Five-dimensional ${\cal N}=2$ gauged supergravity provides a concrete realization of this construction, with the worldvolume geometry determined by its signature and giving rise to pp-wave, hyper-Kähler, and hypersymplectic geometries. We further construct a general class of anisotropic solutions depending on a single variable, with the anisotropy encoded in a constant traceless matrix. A particular branch arises when a subset of scalar fields contribute through their kinetic terms but do not enter the scalar potential; these fields follow geodesics on the scalar manifold. As an explicit application, we recover the Kasner-dilaton-AdS solutions and show that, in the appropriate limit, they satisfy the first-order flow equations.

hep-th↗

Harmonic Eigenspace: A Web-based Application for Navigating and Composing Microtonal Harmony

This paper presents a web-based application for navigating and composing microtonal harmony, built on the Harmonic Eigenspace, a four-dimensional psychoacoustically grounded space in which tetrad chord types are located by their spectral dissonance profiles, computed with Sethares's roughness/dissonance model. The coordinate system is transposition-invariant: a coordinate triple (α, \b{eta}, γ) locates the three upper notes in relation to the root, so a chord quality corresponds to a direction in the space, the invariant ray along which transposition acts, while the root frequency sets the scale. The dissonance field over these coordinates can be computed at any register; the locations of its local minima are register-invariant, as they arise from partial-coincidence ratio conditions. The dissonance volume contains 100 local minima that align with just-intonation intervals and act as landmarks, organising the space into basins around the most consonant tetrads. We embed tetrads from three tonal equal temperaments as discrete lattices within this continuous volume. The application presents this space through two components: the Harmonic Eigenspace as a navigable 3D visualisation of the dissonance volume in which all nodes are playable, and a Modal Studio that extends modal interchange logic to the ten-gradation interval vocabulary of 53-TET. The application also functions as a MIDI controller with MIDI Polyphonic Expression support, usable in any digital audio workstation that supports this format. A listening study with 31 participants used both scenes of the application: listeners first rated isolated 53-TET chords alongside chords familiar from Western practice, such as the maj7 and the m7; they then rated chord progressions composed in the Modal Studio, measuring their acceptance or rejection of microtonal progressions heard for the first time.

cs.SD↗

Optimizing quantum error correction through error attribution

Logical error rate is a standard benchmark for quantum error correction, providing an aggregate measure of the overall effect of physical noise. In actual QEC architectures, errors at different circuit locations can have markedly different effects on logical failure even when their probabilities are equal. Here we introduce an error attribution scheme to quantitatively characterize the influence of individual circuit components on the logical error rate and use this refined information to guide targeted intervention and design for improved QEC performance. Our attribution estimator computes all component sensitivities jointly from the same Monte Carlo samples used to estimate the logical error rate. The resulting sensitivity maps identify hotspots where targeted noise reduction can most effectively improve logical performance. We test the scheme in circuit simulations of rotated surface code memories and high-rate lifted product codes: halving the noise strength on about 5\%-7\% of physical components selected by sensitivity reduces logical error rates by approximately 15\%-25\%, roughly twice the mean reduction from the same intervention on an equal number of randomly selected components. By connecting performance benchmarking to targeted optimization, our scheme provides a practical strategy for improving quantum error correction with limited resources.

quant-ph↗

Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to induce lengthy reasoning, as the model becomes overly reliant on high-entropy visual activations. Alternative approaches that rely on counterfactual-based visual token localization for credit assignment also tend to over-prioritize visual exploration at the expense of decisive reasoning cues for answer derivation, thereby exacerbating the interference from spurious visual nuances. Moreover, these methods employ static counterfactual strategies that fail to co-evolve with the policy during training. In this paper, we introduce DyCPO, a co-evolutionary framework that jointly optimizes reliable token selection and adaptive counterfactual intervention. It constructs a multi-role dependence metric to balance visual exploration and answer-relevance mining in token-wise contrastive learning, while suppressing exploration-only filler tokens and spurious visual noise. Rather than relying on static counterfactual priors, DyCPO dynamically derives counterfactual signals from the model's own successful and failed rollouts, enabling self-diagnostic analysis and co-evolution of the optimization objective with the policy. Extensive experiments on complex video reasoning and general video understanding benchmarks demonstrate consistent performance improvements, establishing DyCPO as a robust token-level credit assignment paradigm for multimodal reinforcement learning.

cs.CV↗

Eisenstein congruences in tame families and the class group of a metabelian extension

Ribet's 1976 proof of the converse to Herbrand's theorem created a new paradigm in algebraic number theory by illustrating that unramified abelian Galois extensions of number fields can be constructed using Galois representations associated to Eisenstein congruences, i.e., cuspidal eigenforms that are congruent to Eisenstein series. In recent years, further progress has been made in investigating the quantity of such Eisenstein congruences, showing that when there are many such congruences, additional and finer information can be gleaned about the arithmetic structure of the related extensions. In this vein, we use new results on the existence of Eisenstein congruences in tame families to explain a computational observation about splitting behavior of primes in a certain metabelian number field.

math.NT↗

16-bit Precision of Convolutional Neural Networks on Microcontroller Units for 8-bit Costs

To deploy deep neural networks on edge hardware, highly efficient inference schemes are necessary that retain high accuracy. This work presents W16A16, a high precision (16-bit), fast speed, low energy quantization method. On a widely applied microcontroller architecture Armv7E-M, our proposed approach achieves faster speed and lower energy consumption on layer- and model-level compared to alternative quantization schemes. We analyze the architecture of Armv7E-M, explain the underlying principles behind the performance advantages of 16-bit approaches, and evaluate the empiric quantization errors for regression and classification tasks, as well as empiric time- and energy consumption in MCU deployment. We observe ca.\ 10 times lower quantization errors compared to 8-bit quantization schemes while achieving similar or better inference times and energy consumption.

cs.LG↗