arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Vacancy-mediated nitrogen diffusion and aggregation via high-fluence electron beam irradiation in HPHT synthesized diamond crystal

The negatively charged nitrogen vacancy (NV-) center in diamond is a promising point defect for highly sensitive quantum sensing. The formation of high-density NV- centers is essential for improving sensitivity. We performed room-temperature electron beam irradiation (EBI) and annealing on nitrogen-doped high-pressure high-temperature diamond crystals, aiming to convert all substitutional nitrogen into NV structures by increasing EBI fluence. While the Ns0 to NV0 and NV- conversion process dominated at low EBI fluences, in a high-fluence region, NV0 and NV- center and Ns0 and Ns+ concentrations decreased with increasing EBI fluence, indicating the formation of unknown nitrogen-related defects such as the H3 center, which is a nitrogen and vacancy aggregation defect. Although H3 centers were observed at high EBI fluence, their annealing temperature of 1375 +- 25 °C was lower than the typically reported temperatures over 1600 °C. We attribute this low-temperature formation to vacancy-mediated nitrogen diffusion and aggregation.

cond-mat.mtrl-sci↗

The Cross-Section of Stock Returns and AI Exposure

We study 380 trillion tokens of realized AI consumption across more than four hundred LLMs. We build a high-frequency AI factor and show that a long-short strategy based on firms' AI exposure earns significantly positive returns. The average strategy return is larger based on intensive, frontier-oriented AI consumption but smaller based on casual or open-weight usage. Internationally, the return spread is significant in developed countries but insignificant in emerging markets. Examining occupational AI exposure, we find more positive exposure in occupations intensive in nonroutine interactive tasks and more negative exposure in those intensive in nonroutine analytical tasks.

cs.CY↗

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain predominantly outcome-centric, relying on final answer correctness or sequence-level preferences. This suffers from sparse credit assignment, making it difficult to optimize the reasoning process essential for clinical applications. Our analysis reveals that cascading errors from early-stage reasoning failures are a leading cause of incorrect predictions in medical visual question answering (VQA) benchmarks. Motivated by this, we propose Medical Reasoning-aware Policy Optimization (MRPO), an RL algorithm that incorporates step-wise process rewards. When the final answer is incorrect, MRPO assigns exponentially larger penalties to tokens in earlier invalid reasoning steps, breaking failure cascades without compromising successful paths. Across four multimodal LLM backbones, MRPO consistently outperforms standard GRPO and a recent RL baseline, and on Qwen3-VL-8B-Thinking even surpasses substantially larger medical MLLMs such as HuatuoGPT-Vision-34B by 4.59 points. Moreover, MRPO reduces early-stage reasoning failures from 58.6% to 13.4%, showing that targeted mitigation of cascading failures improves both reasoning quality and final answer accuracy. Our code is available at https://github.com/dmis-lab/MRPO

cs.CV↗

A multilevel stochastic-gradient neural solver for boundary integral equations

We propose a multilevel stochastic-gradient neural solver (MLSG) for second-kind boundary integral equations. MLSG represents the unknown boundary density using a neural network optimized by stochastic residual minimization over a hierarchy of successively refined Nyström discretizations. Upon transitioning from one level to the next, the network parameters obtained on the previous level initialize training on the current one. This coarse-to-fine strategy retains a continuous, grid-independent density representation and is designed to reduce the total computational effort required to reach a prescribed residual tolerance at the target resolution. The algorithm avoids grid-transfer operators and hierarchical fast-summation machinery, relying instead on batched kernel evaluations and standard network forward and backward passes that map efficiently onto modern GPU architectures. For uniformly stable second-kind discretizations, so strongly nonuniform contraction rates originate in the empirical neural tangent kernel (NTK) rather than in the discretized operator. Within each level, parameter updates can reshape the NTK, while refinement re-samples the tangent kernel on a richer discrete space and reveals directions that were not adequately resolved on coarser levels. A cross-level estimate bounds the warm-start loss in terms of the preceding training tolerance and quadrature error, motivating a tolerance schedule that balances optimization and discretization errors. Experiments on Laplace/Poisson and Helmholtz problems in two and three dimensions, together with an exterior Robin problem on a hypersurface in $\R^4$, demonstrate the method under both parametric and signed-distance surface representations at up to million-scale discretizations.

math.NA↗

BRST-BV approach to fields in Poincare patch of AdS

We use the Poincare parametrization of AdS space to develop a general BRST-BV approach for free fields. A general expression for the BRST-BV Lagrangian of fields with arbitrary masses and symmetry types is obtained. We apply this general framework to study totally symmetric massless, massive, and partially-massless fields with arbitrary integer spin and a continuous-spin field. For these fields, the constrained and unconstrained BRST-BV formulations are developed. We investigate both irreducible and reducible fields. In addition, we demonstrate the matching between the obtained BRST-BV Lagrangian and the metric-like Lagrangian formulated in terms of the modified de Donder derivative. Finally, a realization of AdS space symmetries is obtained within the space of fields and antifields entering the BRST-BV formulation. Modulo the choice of the Poincare parametrization of AdS space, the Lagrangian and the entire approach are manifestly Lorentz invariant.

hep-th↗

Krylov-Lie Algebras for Variational Quantum Algorithms: Geometric, Depth-Aware Insights into Expressivity and Trainability

Variational quantum algorithms (VQAs) are a leading approach to near-term quantum computation, but their utility is limited by barren plateaus and other pathologies in their loss landscapes. Existing landscape theories based on dynamical Lie algebras, Jordan-algebraic Wishart systems, approximate t-designs, and Haar-random circuits are foundational, but they often neglect the finite-depth geometry of realistic ansätze and are therefore ill-suited to the shallow-depth regime, where VQAs are poor approximators of 2-designs and trainability is most feasible. This work introduces Krylov algebras, algebraic structures induced by the Krylov span of a finite generator set acting on one or more seed vectors, as a framework for VQA landscape theory. We show that VQA reachable manifolds can be approximated in a numerically robust, geometrically faithful fashion by Krylov-Lie algebras and groups, and that these structures induce canonical invariant measures for computing expectation values and variances under general sampling measures. In particular, we derive weighted non-Haar variance formulas that recover the usual Lie-algebraic Haar formulas as a special case while isolating non-Haar effects into explicit correction terms. We also show that the common heuristic that sufficiently deep circuit ensembles must converge to Haar fails in general without additional hypotheses, identify concrete obstructions to naive Haar convergence, and recover convergence under natural necessary and sufficient ergodic conditions. Lastly, our formulas further imply that non-Haar contributions to landscape statistics may mitigate barren plateaus by reweighting the visible sectors of the loss landscape, suggesting that VQAs may be more trainable than recent literature has posited.

quant-ph↗

Comparative Evaluation of Encapsulation Methods for Endohedral Doping of Single-Wall Carbon Nanotubes

Single wall carbon nanotubes (SWCNTs) are promising building blocks for nanoelectronic and optoelectronic devices, yet reliable and stable doping, particularly n type, remains challenging due to strong environmental sensitivity and competing extrinsic effects. Encapsulation of charge transfer molecules within the SWCNT cavity offers a promising route to stable doping while preserving the nanotubes outer surface for subsequent processing. Here, we systematically investigate the filling of arc discharge SWCNTs with the electron donor tetrathiafulvalene and electron acceptor tetracyanoquinodimethane, comparing different methods for filling, including melt filling, solution reflux, and vacuum phase sublimation. We follow the entire processing workflow from raw, unfilled powders to aqueous dispersions and employ density gradient ultracentrifugation to separate filled from empty nanotubes as well as metallic from semiconducting ones. Encapsulation efficiency and electronic modification are assessed using absorption spectroscopy, resonant Raman scattering, thermogravimetric analysis, X-ray photoelectron spectroscopy and electron paramagnetic resonance. Finally, we introduce a complementary vacuum-phase method that removes externally adsorbed molecules without extensive solvent washing, enabling cleaner encapsulated systems.

cond-mat.mes-hall↗

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and language development. However, utterance-level timestamps provided in this corpus are often noisy, incomplete, or misaligned with the audio. As a result, utterances cannot always be reliably localized within long recordings, which limits the direct use of these data for training and evaluating speech models. In this work, we propose BEACON (Boundary Estimation via Alignment CONsensus), an ensemble timestamp-curation framework that refines utterance-level timestamps by aggregating knowledge from multiple off-the-shelf ASR models. Specifically, each model's word-level timestamp predictions are first aligned to provided human transcripts, and the final utterance time boundaries are determined by a consensus voting strategy. The framework is corpus-agnostic and applies to any long-form recording paired with a trusted transcript whose timestamps are unreliable or missing, offering a general recipe for timestamp curation. Leveraging this pipeline, we curate and release a 413-hour general-purpose child-speech dataset with corrected utterance-level timestamps, together with a 283-hour quality-controlled subset for ASR training. Fine-tuning on this subset yields up to an average 19.5% relative WER reduction on four out-of-domain child-speech benchmarks.

eess.AS↗

PhysMiner: An Agentic AI Framework for Automated Flow Component Analysis

Uncovering the physical mechanisms of turbulent flows remains a fundamental challenge in fluid mechanics. In particular, conventional velocity-gradient analysis methods suffer from shear contamination, which hinders accurate identification of the dominant physical mechanisms. This study presents PhysMiner, an automated framework integrating the triple decomposition method of the velocity gradient tensor with large language model-driven reasoning for turbulence-physics discovery. The triple decomposition module automatically decomposes flow fields into rigid rotation, pure shearing, and normal straining components, enabling statistical analysis, contour visualization, vortex-line extraction, and threshold-insensitive vortex identification while eliminating shear contamination. These automated capabilities are validated across five benchmarks, ranging from canonical configurations to complex engineering flows. A discover-physics agent combines flow statistics, spatial structures, and literature-derived knowledge to perform pattern recognition and physical inference, while a review Agent iteratively validates physical consistency to ensure reliable conclusions. A continuously evolving Triple Decomposition Library accumulates statistical knowledge from successfully analyzed flows, enabling cross-case comparison and progressive enhancement of inductive capability. The complete PhysMiner pipeline is validated end-to-end on the periodic hill flow, where the framework autonomously generates turbulence modeling recommendations and derives an improved subgrid-scale model with superior Reynolds-stress predictions. PhysMiner is open to the public and establishes a foundation for long-term collaborative advancement in automated turbulence-physics discovery.

physics.flu-dyn↗

SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction

Reconstructing Computer-Aided Design (CAD) modeling sequences from images is crucial for preserving design intent and supporting parametric editing. However, existing methods typically generate full CAD sequences holistically, overlooking the iterative, feedback-driven nature of human design workflows. We address this limitation by introducing the rich stepwise visual supervision: at each modeling step, the system observes the target's orthographic projections, the projections of the incrementally constructed model, and the active sketch, enabling informed action selection. To effectively leverage this on-the-fly feedback, we propose SOV-CAD, a framework that formulates CAD reconstruction as a sequential decision-making task and employs offline reinforcement learning with a Decision Transformer architecture. This design incorporates continuous visual feedback guided by geometric alignment rewards, resulting in a more accurate and human-like modeling process. Extensive experiments show that SOV-CAD surpasses state-of-the-art methods in CAD sequence reconstruction while exhibiting strong data efficiency. Code of SOV-CAD is available at: https://github.com/LukePhong/SOV-CAD

cs.CV↗

Allocation condensation under asymptotically complete mixing

We study a fixed network on which new mass is allocated in proportion to a power of each node's stock and then redistributed. Transport retains a fraction of both old and new mass at each node and divides the rest according to fixed background shares. We show that a small amount of retention can sustain concentration on networks where complete redistribution would leave every node's share of new mass vanishing as the network grows. A node receiving most new mass may retain little relative to the network total, yet much more than redistribution alone would supply. For a power above one, the allocation rule amplifies this relative stock advantage at the next step. We identify retention rates tending to zero with network size that allow widely spread and concentrated stationary allocations to coexist on the same network. One of them stays close to the allocation under complete redistribution. Each node can also dominate a separate concentrated allocation, receiving a share of new mass tending to one. The stocks supporting these allocations attract nearby trajectories, although every stationary stock approaches the same background, with even its largest share tending to zero. These allocations continue to coexist under specified changes in transport that preserve the background. Redistribution can therefore spread the stock widely without spreading new allocations in the same way.

physics.soc-ph↗

Stress calculation in linear scaling DFT: convergence and dynamics

We present the approach needed to calculate stress within density functional theory (DFT) using a localised orbital basis, both for exact diagonalisation and linear scaling approaches, and demonstrate our implementation within the large scale DFT code Conquest. For the linear scaling approach, we test the rate of convergence of stress with density matrix range, and compare it to the convergence of energy and forces for different materials with a range of band gaps. We show that excellent convergence is found for modest cutoffs (below 0.1GPa error for range of 20a0), and show that large-scale isothermal-isobaric molecular dynamics is stable and accurate, showing negligible drift over picoseconds of dynamics.

cond-mat.mtrl-sci↗

On the stability of proximal operators in Wasserstein spaces under different notions of convexity

The proximal operator is a fundamental tool in variational analysis and optimization. In the setting of a Hilbert space, given a proper, lower semicontinuous convex functional, its proximal operator is non-expansive, that is, 1-Lipschitz continuous. In the Wasserstein setting, the contraction properties of this operator have been investigated from different perspectives by Carlen and Craig and by Adve and Mészáros, among others, and are not completely understood. In this paper, we study the stability properties of proximal maps, with a particular focus on non-expansivity, under various notions of convexity of the functional that can be considered in the Wasserstein space.

math.OC↗

glmSTARMA -- An R-Package for fitting autoregressive spatio-temporal models following generalized linear models

The R package glmSTARMA implements autoregressive models for spatio-temporal data at fixed locations, with time-invariant spatial dependency structure. We rely on generalized linear models methodology and unify several approaches for the analysis of spatial count time series. Such models allow the (conditional) mean of the response to depend on past observations, lagged (conditional) expectations, and covariates. The response can be a continuous or a discrete random variable. Additionally, the package develops inference for double generalized linear models, allowing the dispersion parameter(s) of the marginal distributions to be modeled similarly to the mean process. This is a new capability which introduces, for example, spatio-temporal volatility models, such as space-time GARCH processes, and count time series models with spatio-temporal overdispersion and underdispersion. We provide functions for model estimation, simulation, inference, and prediction. Its use is illustrated by data examples.

stat.CO↗

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents

Language-conditioned manipulation requires both precise contact-rich control and robust reasoning over language, scenes, and long horizons. End-to-end Vision-Language-Action (VLA) models provide strong local visuomotor skills, but they are trained on in-distribution task trajectories and often fail under deployment perturbations such as semantic retargeting, goal re-binding, spatial-layout shifts, and unstable local contacts. LLM coding agents provide complementary semantic and compositional reasoning, but purely analytic primitives struggle with irregular grasping, constrained placement, and articulated-object interaction. We present Harness VLA, a memory-augmented agentic framework that exposes a frozen VLA as a retryable contact-rich primitive and composes it with a small fixed library of analytic primitives for grounding, staging, transport, navigation, and release. Rather than expanding the skill library, the harness learns the operating range of these fixed primitives from task-specific execution traces, global success rules, and failure models. By lifting semantic re-grounding, non-contact execution, and VLA re-staging to the planner while reserving the frozen VLA for local contact-rich phases, Harness VLA extends pretrained VLAs beyond their original trajectory distribution without finetuning. Across perturbed tabletop, household kitchen, and clean-to-randomized bimanual manipulation, Harness VLA improves over the strongest relevant baselines by 38.6 and 25.4 percentage points on LIBERO-Pro and RoboCasa365, respectively, and reaches 58.4% on RoboTwin C2R. Real-world demonstrations on dual-Franka robots further show target redirection, grasp recovery, and new task compositions with the same frozen VLA. Code is available at https://github.com/RLinf/RPent.

cs.RO↗

Locally Approximating the Top Eigenvector of Bounded Entry Matrices

We provide a local computation algorithm to approximate the top eigenvector $x \in \mathbb{R}^n$ of a symmetric matrix $A \in \mathbb{R}^{n \times n}$ with entries between $-1$ and $1$, building on the work of Swartworth and Woodruff [SODA 25] who show how to approximate the eigenvalues up to additive-$\varepsilon n$ error using $\tilde{O}(1/\varepsilon^4)$ queries. Our local computation algorithm has a preprocessing complexity of $\tilde{O}(1/\varepsilon^4)$ and per-coordinate query complexity of $\tilde{O}(1/\varepsilon^2)$ for an additive-$\varepsilon n$ approximation whenever {$|λ_{\min}(A)| = O(λ_{\max}(A))$. When $λ_{\min}(A)$ greatly exceeds $λ_{\max}(A)$, our complexity degrades to at most $\tilde{O}(1/\varepsilon^{6.\overline{6}})$ in preprocessing and $\tilde{O}(1/\varepsilon^{3.\overline{3}})$ per query. Furthermore, we show a lower bound of $Ω(n/\varepsilon^2)$ on the total number of queries needed to output an approximately top eigenvector (implying that the per-coordinate query complexity of $Ω(1/\varepsilon^2)$ is necessary). As an application, we use our algorithm to provide local computation algorithms for the sparsest-cut and max-cut problems in the dense graph model of Goldreich, Goldwasser, Ron [JACM 98]. By accessing the top eigenvectors (of an approximate normalized adjacency), we implement local versions of Cheeger's inequality and Trevisan's algorithm [SICOMP 12] to obtain "square-root-opt" approximations in polynomial time (as opposed to exponential-in-$\text{poly}(1/\varepsilon)$ time which is incurred in Goldreich, Goldwasser, Ron.

cs.DS↗

The queer Hero versus the Fool bias of the queer trait: An archetypometric analysis of the collective portrayal of queerness in fictional stories

Visibility in media is pivotal for identity development and for broadening societal views of gender and sexuality. Queer representation has increased in recent years, yet damaging stereotypes and tropes persist. Here, we focus on queer portrayal and its perception by audiences in fictional stories (television, film, and literature) by studying characters by their quantified archetypes which are operationalizations of common conceptions such as Hero, Diva, and Outcast. We use the archetypometrics and Fandom's LGBTQIA+ datasets to study samples of fictional characters along the trait differential spanning straight to queer. We find, quantify, and explain a seeming paradox. The characters with the highest queer score present positive primary archetypes and are typically Heroes rather than Fools, Angels rather than Demons, and Adventurers rather than Traditionalists. But evaluation across many stories for the straight-queer trait itself reveals a strong collective-writing bias towards Fool (away from Hero) and no meaningful loading for the other two dimensions. Our analysis offers a population-scale view of the complexities of queer portrayal, while also pointing to risks in blindly training on many-authored story corpora.

cs.CY↗

Inhomogeneous Strichartz estimates on manifolds with nonpositive curvature and applications

We prove lossless inhomogeneous Strichartz estimates for solutions to the Schrödinger equation on compact manifolds with nonpositive curvature over frequency-dependent time intervals of length $\log λ\cdot λ^{-1} $. As applications, we improve upon the Sobolev norm growth bounds for the cubic NLS established by Planchon, Tzvetkov and Visciglia on 3-dimensional compact manifolds when the manifold also has nonpositive curvature and extend the lossless homogeneous Strichartz estimates on logarithmic time intervals established by the first author and Sogge to Schrödinger operators with critically singular potentials.

math.AP↗