arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,045 records · Page 58Linked to original sources

Cross-Modality Structural Guidance in 3D Latent Diffusion for Robust FLAIR Super-Resolution

High-resolution (HR) MRI acquisition is often hampered by scan time constraints, resulting in anisotropic or low-resolution scans (e.g., thick-slice FLAIR) that limit diagnostic accuracy. While deep learning-based super-resolution (SR) methods show promise, they often hallucinate anatomical details, which can compromise brain structural integrity. To mitigate this limitation, we introduce MR-DiffuSR, a Multi-Resolution Diffusion-based Super-Resolution framework that incorporates HR T1w structural image priors to guide the restoration of thick-slice FLAIR scans and operates in the 3D latent space. Our architecture introduces cross-modality structural swin attention, which derives structural attention maps from the HR T1w and applies them to the low-resolution FLAIR latent features. This design disentangles anatomical structure from modality-specific contrast, effectively preventing hallucinations. Furthermore, we employ a mixed-scale degradation strategy, training the model on a continuum of downsampling factors to ensure robustness to varying slice thicknesses, while optimizing with a DINOv3-based perceptual loss to preserve high-frequency semantic details. Evaluated on the ADNI-4 and ADNI-2 datasets, MR-DiffuSR surpasses both CNN and 2D diffusion approaches, achieving an average PSNR of 32.46 dB, SSIM of 0.97, and LPIPS of 0.07 across all downsampling factors. In downstream white matter hyperintensity segmentation, our model demonstrates exceptional robustness. While baseline performance collapses at 10x downsampling (Dice: 0.51), MR-DiffuSR maintains a Dice score of 0.63, preserving utility even at 7 mm equivalent slice thickness.

cs.CV↗

Decoupled energy-stable Runge-Kutta schemes of arbitrary order for the anisotropic phase-field dendritic crystal growth model

In this paper, we construct an arbitrary-order scheme for the anisotropic phase-field dendritic crystal growth model by introducing a time dependent auxiliary variable. By employing an algebraically stable Runge-Kutta method, the proposed scheme satisfies an unconditional discrete energy dissipation law. To reduce the computational cost, a matrix diagonalization technique is applied to the coupled elliptic system at each time step. This transforms the original system into independent elliptic equations with constant coefficients, which can be solved separately or in parallel. After the decoupling, the auxiliary variable is obtained from a uniquely solvable $q\times q$ algebraic system. For a fixed Fourier-Galerkin space, we further prove $q$th-order convergence in time for the scheme based on a $q$-stage Runge-Kutta method. Numerical experiments in two and three dimensions confirm the theoretical convergence rates and the discrete energy dissipation, and demonstrate the computational efficiency of the decoupled schemes. The effects of anisotropy, latent heat, orientation angle, and initial nuclei on the dendritic morphology are also investigated numerically.

math.NA↗

MPC-Injection: Biasing Off-Policy Locomotion RL Toward Controller-Induced Behavior Basins

Reinforcement learning (RL) for locomotion frequently converges to locally optimal but undeployable behaviors, such as vibrating limbs or scooting on the torso, that maximize return without producing a usable gait. We present MPC-Injection, a low-overhead method that steers RL toward a designer-preferred behavior by inserting transitions generated in the same environment by a model predictive controller (MPC). Unlike reward shaping, MPC-Injection does not require redesigning the task reward, and unlike adversarial imitation learning, it adds no discriminator, no kinematic retargeting, and no auxiliary objective. We analyze how the injected transitions bias the learning, allowing the policy to converge to behaviors that pure RL may fail to reach under simple reward functions. On a 2D walker in simulation and with sim-to-real evaluation on a Go2 quadruped, we show that MPC-Injection produces gaits qualitatively comparable to those of reward shaping and adversarial motion priors. We also show that MPC-Injection can complete a barrel roll that pure RL fails to achieve under the same simple reward and can select between trotting and bounding gaits only through changing the injected MPC data.

cs.RO↗

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

Large pretrained text-to-speech (TTS) models sound almost human for well-resourced languages, but much worse for languages that are rare in their training data. We study this quality gap for Khmer and Korean using VoxCPM2, a 2.4B parameter, tokenizer-free TTS model that joins a MiniCPM-4 language-model backbone with a flow-matching diffusion decoder. We build one shared, language-tagged corpus of 25.5 hours after cleaning and adapt VoxCPM2 with a single Low-Rank Adaptation (LoRA) adapter, trained on both languages at once and added to both the language model and the decoder. The adapter is zero-initialized, so training starts exactly at the original zero-shot model. In native-speaker listening tests, the Khmer Mean Opinion Score (MOS) rises from 3.85 to 4.23 with the best adapter, rank 64. This gain is highly significant under a paired Wilcoxon test with p < 0.001, and it is achieved while training only 0.19 to 3.03 percent of the parameters. Two findings stand out. First, the training loss and human ratings disagree on the best rank. The loss is lowest at rank 128, but MOS peaks at rank 64. Second, the same adapter gives no significant gain for Korean, which the base model already covers well, and a high rank even hurts quality. This shows that adaptation helps mainly where the base model is truly weak.

cs.CL↗

All you need is log

How different are several probability distributions from one another? For two distributions the standard answer is the family of Rényi divergences, singled out by two natural requirements: processing the data never makes distributions easier to tell apart, and independent repetitions add. Many problems in learning and statistics compare more than two distributions at once, such as testing among several hypotheses or bounding generalization against several priors. The same two requirements leave one kind of building block, built on a coincidence probability: how unlikely it is that independent samples, one from each distribution, all show the same empirical distribution. The logarithm is forced because repetitions add, which is already visible for a single experiment repeated. This characterization is known in greater generality, and this paper is about the meaning of its building blocks. On a finite alphabet, each building block indexed by a rational point of the simplex is the exponential rate of that coincidence as the samples grow in fixed proportions. Each is also the limiting free energy of Bayesian inference over distributions. At any amount of data, the free energy of the posterior is the coincidence measure plus two costs: the expected distance from a posterior draw to the most likely distribution, and the information gained per unit of data. Both costs vanish as data accumulate. When the comparison is conditioned on side information, every kind of building block has a conditional counterpart, and the coincidence ones alone do not suffice.

cs.IT↗

Vacancy-mediated nitrogen diffusion and aggregation via high-fluence electron beam irradiation in HPHT synthesized diamond crystal

The negatively charged nitrogen vacancy (NV-) center in diamond is a promising point defect for highly sensitive quantum sensing. The formation of high-density NV- centers is essential for improving sensitivity. We performed room-temperature electron beam irradiation (EBI) and annealing on nitrogen-doped high-pressure high-temperature diamond crystals, aiming to convert all substitutional nitrogen into NV structures by increasing EBI fluence. While the Ns0 to NV0 and NV- conversion process dominated at low EBI fluences, in a high-fluence region, NV0 and NV- center and Ns0 and Ns+ concentrations decreased with increasing EBI fluence, indicating the formation of unknown nitrogen-related defects such as the H3 center, which is a nitrogen and vacancy aggregation defect. Although H3 centers were observed at high EBI fluence, their annealing temperature of 1375 +- 25 °C was lower than the typically reported temperatures over 1600 °C. We attribute this low-temperature formation to vacancy-mediated nitrogen diffusion and aggregation.

cond-mat.mtrl-sci↗

Soft Contributions Stabilize NNLO QCD Corrections to Quarkonium Production and Decay

Next-to-next-to-leading order (NNLO) QCD corrections to quarkonium production and decay are known to exhibit perturbative instabilities within non-relativistic QCD. We identify the origin of this problem and propose a simple remedy. Applying our approach to $S$-wave color-singlet quarkonium processes, we achieve substantially improved perturbative convergence and agreement with experimental data.

hep-ph↗

Dissipative Effects in Transmission Line Analogues of Hawking Radiation

Hawking radiation is a fundamental result of quantum field theory in curved spacetime, yet its direct observation remains beyond current experimental capabilities. Circuit quantum electrodynamics provides a practical platform for realizing analogue systems where Hawking-like radiation may be studied under controlled laboratory conditions. In this work, we analyze two superconducting-circuit analogues of Schwarzschild black holes: a tunable dc-SQUID transmission line and a SNAIL-based transmission line supporting solitonic solutions of the KdV equation. We investigate the conditions under which these architectures can generate an observable Hawking temperature and study the impact of dissipation and thermal noise using an open quantum systems approach. To assess the observability of the Hawking signal, we propose complementing particle number measurements with estimates of the Hilbert-Schmidt distance to the thermal bath. Our analysis establishes practical detectability thresholds and shows that Hawking temperatures above approximately 73 mK remain distinguishable under realistic experimental conditions. While the tunable transmission line architecture can reach temperatures of about 113 mK and therefore appears more viable, the solitonic model requires further optimization and more demanding experimental conditions.

quant-ph↗

The Prepare and Broadcast Scenario

We introduce the dimension-restricted prepare and broadcast (PAB) scenario, which generalizes standard prepare-and-measure frameworks. Here, the system prepared by a sender undergoes a broadcasting transformation before being locally measured by multiple receivers. We develop a hierarchy of classical, quantum, and nonsignalling models describing this scenario, characterize their corresponding correlation sets, and derive new families of Bell-like inequalities together with linear and semidefinite programming methods for their certification. First, assuming shared randomness, we prove that the hierarchy collapses into a single set whenever we consider only one measurement per party. Then, considering multiple possible measurements, we show that PAB scenarios allow the activation of nonclassicality, revealing genuinely nonclassical features in resources that admit classical descriptions in standard prepare-and-measure or Bell settings.

quant-ph↗

The Cross-Section of Stock Returns and AI Exposure

We study 380 trillion tokens of realized AI consumption across more than four hundred LLMs. We build a high-frequency AI factor and show that a long-short strategy based on firms' AI exposure earns significantly positive returns. The average strategy return is larger based on intensive, frontier-oriented AI consumption but smaller based on casual or open-weight usage. Internationally, the return spread is significant in developed countries but insignificant in emerging markets. Examining occupational AI exposure, we find more positive exposure in occupations intensive in nonroutine interactive tasks and more negative exposure in those intensive in nonroutine analytical tasks.

cs.CY↗

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain predominantly outcome-centric, relying on final answer correctness or sequence-level preferences. This suffers from sparse credit assignment, making it difficult to optimize the reasoning process essential for clinical applications. Our analysis reveals that cascading errors from early-stage reasoning failures are a leading cause of incorrect predictions in medical visual question answering (VQA) benchmarks. Motivated by this, we propose Medical Reasoning-aware Policy Optimization (MRPO), an RL algorithm that incorporates step-wise process rewards. When the final answer is incorrect, MRPO assigns exponentially larger penalties to tokens in earlier invalid reasoning steps, breaking failure cascades without compromising successful paths. Across four multimodal LLM backbones, MRPO consistently outperforms standard GRPO and a recent RL baseline, and on Qwen3-VL-8B-Thinking even surpasses substantially larger medical MLLMs such as HuatuoGPT-Vision-34B by 4.59 points. Moreover, MRPO reduces early-stage reasoning failures from 58.6% to 13.4%, showing that targeted mitigation of cascading failures improves both reasoning quality and final answer accuracy. Our code is available at https://github.com/dmis-lab/MRPO

cs.CV↗

A multilevel stochastic-gradient neural solver for boundary integral equations

We propose a multilevel stochastic-gradient neural solver (MLSG) for second-kind boundary integral equations. MLSG represents the unknown boundary density using a neural network optimized by stochastic residual minimization over a hierarchy of successively refined Nyström discretizations. Upon transitioning from one level to the next, the network parameters obtained on the previous level initialize training on the current one. This coarse-to-fine strategy retains a continuous, grid-independent density representation and is designed to reduce the total computational effort required to reach a prescribed residual tolerance at the target resolution. The algorithm avoids grid-transfer operators and hierarchical fast-summation machinery, relying instead on batched kernel evaluations and standard network forward and backward passes that map efficiently onto modern GPU architectures. For uniformly stable second-kind discretizations, so strongly nonuniform contraction rates originate in the empirical neural tangent kernel (NTK) rather than in the discretized operator. Within each level, parameter updates can reshape the NTK, while refinement re-samples the tangent kernel on a richer discrete space and reveals directions that were not adequately resolved on coarser levels. A cross-level estimate bounds the warm-start loss in terms of the preceding training tolerance and quadrature error, motivating a tolerance schedule that balances optimization and discretization errors. Experiments on Laplace/Poisson and Helmholtz problems in two and three dimensions, together with an exterior Robin problem on a hypersurface in $\R^4$, demonstrate the method under both parametric and signed-distance surface representations at up to million-scale discretizations.

math.NA↗

Bringing Agentic Search to Earth Observation Data Discovery

NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search framework for geoscience data discovery that takes a natural-language research query and returns matching datasets and tools. We demonstrate that, in the era of large language models, the latent value of knowledge graphs (KGs) can be substantially amplified through agentic search. From the NASA Earth Observation Knowledge Graph (NASA EO-KG) we derive NASA-EO-Bench, an open benchmark of 47k query-dataset pairs (21k task-based queries). A neural scorer fine-tuned on NASA-EO-Bench beats cosine and BM25 baselines. Further combining it with BM25 via score fusion raises both Recall@10 (R@10) and MRR to over 5x the unadapted cosine baseline. On top of this supervised pipeline, a zero-shot reranking stage lifts MRR by 16%, significant under a paired bootstrap, with no additional training, and autonomous web and arXiv tool use adds a further gain, showing that LLM reasoning is complementary to supervised retrieval.

cs.IR↗

BRST-BV approach to fields in Poincare patch of AdS

We use the Poincare parametrization of AdS space to develop a general BRST-BV approach for free fields. A general expression for the BRST-BV Lagrangian of fields with arbitrary masses and symmetry types is obtained. We apply this general framework to study totally symmetric massless, massive, and partially-massless fields with arbitrary integer spin and a continuous-spin field. For these fields, the constrained and unconstrained BRST-BV formulations are developed. We investigate both irreducible and reducible fields. In addition, we demonstrate the matching between the obtained BRST-BV Lagrangian and the metric-like Lagrangian formulated in terms of the modified de Donder derivative. Finally, a realization of AdS space symmetries is obtained within the space of fields and antifields entering the BRST-BV formulation. Modulo the choice of the Poincare parametrization of AdS space, the Lagrangian and the entire approach are manifestly Lorentz invariant.

hep-th↗

Krylov-Lie Algebras for Variational Quantum Algorithms: Geometric, Depth-Aware Insights into Expressivity and Trainability

Variational quantum algorithms (VQAs) are a leading approach to near-term quantum computation, but their utility is limited by barren plateaus and other pathologies in their loss landscapes. Existing landscape theories based on dynamical Lie algebras, Jordan-algebraic Wishart systems, approximate t-designs, and Haar-random circuits are foundational, but they often neglect the finite-depth geometry of realistic ansätze and are therefore ill-suited to the shallow-depth regime, where VQAs are poor approximators of 2-designs and trainability is most feasible. This work introduces Krylov algebras, algebraic structures induced by the Krylov span of a finite generator set acting on one or more seed vectors, as a framework for VQA landscape theory. We show that VQA reachable manifolds can be approximated in a numerically robust, geometrically faithful fashion by Krylov-Lie algebras and groups, and that these structures induce canonical invariant measures for computing expectation values and variances under general sampling measures. In particular, we derive weighted non-Haar variance formulas that recover the usual Lie-algebraic Haar formulas as a special case while isolating non-Haar effects into explicit correction terms. We also show that the common heuristic that sufficiently deep circuit ensembles must converge to Haar fails in general without additional hypotheses, identify concrete obstructions to naive Haar convergence, and recover convergence under natural necessary and sufficient ergodic conditions. Lastly, our formulas further imply that non-Haar contributions to landscape statistics may mitigate barren plateaus by reweighting the visible sectors of the loss landscape, suggesting that VQAs may be more trainable than recent literature has posited.

quant-ph↗

Comparative Evaluation of Encapsulation Methods for Endohedral Doping of Single-Wall Carbon Nanotubes

Single wall carbon nanotubes (SWCNTs) are promising building blocks for nanoelectronic and optoelectronic devices, yet reliable and stable doping, particularly n type, remains challenging due to strong environmental sensitivity and competing extrinsic effects. Encapsulation of charge transfer molecules within the SWCNT cavity offers a promising route to stable doping while preserving the nanotubes outer surface for subsequent processing. Here, we systematically investigate the filling of arc discharge SWCNTs with the electron donor tetrathiafulvalene and electron acceptor tetracyanoquinodimethane, comparing different methods for filling, including melt filling, solution reflux, and vacuum phase sublimation. We follow the entire processing workflow from raw, unfilled powders to aqueous dispersions and employ density gradient ultracentrifugation to separate filled from empty nanotubes as well as metallic from semiconducting ones. Encapsulation efficiency and electronic modification are assessed using absorption spectroscopy, resonant Raman scattering, thermogravimetric analysis, X-ray photoelectron spectroscopy and electron paramagnetic resonance. Finally, we introduce a complementary vacuum-phase method that removes externally adsorbed molecules without extensive solvent washing, enabling cleaner encapsulated systems.

cond-mat.mes-hall↗

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and language development. However, utterance-level timestamps provided in this corpus are often noisy, incomplete, or misaligned with the audio. As a result, utterances cannot always be reliably localized within long recordings, which limits the direct use of these data for training and evaluating speech models. In this work, we propose BEACON (Boundary Estimation via Alignment CONsensus), an ensemble timestamp-curation framework that refines utterance-level timestamps by aggregating knowledge from multiple off-the-shelf ASR models. Specifically, each model's word-level timestamp predictions are first aligned to provided human transcripts, and the final utterance time boundaries are determined by a consensus voting strategy. The framework is corpus-agnostic and applies to any long-form recording paired with a trusted transcript whose timestamps are unreliable or missing, offering a general recipe for timestamp curation. Leveraging this pipeline, we curate and release a 413-hour general-purpose child-speech dataset with corrected utterance-level timestamps, together with a 283-hour quality-controlled subset for ASR training. Fine-tuning on this subset yields up to an average 19.5% relative WER reduction on four out-of-domain child-speech benchmarks.

eess.AS↗

PhysMiner: An Agentic AI Framework for Automated Flow Component Analysis

Uncovering the physical mechanisms of turbulent flows remains a fundamental challenge in fluid mechanics. In particular, conventional velocity-gradient analysis methods suffer from shear contamination, which hinders accurate identification of the dominant physical mechanisms. This study presents PhysMiner, an automated framework integrating the triple decomposition method of the velocity gradient tensor with large language model-driven reasoning for turbulence-physics discovery. The triple decomposition module automatically decomposes flow fields into rigid rotation, pure shearing, and normal straining components, enabling statistical analysis, contour visualization, vortex-line extraction, and threshold-insensitive vortex identification while eliminating shear contamination. These automated capabilities are validated across five benchmarks, ranging from canonical configurations to complex engineering flows. A discover-physics agent combines flow statistics, spatial structures, and literature-derived knowledge to perform pattern recognition and physical inference, while a review Agent iteratively validates physical consistency to ensure reliable conclusions. A continuously evolving Triple Decomposition Library accumulates statistical knowledge from successfully analyzed flows, enabling cross-case comparison and progressive enhancement of inductive capability. The complete PhysMiner pipeline is validated end-to-end on the periodic hill flow, where the framework autonomously generates turbulence modeling recommendations and derives an improved subgrid-scale model with superior Reynolds-stress predictions. PhysMiner is open to the public and establishes a foundation for long-term collaborative advancement in automated turbulence-physics discovery.

physics.flu-dyn↗