arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Estimates of ground state energies for the quantum SK and 2D-EA models, using deGennes-Suzuki-Kubo mean-field annealing dynamics

We perform large scale quantum annealing of the Sherrington-Kirkpatrick (SK) spin glass up to a system size $N=40000$ to estimate its ground state energy using the deGennes-Suzuki-Kubo mean-field quantum Ising dynamics, extending the earlier results (reported in Eur. Phys. J. B {\bf 98}, 226 (2025)). Here we numerically solve the deGennes-Suzuki-Kubo annealing dynamics to obtain the spin configurations and subsequently the ground state energy for a given system size at the end of the annealing, starting from a quantum paramagnetic state. The method shows high efficiency, with an overall algorithmic cost of $O(N^3)$ in estimating the energy of the ground state. We later extend this quantum annealing study to estimate the ground state energies (starting again from the quantum paramagnetic phase, annealing down to any desired low value of the transverse field) for the Edwards-Anderson (EA) spin glass model on a square lattice.

cond-mat.dis-nn↗

Vertex-transitive quantum graphs

We define a quantum graph to be vertex-transitive if the join of its automorphism group is the maximum quantum relation on its quantum vertex set, in direct analogy with the classical case. All simple quantum graphs in $M_2(\mathbb C)$ are vertex-transitive, but many simple quantum graphs in $M_3(\mathbb C)$ are not vertex-transitive. We provide a complete classification of vertex-transitive quantum graphs in $M_3(\mathbb C)$ up to isomorphism. To do this, we introduce a polynomial invariant for quantum graphs in $M_n(\mathbb C)$, which we call the panoramic polynomial.

math.OA↗

Vision-Language Model Ensembles Achieve Human-Expert Accuracy for Galaxy Merger Classification

Models (VLMs) combined using a Bayesian statistical framework can classify galaxy merger morphologies with accuracy comparable to trained human experts. We deploy 15 VLM classifier configurations, spanning four model architectures (Gemma-4 E2B, Gemma-4 E4B, Qwen2.5-VL, and Qwen3-VL) tested with up to four prompt engineering strategies each. We evaluate their performance against a truth-known sample of 41 VELA+SUNRISE mock galaxy images from Lambrides et al. (2021a). As a proof-of-concept, all validation is performed on these mock images; application to real observational samples will require additional calibration and observational validation. The VLM ensemble achieves 83.3% accuracy on confident classifications (merger probability pM >= 0.8 or pM <= 0.2) and 58.3% completeness, with 5 misclassified galaxies, compared to 85% for both human accuracy and completeness. The ensemble recovers the population merger fraction to within 0.66 sigma of the truth (fM = 0.52 +/- 0.09 vs. true value of 0.585). Bayesian weighting improves overall accuracy by 17.1 percentage points over simple majority voting, with sensitivity improving by 29.2 percentage points. The ensemble produces 5 misclassifications (2 false positives, 3 false negatives), comparable to the 6 misclassifications (5 false positives, 1 false negative) reported for human classifiers by L21. The error-profile differences are not statistically significant for this sample. VLMs also produce more moderate per-galaxy merger probability distributions (27% uncertain) than the more polarized human distributions (15% uncertain), though this difference is also consistent with statistical fluctuation. These results establish VLMs as scalable, reproducible alternatives to human classifiers within a Bayesian probabilistic merger-fraction framework, for large-survey applications.

astro-ph.GA↗

A Unified Benchmark for Dynamic Medical Treatment Reinforcement Learning

Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals. Existing RL formulations and simulated environments, however, are based on discrete-time MDPs with fixed decision intervals. Thus, it remains difficult to evaluate whether RL methods can handle time-interval-dependent disease progression, personalized treatment response, and safety between consecutive measurement points. To address this gap, we introduce MedGym, a benchmark environment for dynamic treatment recommendation. MedGym models longitudinal patient evolution in a continuous-time framework and constructs a configurable medical RL benchmark from clinical data by using Physics-Informed Neural Networks. The resulting benchmark enables direct comparison between discrete-time and continuous-time methods under irregular treatment timing and patient-specific dynamics. Furthermore, MedGym supports evaluation from clinically important perspectives, such as personalization and trajectory-level safety. By providing a standardized and configurable benchmark for continuous-time dynamic treatment, MedGym enables more realistic and informative evaluation of medical RL methods.

cs.LG↗

Enhancing the Angular Resolution of Large Array of imaging atmospheric Cherenkov Telescope (LACT) at Ultra-High Energies

The Large Array of Imaging Atmospheric Cherenkov Telescopes (LACT) is dedicated to high-resolution morphological studies of PeVatrons. In this work, we present a fundamental investigation into stereoscopic direction reconstruction for the LACT array, specifically addressing the challenges of ultra-high-energy observations. We demonstrate that the standard Hillas parameterization introduces a significant reconstruction bias under severe image leakage. To mitigate this, we introduce an approach utilizing a 2D Gaussian fit, achieving an exceptional angular resolution of better than $0.06^\circ$ at $100\text{ TeV}$ within the central $0^\circ\text{--}1^\circ$ offset bin, and maintaining better than $0.12^\circ$ across offsets up to $4^{\circ}$. Building on this robust baseline, we evaluate advanced weighting schemes by utilizing a LightGBM-based quantile regression model to independently estimate single-image quality. Applying these quality-based weights yields a consistent improvement of $0.02^\circ$ to $0.03^\circ$ for high-energy, large-offset events using both the \textit{HillasWeightedSum} and \textit{HillasWeightedDisp} methods. Finally, to establish a theoretical performance ceiling, we explore a pixel-wise likelihood reconstruction technique utilizing Neural Ratio Estimation. While its practical realization depends heavily on minimizing the gap between Monte Carlo simulations and observational data, this exploratory approach demonstrates the potential to yield an overall improvement of approximately 15\% to 40\% at $100~\rm TeV$ across the entire field of view. Such high angular resolution is critical for disentangling complex emission regions and mapping the internal structures of PeVatrons.

astro-ph.HE↗

Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners

Current benchmarks for embodied vision-language planning inadvertently favor linguistic next-token prediction over physically grounded next-state reasoning. This rewards models that mimic statistical language priors rather than track true causal dependencies, reducing complex physical planning to shallow sequence modeling. Hence, achieving genuine physical autonomy requires a fundamental shift from linguistically grounded token prediction toward physically grounded causal reasoning. To this end, we introduce Causal-Plan-Bench, a high-fidelity diagnostic suite spanning four causal dimensions, curated via multi-stage verification. To endow models with this capability, a four-stage annotation pipeline extracts structured interaction records from egocentric videos to construct Causal-Plan-1M, a dense million-scale corpus of explicit causal reasoning traces. Extensive evaluation reveals a striking gap: leading models struggle to demonstrate genuine physical agency -- even GPT-6-astra scores only 43.04. In contrast, our tailored training recipe enables Causal Planner to internalize the complex physical logic required for accurate next-state estimation. Built upon Qwen3-VL-8B, Causal Planner raises its backbone's score from 33.23 to 45.28, a 36.3% relative gain, and improves on three external benchmarks without benchmark-specific adaptation. We further observe an empirical Causal-Supervision Scaling Trend. Paired no-vision controls also reveal substantial visual dependence, while cross-judge comparisons and human scoring assess the reliability of automated evaluation. More importantly, we initiate the first effort to turn agents from superficial token predictors into physically grounded causal reasoners, bridging language modeling and world modeling.

cs.AI↗

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spatially inconsistent, and indifferent to the human subject driving the scene. In this work, we present Auteur, a method for language-driven, human-centric camera framing in generative video. Our core insight is that professional filmmakers conceive shots not as world-space trajectories but as framings defined relative to the actor, encoding shot size, angle, and composition as functions of human pose and motion. We formalize this intuition as a human-centric camera parameterization and introduce a Domain-Specific Language (DSL) that is convertible to standard 6-DoF camera parameters. A fine-tuned multimodal large language model then acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are deterministically interpolated into continuous camera trajectories, which are then provided as input to video generators. We train and evaluate Auteur on a new dataset of 34K aligned text, human motion, and DSL-annotated camera trajectories drawn from procedural synthesis and real-world movie footage from the CondensedMovies dataset. Auteur enables cinematographic framing of human-centered scenes, a capability largely absent in prior generative models. To assess this behavior, we propose new framing-focused metrics, and our experiments show that Auteur consistently outperforms existing methods. Project page is https://cyberiada.github.io/Auteur/

cs.CV↗

Semi-analytical two-loop QCD corrections to $e^+e^-\to J/ψ+χ_{cJ}$ at B factories

In this work, we compute the next-to-next-to-leading-order (NNLO) QCD corrections to the process $e^+e^-\to J/ψ+χ_{cJ}$ at B factories within the NRQCD factorization framework. The helicity amplitudes are obtained via asymptotic expansions around $r=0$ and $r=1$, with $r=16m_c^2/s$. Our asymptotic expressions reproduce the exact numerical results with high accuracy across the kinematic range $0\le r \le 1$, achieving a relative error below $10^{-5}$, which is sufficient for phenomenological applications. Notably, the large logarithmic terms are obtained analytically, and the explicit expressions established in this work provide a useful basis for further theoretical investigations. We compute the unpolarized cross sections. The $\mathcal{O}(α_s)$ correction is found to be large, while the $\mathcal{O}(α_s^2)$ correction for $χ_{c0}$ production amounts to $33\%$ of the leading-order (LO) cross section, significantly reducing the scale uncertainties. For $χ_{c1}$, the $\mathcal{O}(α_s)$ and $\mathcal{O}(α_s^2)$ corrections correspond to $35\%$ and $-15\%$, respectively. For $χ_{c2}$, the corresponding corrections are $24\%$ and $-38\%$. The large cancellation between the corrections for $χ_{c2}$ brings the NNLO cross section close to the LO prediction. Our prediction for $χ_{c0}$ is consistent with both the {\tt Belle} and {\tt BaBar} data within uncertainties. We also predict the angular distribution parameters $α^J_θ$, which are independent of nonperturbative inputs. A sharp discrepancy between the theory and the {\tt Belle} measurement is observed for $α^0_θ$, calling for further experimental and theoretical investigations. Moreover, future measurements of the angular distribution parameters for $χ_{c1}$ and $χ_{c2}$ will provide important tests of the theoretical framework.

hep-ph↗

On vacua and bounded masses in the general 2HDM

Two Higgs doublets models with a scalar potential that breaks electroweak symmetry spontaneously can have either one or two local minima. While potentials with one minimum can have a decoupling regime where all the new scalars are heavy, we show that, for potentials with $two$ local minima, the masses of all the scalars are bounded if the dimensionless quartic couplings obey perturbativity constraints.

hep-ph↗

Dirac operators for infinite-dimensional color Lie algebras

We develop the Dirac formalism for infinite-dimensional quadratic $\mathbb{Z}$-graded color Lie algebras with finite-dimensional components. Cubic Dirac operators are defined in completions of the quantum Weil algebra determined by the $\mathbb{Z}$-grading. The same grading fixes the normal-ordering convention. Normal ordering introduces a cohomological obstruction to the construction, measured by a color analogue of the Kac-Peterson class. When this class is trivial, we construct cubic and relative cubic Dirac operators satisfying the expected invariance properties and Parthasarathy-type square formulas. We further extend the Chern-Weil homomorphism to completed $\mathfrak{g}$-differential algebras and use it to identify the classical precursor of the cubic Dirac operator with the Chern-Simons element associated with the invariant quadratic polynomial determined by the quadratic structure. As applications, we consider symmetrizable Kac-Moody superalgebras. In this setting, the Kac-Peterson class is trivial, with primitive given by the Weyl vector, which yields the linear correction defining the cubic Dirac operator. We then use the relative Dirac operator to extract representation-theoretic information from highest weight supermodules. As an explicit example, for the affine Kac-Moody superalgebra associated with $\mathfrak{osp}(1\vert 2n)$, we compute the kernel of $\operatorname{D}_{\mathfrak{g},\mathfrak{g}_{\bar{0}}}$ on integrable highest weight supermodules. Finally, for unitarizable highest weight supermodules, we explain why the usual Dirac inequality is not available in the affine setting.

math.RT↗

TAGA: Terrain-aware Active Gaze Learning for Generalizable Agile Humanoid Locomotion

Agile humanoid locomotion across diverse challenging terrain demands both wide perceptual coverage and precise local geometry understanding. Motivated by the way humans selectively look at relevant terrain during locomotion, we introduce TAGA, a Terrain-aware Active Gaze learning framework for Attention-based humanoid control. By fusing vision, proprioception, and motion commands, our framework guides the model to learn anticipatory cues and actively attend to specific areas of the height scan, selectively using these informative regions for the downstream network. This adaptively increases the information density of observations under tight onboard computational constraints, thus enabling fine-grained perceptive locomotion over larger-scale terrains. We find that such gaze behaviors can naturally emerge through reinforcement learning alone, without requiring additional supervision or explicit guidance, significantly improving training efficiency. As a result, the trained policy demonstrates robust and generalizable locomotion in simulation and on hardware, including reliable terrain-aware foothold selection, elevated-platform traversal, competitive sparse-foothold traversal, and the largest reported real-world gap traversal distance of 1.2m among perceptive humanoid locomotion systems, while maintaining stability under severe perceptual disturbances and environmental interference.

cs.RO↗

Deep Learning-based 3D Oral Cavity Reconstruction Using 2D Intraoral Images

Oral 3D modelling is one of the most essential stages in dentistry, and many different approaches, such as impression taking and intraoral scanning, are commonly used for this phase, each with notable limitations. Impression taking, which involves placing alginate or silicone material in a tray and inserting it into the patient's oral cavity to form a negative mold, suffers from significant patient discomfort, material deformation errors, and difficulties in storage and transportation. Intraoral scanners, which directly scan oral structures in real time using structured light or laser technology, produce state-of-the-art results but are associated with substantially high equipment costs. To address these limitations, this paper proposes a software-based approach that reconstructs a 3D oral model using only ten 2D intraoral images captured from different angles, requiring no dedicated hardware devices. The proposed method reduces cost, eliminates the need for physical scanning equipment, minimises patient discomfort, and enables automated 3D reconstruction. The model is trained on the publicly available Teeth3DS dataset, comprising 950 upper jaw samples, and employs MobileNetV2 as the image encoder combined with Multi-head Attention for multi-view feature fusion. The proposed model achieves an accuracy of 77.49%, measured by nearest-neighbor matching with a distance threshold of 0.035. However, predicted vertices tend to concentrate in high-density regions of the ground truth, resulting in uneven point distribution across the reconstructed model.

cs.CV↗

3D Oral Modelling with Improved Vertex Distribution Using Matching-Based Learning

In our previous work, a deep learning-based framework for 3D intraoral reconstruction was proposed. The model directly predicts explicit 3D point cloud coordinates from ten fixed-angle intraoral images, employing MobileNetV2 and Multi-head Attention for multi-view feature fusion, with a combined L1 Loss and Chamfer Distance as the loss function. Although the model achieved an accuracy of 77.49%, predicted vertices tended to concentrate in high-density regions of the ground truth, leaving other regions largely uncovered. In this paper, an improved loss function is proposed to address this limitation. Hungarian matching with filtering and Repulsion Loss are introduced to enforce more uniform vertex distribution across the reconstructed model. The proposed model achieves an accuracy of 68.02%, which is numerically lower than the previous model. However, the vertex clustering issue observed in the prior work is substantially alleviated, with predicted vertices distributed more evenly across the entire reconstructed surface.

cs.CV↗

Hyperbolic dynamic friction models for viscoelastic sliding and rolling contact

This paper considers the sliding and rolling contact between viscoelastic bodies. Combining linear viscoelastic rheologies for bristle-like elements with nonlinear dynamic friction models, and under the assumption of small deformations and displacement gradients, it derives a class of viscoelasto-kinematic equations, formulated as a system of semilinear partial differential equations (PDEs) governing the evolution of the frictional force, bristle deformations, and internal state variables at the interface between the contacting bodies. The resulting system is analysed mathematically, demonstrating that linear viscoelasticity preserves the hyperbolic character of the PDE systems typically encountered in rolling contact. The proposed theory is illustrated through representative examples of both sliding and rolling contact, highlighting that these two processes, whilst often treated as distinct, may in fact exhibit closely related underlying dynamics. Overall, the framework provides a general theoretical setting applicable to a broad class of viscoelastic frictional systems.

physics.app-ph↗

All-multiplicity monodromy and KLT relations for AdS string integrals

We propose and study all-multiplicity building blocks for tree-level string amplitudes in AdS. These are worldsheet integrals obtained by dressing the corresponding flat-space disc and sphere integrals with multivariable multiple polylogarithms and their single-valued analogues, respectively. We derive monodromy relations for the open-string building blocks and a KLT factorisation for their closed-string counterparts. This extends the non-commutative AdS uplift of lower-point flat-space structures to general $n$-point kinematics.

hep-th↗

Recovering the Zipfian Distribution in Unsupervised Term Discovery

Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate types. True lexicons follow a Zipfian distribution, yet the dominant centre-based clustering approach -- K-means -- produces a more uniform distribution due to an inductive bias toward spherical clusters. In this paper we revisit graph-based clustering as a bottom-up alternative, where segment embeddings are connected by pairwise similarity and partitioned using the Leiden algorithm. We show that graph clustering substantially outperforms centre-based approaches (K-means, GMM, BIRCH) in both word- and syllable-level lexicon discovery across three languages, producing more Zipf-like distributions. Another bottom-up approach, agglomerative clustering with average linkage, also performs well, although it is computationally less efficient and allows for less control over the resulting distribution. Our work calls into question the dominance of centre-based clustering for term discovery, and promotes graph clustering as an attractive alternative.

eess.AS↗

Derivatives of Tensor Products and Applications to Spaces $C(K^n,X)$

In this paper, we develop an abstract theory of derivatives for Banach spaces based on objects that we call \emph{bidual assignments}. This framework encompasses both the Semadeni derivative and the recently introduced Semadeni--Pełczyński derivative. More generally, suitable ideals of subsets of dual spaces give rise to a broad family of derivatives within this setting. We establish direct-sum and tensor-product formulas for these derivatives, showing that they behave naturally with respect to direct sums and injective tensor products. We then obtain explicit descriptions of derivatives associated with compact trees and finite products of compact lines. In particular, we compute iterated derivatives and use them to derive isomorphic invariants for vector-valued spaces of continuous functions. As one consequence, if $K=\prod_{i=1}^n K_i$ and $L=\prod_{j=1}^m L_j$, where $n,m\geq1$ and all the factors are compact lines of uncountable character, then, for $1\leq p,q<\infty$, \[ C(K,\ell_p)\sim C(L,\ell_q) \] implies that $n=m$ and $p=q$. We also establish classification results for spaces of the form $C(K^n,X)$. In particular, for uncountable ordinals $α$ and $β$, an integer $n\geq1$, and Banach spaces $X$ satisfying suitable rigidity assumptions, we prove that \[ C([0,α]^n,X)\sim C([0,β]^n,X) \quad\text{if and only if}\quad C([0,α])\sim C([0,β]). \] This extends Kislyakov's classification of the spaces $C([0,α])$ and its vector-valued extension due to Galego to finite powers of ordinal intervals.

math.FA↗

GUIDE: Goal-Initialized Directional Understanding for End-to-End Legged Navigation

End-to-end reinforcement learning (RL) has shown strong potential for legged robot navigation, yet existing approaches commonly rely on continuously updated robot-to-goal states from external state estimation modules, leaving part of the navigation problem outside the learned policy. In this work, we seek to push the limits of end-to-end sim-to-real RL navigation by investigating whether a legged robot can internally maintain the spatial context required for long-horizon navigation. To this end, we study goal-initialized navigation, where the goal is provided only once at the beginning of an episode, with no subsequent external relative-goal updates. We present GUIDE, an end-to-end RL framework that jointly learns navigation and internal directional awareness from onboard observations. GUIDE leverages multi-frequency proprioceptive history to capture egomotion and auxiliary spatial-anchor prediction to maintain task-relevant spatial states, while temporal depth observations provide local environmental geometry. The navigation policy is trained entirely in simulation and transferred zero-shot to the real world. Across cluttered environments and structured mazes, GUIDE reliably avoids obstacles, escapes dead ends, and reaches distant goals using only onboard sensing. These results demonstrate that robust sim-to-real legged navigation can be achieved without continuously providing external robot-to-goal estimation, opening a promising direction for future research on more self-contained end-to-end navigation.

cs.RO↗