arXiv ScienceSearch

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Dirac Observables for Gowdy Cosmologies regular at the Big Bang

Gowdy cosmologies are exact, spatially inhomogeneous solutions of the vacuum Einstein equations which describe nonlinear gravitational waves coalescing at the Big Bang singularity. With toroidal spatial sections they provenly have the Asymptotic Velocity Domination property, in that close to the Big Bang dynamical spatial gradients fade out and the dynamics is governed by a Carroll-type gravity theory. Here we construct an infinite set of Dirac observables for Gowdy cosmologies, valid off-shell, strongly, and without gauge fixing. These observables stay regular at the Big Bang and can be matched to much simpler Dirac observables of the Carroll-type gravity theory. Conversely, in an adapted foliation there is a systematic anti-Newtonian expansion (in inverse powers of the reduced Newton constant) of the full Dirac observables whose leading terms are the Carroll ones. In particular, this provides an off-shell generalization of the Asymptotic Velocity Domination property.

gr-qc

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level visual details, such as smoke, specular reflections, and fluid motion, rather than semantic structures essential for surgical understanding. We present SurgMotion, a video-native foundation model that shifts the learning paradigm from pixel-level reconstruction to latent motion prediction. Built on the Video Joint Embedding Predictive Architecture (V-JEPA), SurgMotion introduces three key technical innovations tailored to surgical videos: (1) motion-guided latent masked prediction to prioritize semantically meaningful regions, (2) spatiotemporal affinity self-distillation to enforce relational consistency, and (3) spatiotemporal feature diversity regularization (SFDR) to prevent representation collapse in texture-sparse surgical scenes. To enable large-scale pretraining, we curate SurgMotion-15M, the largest surgical video dataset to date, comprising 3,658 hours of video from 50 sources across 13 anatomical regions. Extensive experiments across 17 benchmarks demonstrate that SurgMotion significantly outperforms state-of-the-art methods on surgical workflow recognition, achieving 14.6 percent improvement in F1 score on EgoSurgery and 10.3 percent on PitVis; on action triplet recognition with 39.54 percent mAP-IVT on CholecT50; as well as on skill assessment, polyp segmentation, and depth estimation. These results establish SurgMotion as a new standard for universal, motion-oriented surgical video understanding.

cs.CV

Spontaneous Parity Breaking in Quantum Antiferromagnets on the Triangular Lattice

Frustration on the triangular lattice has long been a source of intriguing and often debated phases in many-body systems. Although symmetry analysis has been employed, the role of the seemingly trivial parity symmetry has received little attention. In this work, we show that phases induced by frustration are systematically shaped by an implicit rule-of-thumb associated with spontaneous parity breaking in weak longitudinal field. This principle enables us to anticipate and rationalize the regimes and conditions under which nontrivial phases emerge. For the spin-$S$ antiferromagnetic XXZ model, we demonstrate that a controversial parity-broken phase appears at intermediate values of $S$. In bilayer systems, enhanced frustration leads to additional phases, such as supersolids, whose properties can be classified by their characteristic parity features. Benefiting from our improved tensor network contraction techniques, we confirm these results through large-scale tensor-network calculations. This study offers an alternative viewpoint and a systematic approach for examining the interplay between spin, symmetry, and frustration in many-body systems.

cond-mat.str-el

NavDreamer: Video Models as Zero-Shot 3D Navigators

Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture temporal dynamics and physical laws. We propose NavDreamer, a video-based framework for 3D navigation that leverages generative video models as a universal interface between language instructions and navigation trajectories. Our main hypothesis is that video's ability to encode spatiotemporal information and physical dynamics, combined with internet-scale availability, enables strong zero-shot generalization in navigation. To mitigate the stochasticity of generative predictions, we introduce a sampling-based optimization method that utilizes a VLM for trajectory scoring and selection. An inverse dynamics model is employed to decode executable waypoints from generated video plans for navigation. To systematically evaluate this paradigm in several video model backbones, we introduce a comprehensive benchmark covering object navigation, precise navigation, spatial grounding, language control, and scene reasoning. Extensive experiments demonstrate robust generalization across novel objects and unseen environments, with ablation studies revealing that navigation's high-level decision-making nature makes it particularly suited for video-based planning.

cs.RO

Althea: The Fact-Checking--Metalearning Tradeoff in AI-Assisted Verification

Fact-checking systems must be scalable and epistemically trustworthy. We introduce Althea, a retrieval-augmented system for user-driven claim evaluation that matches standard pipelines on AVeriTeC while improving supported/refuted discrimination. A longitudinal survey experiment (N=961) treats a ten-day follow-up as a fading test: after modeling a verification procedure, we remove the system and ask whether users reproduce it unaided, testing metalearning rather than one-time accuracy. We compare two AI-assisted treatments, Exploratory (guided reasoning) and Summary (synthesized verdicts), against two baselines, unrelated news and Self-search. The treatments yield the strongest immediate accuracy and confidence gains but do not survive the fading test: on unseen claims they perform no better than news, while Self-search, with no procedure to fade, retains a large advantage. This reveals a factchecking-metalearning tradeoff: conditions that most improve immediate accuracy are least likely to produce metalearning, cautioning against treating AI-delivered verdicts as a source of durable literacy gains.

cs.HC

HyperDet: 3D Object Detection with Hyper 4D Radar Point Clouds

How far can 3D object detection go using 4D radar alone? Despite offering weather-robust and velocity- aware sensing for autonomous perception, modern 4D radar still yields sparse, noisy, and unstable point clouds, limiting radar-only 3D detection. We present HyperDet, a detector- agnostic input enhancement pipeline that constructs task- aware hyper 4D radar point clouds by combining measured observations with completed foreground geometry. HyperDet first refines short-window surround-view radar observations through spatio-temporal accumulation and cross-sensor val- idation, while Doppler-guided motion compensation reduces dynamic object trails when motion can be estimated reliably. It then performs foreground generative enhancement using LiDAR-guided pseudo-radar supervision available only during training, enriching object geometry while preserving measured radar background and radar-native attributes. During detec- tor training, radar-aware object-level augmentation maintains Doppler consistency under geometric relocation. At inference, HyperDet requires radar input alone and can be directly paired with standard 3D detectors. Experiments on two public surround-view 4D radar datasets demonstrate consistent im- provements over matched temporal accumulation across stan- dard 3D detectors, validating input-level radar enhancement as an effective approach to radar-only 3D detection.

cs.RO

VeRA: Renewing Reasoning Benchmarks with Executable Specifications

Reasoning benchmarks need renewal along two axes: freshness and headroom. VeRA makes both executable and auditable by turning each item into a task family: a natural-language template, an input generator, and a deterministic answer program. VeRA-E draws fresh instances within a family; VeRA-H modifies the family toward harder tasks; and VeRA-H Pro selects one judge-ranked candidate from up to five validated proposals per seed. Execution checks, seed anchoring, answer discrimination, and independent human solving validate specifications and items. Accepted programs generate further labeled instances through local computation. Across 16 models, AIME-2024 accuracy decreases from 84.46% on seeds to 70.25% on VeRA-E variants, exposing a gap between fixed-item success and fresh-instance robustness. On AIME-2024-II, the human-audited VeRA-H Pro release lowers accuracy from 84.91% to 58.57%. Across the three hardening sources, H Pro has lower mean accuracy than H. On AMO-Bench, both releases average higher accuracy than the seeds under the evaluated budget. Initial auditing accepts 75.4% of hardened candidates; targeted repair raises usable yield to 95.1%. Executable families thus support repeatable benchmark renewal, with validation improving task quality and selection shaping the delivered challenge.

cs.AI

Flexoelectricity-driven softening of bend elasticity leads to spontaneous chiral symmetry breaking in a polar fluid

The origin of the recently observed spontaneous chiral symmetry breaking in polar fluids composed of achiral molecules is an unsolved problem, raising fundamental questions about how heliconical structures emerge in such systems. Here, we investigate the pretransitional fluctuations leading to the formation of the spontaneously chiral twist-bend ferroelectric nematic phase using dielectric spectroscopy, light scattering, and small-angle X-ray scattering. We observe simultaneous softening of the bend elastic constant and the emergence of a collective dielectric mode on approaching the transition. By developing a theoretical model, we show that these phenomena are signatures of a flexoelectricity-driven transition arising from the coupling between electric polarization and bend deformation.

cond-mat.soft

Spectral boundaries of deterministic matrices deformed by rotationally invariant random non-Hermitian ensembles

One of the great miracles of random matrix theory is that, in the $N \to \infty$ limit, many otherwise intractable matrix problems with horrendously complicated finite-$N$ expressions admit remarkably simple and elegant asymptotic solutions. In this paper, we illustrate this phenomenon in the context of spectral boundaries (or spectral edges) for deformed random matrices. Specifically, we consider matrices of the form $\mathbf{A} + \mathbf{B}$, where $\mathbf{A}$ is a deterministic $N\times N$ matrix (not necessarily Hermitian) and $\mathbf{B}$ is a rotationally invariant random matrix. In the large-$N$ limit, we show that the complex eigenvalue distribution of $\mathbf{A} + \mathbf{B}$ satisfies remarkably simple boundary equations that depend on the $\mathcal{R}_1$ and $\mathcal{R}_2$ transforms of $\mathbf{B}$. We illustrate our results on several explicit random matrix ensembles and support them with numerical simulations.

cond-mat.dis-nn

Global causality constraints in rotating scalar-tensor spacetimes

Modified gravity is often formulated as an effective field theory (EFT), where higher-order corrections parametrize departures from General Relativity. We argue that such corrections should be constrained by the global causal structure of curved spacetime, in addition to the usual flat-space requirements such as positivity and unitarity. We propose that within the domain of validity of the EFT, the onset of closed timelike curves should not happen in a parametrically more accessible region than in the corresponding GR background. We test this diagnostic in the quadratic k-essence sector of scalar-tensor gravity. For stationary and axisymmetric spacetimes, the invariant test for closed axial orbits is the sign of the azimuthal component of the metric \(g_{φφ}\). We supplement this test by requiring a local time function in the space of Killing vectors. We apply these conditions to quadratic k-essence on Kerr--(A)dS backgrounds, with and without scalar charge. The zero-charge branch is exact Kerr--(A)dS, and we treat the charged branch perturbatively in scalar charge and in Hartle--Thorne slow rotation. Expanding for small spin \(χ=a/(GM)\ll1\), frame dragging begins at \(\mathcal O(χ)\), while the quadrupolar backreaction relevant for circular closed timelike curves enters at second order in both rotation and charge. We find that, in the truncation used here, any occurrence of \(g_{φφ}<0\) also lies outside EFT control. A higher-order calculation or a fully nonlinear treatment is therefore needed. Finally, we discuss how quasinormal modes and black-hole echoes could probe such causal structure.

gr-qc

Energy dependence of cross sections in proton-proton and antiproton-proton collisions

Energy dependence of global scattering parameters, mostly of total cross section, is studied for proton-proton and antiproton-proton collisions. Results are presented for physical analysis updated with taken into account the recent data from accelerator experiments as well as from cosmic ray measurements. The analytic parameterizations suggested within Axiomatic Quantum Field Theory (AQFT) provide the quantitative description of energy dependence of global scattering parameters for rather wide energy range. Detailed scan on low boundary of the fitting range for energy dependence of global scattering parameters allows the observation of the onsets for regions in which Pomeranchuk theorem and / or Froissart-Martin one is valid. It is obtained that global scattering parameters show the behavior corresponded to any formulations of Pomeranchuk theorem and closed to (modified) Froissart-Martin limit in functional sense in multi-TeV energy region. Bosonic condensation is considered as one of the possible dynamical mechanisms which would be provide the total cross section approaches to (modified) Froissart-Martin limit at quantitative level but not functionally only.

hep-ph

When can we trust untrusted monitoring? A safety case sketch across collusion strategies

AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitoring -- using one untrusted model to oversee another -- is one approach to reducing risk. Justifying the safety of an untrusted monitoring deployment is challenging because developers cannot safely deploy a misaligned model to test their protocol directly. In this paper, we develop upon existing methods for rigorously demonstrating safety based on pre-deployment testing. We relax assumptions that previous AI control research made about the collusion strategies a misaligned AI might use to subvert untrusted monitoring. We develop a taxonomy covering passive self-recognition, causal collusion (hiding pre-shared signals), acausal collusion (hiding signals via Schelling points), and combined strategies. We create a safety case sketch to clearly present our argument, explicitly state our assumptions, and highlight unsolved challenges. We identify conditions under which passive self-recognition could be a more effective collusion strategy than those studied previously. Our work builds towards more robust evaluations of untrusted monitoring.

cs.AI

Burgess-type volume dependent bounds for character sums over $\mathbb{F}_{p^n}$

We establish a Burgess-type bound for short multiplicative character sums over finite fields $\mathbb{F}_{p^n}$. Let \[ B=\left\{\sum_{i=1}^{n}x_iω_i: N_i+1\le x_i\le N_i+H_i,1\le i\le n\right\}\subseteq\mathbb{F}_{p^n}, \] where $1\le H_i\le p$ for all $1\le i\le n$, and the side lengths satisfy $H_1\le H_2\le\cdots\le H_n.$ We prove that if the side lengths satisfy certain lower bounds in terms of the two largest side lengths, then a nontrivial cancellation occurs in the character sum over the boxes. This generalizes the work of Gabdullin \cite{GB} in dimensions $n=2,3$ to arbitrary dimension. This also generalizes the character sum estimate of Konyagin \cite{Kon} where each of the side lengths of the boxes are greater than $p^{1/4}$. The proof combines techniques from the geometry of numbers, multiplicative energy estimates, and Katz's bounds for multiplicative character sums.

math.NT

Excited-state quantum phase transitions and chaos in a three-level Lipkin model

Excited-state quantum phase transitions (ESQPTs) have been extensively studied in two-level models, but their characterization remains challenging in systems displaying mixed regular and chaotic dynamics. In this work, we investigate ESQPTs within the three-level Lipkin-Meshkov-Glick model, where an enlarged Hilbert space and multiple separatrices give rise to rich spectral structures strongly influenced by chaos. To investigate the different dynamical regions, we have calculated Poincaré sections and Peres lattices. In addition, by combining chaos-sensitive measures with standard ESQPT diagnostics, we provide a static analysis of ESQPT signatures in this model and establish a robust framework for future studies of its dynamical behavior. The degree of chaos and the Kullback-Leibler divergence are found to be very effective chaos-sensitive measures, which are complementary to ESQPT diagnostics such as the mean field limit and the participation ratio. Hence we provide a standard framework to work with ESQPTs in chaotic three-level systems.

quant-ph

SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits this convergence by aligning frozen pretrained vision and language models with lightweight alignment layers, but typically relies on contrastive losses and millions of paired samples. In this work, we ask whether meaningful alignment can be achieved with substantially less supervision. We introduce a semi-supervised setting in which pretrained unimodal encoders are aligned using a small number of image-text pairs together with large amounts of unpaired data. To address this challenge, we propose SOTAlign, a two-stage framework that first recovers a coarse shared geometry from limited paired data using a linear teacher, and then refines the alignment on unpaired samples via an optimal-transport-based divergence that transfers relational structure without overconstraining the target space. SOTAlign effectively leverages unpaired images and text, learning robust joint embeddings across datasets and encoder pairs, and significantly outperforming supervised and semi-supervised baselines. Code is available at https://github.com/ExplainableML/SOTAlign.

cs.LG

Small hosts, big appetites: unveiling rapid and early low-mass black hole growth in cosmological zoom-in simulations of dwarf galaxies

Dwarf galaxies are ideal laboratories to probe the interplay between galaxy formation and the growth of black holes (BHs) in the early Universe. Mounting observational evidence reveals the presence of BHs in low-mass galaxies across cosmic time, with $\textit{JWST}$ uncovering a likely population of $\textit{overmassive}$ BHs at $2 \lesssim z \lesssim 11$. Simulations struggle to reproduce this high-redshift regime, motivating revisions to models of BH accretion and feedback from active galactic nuclei (AGN). To address this, we present high-resolution cosmological zoom-in simulations of a dwarf galaxy based on FABLE physics, introducing novel sink-based BH accretion models and relaxing the fiducial assumption of strong supernova feedback. BHs accrete more efficiently in the sink-based runs compared to the `traditional' Bondi-based counterparts, with AGN feedback leading to early, rapid quenching maintained by fast, hot and metal-enriched outflows. These outflows pollute the outer circumgalactic medium, yielding flat metallicity gradients down to $z=0$. We further assess the performance of two widely used virial estimators and find significant departures from the true dynamical mass, especially during the high-redshift dwarf assembly. Since our galaxy is dark-matter-dominated at all times and radii, BH growth, tied to the baryon cycle, shows no clear correlation with global dynamical properties. Efficient AGN feedback is produced by overmassive BHs relative to extrapolated local $M_\bullet - M_\star$ relations, raising the possibility that dormant, overmassive BHs in local quenched dwarfs and those probed by $\textit{JWST}$ may reflect a common mode of early and rapid BH growth in low-mass galaxies.

astro-ph.GA

MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine

Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remains: top-performing systems are either closed-source or computationally prohibitive, precluding the on-premises deployment required for patient privacy and PHI compliance. We introduce MEDGPT-OSS, an open-weight, 20B-parameter generalist vision-language model designed to facilitate open research in clinical AI. Rather than relying on architectural complexity, MEDGPT-OSS pairs the GPT-oss language backbone with a visual front-end via a optimized, three-stage training curriculum. By progressively domain-adapting these modules through rigorous data curation and long-context multimodal alignment, we demonstrate that a 20B model can bridge the capacity gap. It successfully outperforms larger open medical models on out-of-distribution (OOD) multimodal reasoning and complex text-only clinical tasks. By unifying diverse modalities under a single instruction-following interface, MEDGPT-OSS maintains a parameter-efficient footprint fully compatible with commodity GPUs. We release the complete training recipe, open-weight checkpoints, and a rigorous evaluation harness to serve as a verifiable foundation for privacy-preserving, institution-specific clinical AI research.

cs.CL

VectorMaton: Efficient Vector Search with Pattern Constraints via an Enhanced Suffix Automaton

Approximate nearest neighbor search (ANNS) has become a cornerstone in modern vector database systems. Given a query vector, ANNS retrieves the closest vectors from a set of base vectors. In real-world applications, vectors are often accompanied by additional information, such as sequences or structured attributes, motivating the need for fine-grained vector search with constraints on this auxiliary data. Existing methods support attribute-based filtering or range-based filtering on categorical and numerical attributes, but they do not support pattern predicates over sequence attributes. In relational databases, predicates such as LIKE and CONTAINS are fundamental operators for filtering records based on substring patterns. As vector databases increasingly adopt SQL-style query interfaces, enabling pattern predicates over sequence attributes (e.g., texts and biological sequences) alongside vector similarity search becomes essential. In this paper, we formulate a novel problem: given a set of vectors each associated with a sequence, retrieve the nearest vectors whose sequences contain a given query pattern. To address this challenge, we propose VectorMaton, an automaton-based index that integrates pattern filtering with efficient vector search, while maintaining an index size comparable to the dataset size. Extensive experiments on real-world datasets demonstrate that VectorMaton consistently outperforms all baselines, achieving up to 10x higher query throughput at the same accuracy and up to 18x reduction in index size.

cs.DB