arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,063 records · Page 59Linked to original sources

SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction

Reconstructing Computer-Aided Design (CAD) modeling sequences from images is crucial for preserving design intent and supporting parametric editing. However, existing methods typically generate full CAD sequences holistically, overlooking the iterative, feedback-driven nature of human design workflows. We address this limitation by introducing the rich stepwise visual supervision: at each modeling step, the system observes the target's orthographic projections, the projections of the incrementally constructed model, and the active sketch, enabling informed action selection. To effectively leverage this on-the-fly feedback, we propose SOV-CAD, a framework that formulates CAD reconstruction as a sequential decision-making task and employs offline reinforcement learning with a Decision Transformer architecture. This design incorporates continuous visual feedback guided by geometric alignment rewards, resulting in a more accurate and human-like modeling process. Extensive experiments show that SOV-CAD surpasses state-of-the-art methods in CAD sequence reconstruction while exhibiting strong data efficiency. Code of SOV-CAD is available at: https://github.com/LukePhong/SOV-CAD

cs.CV↗

Allocation condensation under asymptotically complete mixing

We study a fixed network on which new mass is allocated in proportion to a power of each node's stock and then redistributed. Transport retains a fraction of both old and new mass at each node and divides the rest according to fixed background shares. We show that a small amount of retention can sustain concentration on networks where complete redistribution would leave every node's share of new mass vanishing as the network grows. A node receiving most new mass may retain little relative to the network total, yet much more than redistribution alone would supply. For a power above one, the allocation rule amplifies this relative stock advantage at the next step. We identify retention rates tending to zero with network size that allow widely spread and concentrated stationary allocations to coexist on the same network. One of them stays close to the allocation under complete redistribution. Each node can also dominate a separate concentrated allocation, receiving a share of new mass tending to one. The stocks supporting these allocations attract nearby trajectories, although every stationary stock approaches the same background, with even its largest share tending to zero. These allocations continue to coexist under specified changes in transport that preserve the background. Redistribution can therefore spread the stock widely without spreading new allocations in the same way.

physics.soc-ph↗

Stress calculation in linear scaling DFT: convergence and dynamics

We present the approach needed to calculate stress within density functional theory (DFT) using a localised orbital basis, both for exact diagonalisation and linear scaling approaches, and demonstrate our implementation within the large scale DFT code Conquest. For the linear scaling approach, we test the rate of convergence of stress with density matrix range, and compare it to the convergence of energy and forces for different materials with a range of band gaps. We show that excellent convergence is found for modest cutoffs (below 0.1GPa error for range of 20a0), and show that large-scale isothermal-isobaric molecular dynamics is stable and accurate, showing negligible drift over picoseconds of dynamics.

cond-mat.mtrl-sci↗

On the stability of proximal operators in Wasserstein spaces under different notions of convexity

The proximal operator is a fundamental tool in variational analysis and optimization. In the setting of a Hilbert space, given a proper, lower semicontinuous convex functional, its proximal operator is non-expansive, that is, 1-Lipschitz continuous. In the Wasserstein setting, the contraction properties of this operator have been investigated from different perspectives by Carlen and Craig and by Adve and Mészáros, among others, and are not completely understood. In this paper, we study the stability properties of proximal maps, with a particular focus on non-expansivity, under various notions of convexity of the functional that can be considered in the Wasserstein space.

math.OC↗

glmSTARMA -- An R-Package for fitting autoregressive spatio-temporal models following generalized linear models

The R package glmSTARMA implements autoregressive models for spatio-temporal data at fixed locations, with time-invariant spatial dependency structure. We rely on generalized linear models methodology and unify several approaches for the analysis of spatial count time series. Such models allow the (conditional) mean of the response to depend on past observations, lagged (conditional) expectations, and covariates. The response can be a continuous or a discrete random variable. Additionally, the package develops inference for double generalized linear models, allowing the dispersion parameter(s) of the marginal distributions to be modeled similarly to the mean process. This is a new capability which introduces, for example, spatio-temporal volatility models, such as space-time GARCH processes, and count time series models with spatio-temporal overdispersion and underdispersion. We provide functions for model estimation, simulation, inference, and prediction. Its use is illustrated by data examples.

stat.CO↗

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents

Language-conditioned manipulation requires both precise contact-rich control and robust reasoning over language, scenes, and long horizons. End-to-end Vision-Language-Action (VLA) models provide strong local visuomotor skills, but they are trained on in-distribution task trajectories and often fail under deployment perturbations such as semantic retargeting, goal re-binding, spatial-layout shifts, and unstable local contacts. LLM coding agents provide complementary semantic and compositional reasoning, but purely analytic primitives struggle with irregular grasping, constrained placement, and articulated-object interaction. We present Harness VLA, a memory-augmented agentic framework that exposes a frozen VLA as a retryable contact-rich primitive and composes it with a small fixed library of analytic primitives for grounding, staging, transport, navigation, and release. Rather than expanding the skill library, the harness learns the operating range of these fixed primitives from task-specific execution traces, global success rules, and failure models. By lifting semantic re-grounding, non-contact execution, and VLA re-staging to the planner while reserving the frozen VLA for local contact-rich phases, Harness VLA extends pretrained VLAs beyond their original trajectory distribution without finetuning. Across perturbed tabletop, household kitchen, and clean-to-randomized bimanual manipulation, Harness VLA improves over the strongest relevant baselines by 38.6 and 25.4 percentage points on LIBERO-Pro and RoboCasa365, respectively, and reaches 58.4% on RoboTwin C2R. Real-world demonstrations on dual-Franka robots further show target redirection, grasp recovery, and new task compositions with the same frozen VLA. Code is available at https://github.com/RLinf/RPent.

cs.RO↗

Locally Approximating the Top Eigenvector of Bounded Entry Matrices

We provide a local computation algorithm to approximate the top eigenvector $x \in \mathbb{R}^n$ of a symmetric matrix $A \in \mathbb{R}^{n \times n}$ with entries between $-1$ and $1$, building on the work of Swartworth and Woodruff [SODA 25] who show how to approximate the eigenvalues up to additive-$\varepsilon n$ error using $\tilde{O}(1/\varepsilon^4)$ queries. Our local computation algorithm has a preprocessing complexity of $\tilde{O}(1/\varepsilon^4)$ and per-coordinate query complexity of $\tilde{O}(1/\varepsilon^2)$ for an additive-$\varepsilon n$ approximation whenever {$|λ_{\min}(A)| = O(λ_{\max}(A))$. When $λ_{\min}(A)$ greatly exceeds $λ_{\max}(A)$, our complexity degrades to at most $\tilde{O}(1/\varepsilon^{6.\overline{6}})$ in preprocessing and $\tilde{O}(1/\varepsilon^{3.\overline{3}})$ per query. Furthermore, we show a lower bound of $Ω(n/\varepsilon^2)$ on the total number of queries needed to output an approximately top eigenvector (implying that the per-coordinate query complexity of $Ω(1/\varepsilon^2)$ is necessary). As an application, we use our algorithm to provide local computation algorithms for the sparsest-cut and max-cut problems in the dense graph model of Goldreich, Goldwasser, Ron [JACM 98]. By accessing the top eigenvectors (of an approximate normalized adjacency), we implement local versions of Cheeger's inequality and Trevisan's algorithm [SICOMP 12] to obtain "square-root-opt" approximations in polynomial time (as opposed to exponential-in-$\text{poly}(1/\varepsilon)$ time which is incurred in Goldreich, Goldwasser, Ron.

cs.DS↗

The queer Hero versus the Fool bias of the queer trait: An archetypometric analysis of the collective portrayal of queerness in fictional stories

Visibility in media is pivotal for identity development and for broadening societal views of gender and sexuality. Queer representation has increased in recent years, yet damaging stereotypes and tropes persist. Here, we focus on queer portrayal and its perception by audiences in fictional stories (television, film, and literature) by studying characters by their quantified archetypes which are operationalizations of common conceptions such as Hero, Diva, and Outcast. We use the archetypometrics and Fandom's LGBTQIA+ datasets to study samples of fictional characters along the trait differential spanning straight to queer. We find, quantify, and explain a seeming paradox. The characters with the highest queer score present positive primary archetypes and are typically Heroes rather than Fools, Angels rather than Demons, and Adventurers rather than Traditionalists. But evaluation across many stories for the straight-queer trait itself reveals a strong collective-writing bias towards Fool (away from Hero) and no meaningful loading for the other two dimensions. Our analysis offers a population-scale view of the complexities of queer portrayal, while also pointing to risks in blindly training on many-authored story corpora.

cs.CY↗

Inhomogeneous Strichartz estimates on manifolds with nonpositive curvature and applications

We prove lossless inhomogeneous Strichartz estimates for solutions to the Schrödinger equation on compact manifolds with nonpositive curvature over frequency-dependent time intervals of length $\log λ\cdot λ^{-1} $. As applications, we improve upon the Sobolev norm growth bounds for the cubic NLS established by Planchon, Tzvetkov and Visciglia on 3-dimensional compact manifolds when the manifold also has nonpositive curvature and extend the lossless homogeneous Strichartz estimates on logarithmic time intervals established by the first author and Sogge to Schrödinger operators with critically singular potentials.

math.AP↗

Implicit Behavior Coordination from Sub-Task Demonstrations by Exploiting Overlap-Induced Multimodality

Long-horizon robotic rearrangement is commonly formulated as a skill-sequencing problem, where distinct behaviors are explicitly represented and coordinated by a planner or high-level policy. We investigate whether such explicit behavior identities and sequencing interfaces are necessary at all. We introduce implicit behavior coordination from sub-task demonstrations, where separately collected behaviors are coordinated without behavior identity labels, complete-task demonstrations, or task-ordering supervision. Our key observation is that overlap between sub-task demonstrations induces multimodal action distributions that need not be resolved through explicit behavior partitioning. Instead, this overlap-induced multimodality can be exploited as a coordination resource. We instantiate this idea with a shared Flow Matching policy that preserves multiple action modes and critic-guided in-sample planning that propagates task value across demonstrations and selects task-relevant modes. Experiments in Habitat and on a real robot show that implicit behavior coordination remains effective under reduced cross-behavior overlap, larger behavior mixtures, longer horizons, and execution failures, supporting the idea that long-horizon coordination can emerge directly from sub-task demonstrations without explicitly recovering or sequencing behavior identities.

cs.RO↗

Macroscopic Zero-Mode Manifold Isolated by Quantum Chaos

Chaotic many-body spectra are expected to densely fill their energy window. We show that constrained spin chains with chiral symmetry evade this expectation by hosting an exponentially large manifold of symmetry-protected exact zero modes separated from the surrounding spectrum by a sharp gap at zero energy. The gap is generated by chaotic level repulsion, with width set by the number of zero modes times the mean level spacing. We verify this mechanism in an East-West kinetically constrained chain, develop a minimal random-matrix description, and show how the gap can be detected through linear-response spectroscopy.

cond-mat.str-el↗

Learning to Navigate with Minimal Parameters: Decomposing Visual Navigation Through Closed-Form Geometric Interfaces

Visual navigation policies have grown to hundreds of millions of parameters trained on billions of frames, with geometry, mapping, and control learned implicitly. We propose a decomposed point-goal navigation system in which operations with known closed-form structure, such as projective geometry, occupancy, and coordinate transforms, are computed analytically and serve as interfaces between three small learned modules: an egress predictor that grounds the episode goal as a local subgoal in the current view, a navigation predictor that estimates a goal-conditioned posterior over where trajectories travel, and an endpoint-pinned residual diffusion generator that samples trajectory shapes from this posterior. Only 0.58M out of 23M parameters are trained, on 44k frames, in under one GPU-hour. Across 6060 point-goal episodes in 60 environments, the system attains competitive success rates with the lowest collision rate among evaluated methods. We further show that under this decomposition, the frozen image encoder can be replaced by a 0.54M MobileNetV2 at a -2.0 SR cost, bringing the full system under 1.2M parameters. It also transfers to no-goal exploration by retraining only the 123k-parameter egress head, and its failure modes under sensor corruption are transparent and analytically correctable. We deploy and evaluate the system zero-shot on a low-cost UGV, running navigation and localization on a Jetson Orin Nano in real-time.

cs.RO↗

Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents

Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict. A critic reads each noisy observation and passes only the usable content to the ledger, discarding the rest, so the planner works from a compact context throughout the task. Each run records the state, action, and verdict at every step, and supervised fine-tuning and decision-level reinforcement learning on these trajectories further improve the planner. Omni-Decision achieves state-of-the-art accuracy of 81.4% on OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.

cs.AI↗

A hybrid analytical-PINN model for subsurface simulation of geothermal heat exchangers in heterogeneous underground

Accurate and efficient prediction of subsurface temperature fields is essential for the design and operation of borehole heat exchanger (BHE) systems. Here we develop a parametric hybrid analytical and physics-informed neural network (PINN) framework for long-term multi-BHE simulations in heterogeneous underground. The method analytically extracts the singular line source response and enables the effective training of neural correction associated with subsurface heterogeneity. An explicit parametrization of the thermal conductivity allows physics-informed learning of a single feedforward neural network to generalize across different subsurface conditions. By formulating the correction in borehole-centered relative coordinates, the learned correction can be reused as a universal corrector through spatial and temporal superposition principles. Numerical experiments based on the infinite line source (ILS), finite line source (FLS) and moving finite line source (MFLS) models show that the hybrid method outperforms analytical approximations with stable accuracy over long simulation horizons and achieves orders-of-magnitude speedups over traditional solvers. The proposed framework therefore combines the efficiency of analytical models with the ability of numerical methods to capture heterogeneous subsurface physics, providing a fast and accurate approach for repeated long-term simulation of multi-BHE systems.

cs.LG↗

A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation

The key-value (KV) cache is the dominant memory bottleneck in long-context language model inference. Existing compression methods apply low-rank factorization or quantization independently, without jointly allocating rank and precision under a shared storage budget. We introduce JoLT, a training-free compressor that treats grouped prefill caches as fourth-order tensors and applies partial Tucker decomposition along the token and feature modes, the two axes that carry low-rank structure, while leaving the head and layer modes intact. A rotated low-bit quantizer captures the truncation residual, and a single Lagrangian dual allocates per-group Tucker ranks and residual bit-widths under a global byte constraint. FlashJoLT replaces the exact token-mode SVD with a randomized approximation that matches JoLT within the free zone at a fraction of the compression cost, and a fused Triton decode kernel evaluates attention directly over the stored factors without materializing dense KV tensors. Across five models from four architecture families, covering multi-head attention, grouped-query attention, and mixture-of-experts architecture, JoLT achieves 2 - 3x compression with less than 0.2% perplexity degradation, without retraining. On RULER at 64K context with LLaMA-3.1-8B, retrieval accuracy remains near-lossless through 3x and declines by only 0.90 and 2.40pp at 4x and 5x, respectively. JoLT demonstrates that tensor-aware low-rank decomposition and quantized residuals, unified under a single storage budget, achieve near-lossless KV-cache compression across diverse model architectures without retraining.

cs.LG↗

Collective-State Preparation in a Subwavelength Triangular Trimer Using SUPER Excitation

The Swing-UP of quantum EmitteR population (SUPER) scheme has recently been proposed as a deterministic method for the preparation of collective radiative states in two strongly dipole-coupled quantum emitters (Phys. Rev. Res. \textbf{8}, 013179 (2026)). Here, we apply this approach to an equilateral subwavelength triangular trimer of dipole-coupled two-level quantum emitters (QEs), loosely inspired by biological light-harvesting ring geometries, and demonstrate that tailored, time-overlapping, red-detuned ultrashort SUPER pulses can selectively prepare collective states in a $C_3$-symmetric system, including energetically degenerate eigenstates that are resolved through site-dependent optical phases. We find that both the state selectivity and the preparation efficiency depend strongly on the inter-emitter spacing. In particular, at deep-subwavelength separations, the symmetric collective state can be deterministically prepared with near-unity efficiency (approximately $94\%$), whereas the inversion efficiency and state selectivity are significantly lower at larger inter-emitter separations. Furthermore, this state preparation technique inherits a certain degree of robustness against reasonable static position imperfections and on-site frequency inhomogeneities of the individual QEs. Our results demonstrate that deep-subwavelength triangular trimers and, more broadly, highly compact ring geometries are excellent candidates for the deterministic preparation of collective radiative states via SUPER excitation. These predictions could be realized with solid-state emitters and molecules. Our findings offer a route toward the direct probing of the `pure' electromagnetic layer of interaction in biological and bio-inspired synthetic nanophotonic ring configurations, with possible relevance in photonics, quantum information processing, and metrology.

physics.optics↗

Simulating the convection in red super-giant stars: wobbling jets in common envelope evolution

We use our newly constructed three-dimensional red supergiant (RSG) stellar model, which also mimics nuclear energy production and photospheric emission, to calculate the stochastic component of the angular momentum of the mass that a companion spiraling within the RSG's envelope accretes during common envelope evolution (CEE). The accreted mass has a fixed-direction angular-momentum component arising from the density gradient in the RSG envelope and orbital motion. The angular momentum component with a stochastically varying direction results from vigorous envelope convection. We do not include the companion's influence on the RSG envelope during the CEE and consider an undisturbed, non-rotating RSG stellar model. We find that the fluctuating angular momentum amplitude can be several times the fixed-axis angular momentum. The total specific angular momentum of the accreted mass easily forms intermittent accretion disks around neutron stars and black holes, but it is only marginally sufficient, or not at all, to form accretion disks around main-sequence stellar companions. The intermittent accretion disks we expect to form will launch wobbling jets with varying axes. We discuss aspects of wobbling jets in the CEE and the grazing envelope evolution (GEE), which might precede the CEE or replace it altogether. Studies have claimed that jets are a crucial ingredient in many cases of CEE, and the standard CEE should include jets that the companion launches, before (like the GEE), during, and/or at the exit from the CEE. Our study supports this claim and emphasizes the importance of wobbling jets.

astro-ph.SR↗

DreamSat-Pose: Spacecraft Pose Estimation from Single-View 3D Reconstructions and Learned 2D-3D Feature Matching

6-DoF pose estimation is a critical task in autonomous rendezvous and proximity operations. In the case of an unknown target, this task becomes challenging as it shall be paired with the reconstruction of the target shape model. In this article, we propose a novel framework for single-shot shape and pose estimation of unknown spacecraft objects. Given a single image, we first reconstruct a 3D shape model of the target, then estimate the relative six-degrees-of-freedom pose by learning dense 2D-3D correspondences. The image features are extracted using a frozen DINOv3 vision transformer, while the geometric features are computed from the reconstructed point cloud using a trainable dynamic graph convolutional neural network encoder. A dual-stream transformer matcher refines descriptors through alternating self- and cross-attention, producing soft correspondences that are passed to a Perspective-$n$-Point solver for pose recovery. We evaluate the method on the SPE3R dataset and consider FoundationPose as a representative baseline for current state-of-the-art capabilities. Results show reliable pose estimates achieving 0.157 degrees mean pointing error using only a single image and reconstructed geometry, demonstrating strong generalization to unseen spacecraft.

cs.CV↗