arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

DETECT: Real-Time Identification of Transients with DESI Spectroscopic Redshift

Wide-field surveys now report thousands of transients per month, while spectra are obtained for only a few per cent of them. We present DETECT (DESI--Transient Event Cross-matching Tool), a pipeline that turns each Transient Name Server (TNS) alert into a distance-informed candidate within an hour by placing it in the context of the Dark Energy Spectroscopic Instrument (DESI) archive. DETECT cross-matches every new report against the 22 million spectra of DESI EDR and DR1 through a HEALPix index, associates the transient with a host galaxy by a directional-light-radius rule whose morphology-dependent thresholds are calibrated on 4,474 supernovae with known redshifts (92.5% completeness, 0.4% wrong hosts), and passes ambiguous cases to a reviewer through a web interface. With a validated host redshift it derives the peak absolute magnitude, the projected offset and basic host properties of each event, and ranks it for spectroscopic follow-up. Applied retrospectively to ~1.2x10^5 TNS reports from 2020--2024, DETECT recovers a spectroscopic host for 15% of them, from these we build a Gold Sample of ~5,400 well-sampled light curves whose luminosity distributions by class reproduce those of untargeted surveys. Because the absolute magnitude is known at discovery, intrinsically overluminous events stand out immediately: in the prospective 2025 run this is how the lensed superluminous supernova SN~2025wny was flagged within hours of its report, and how faint, fast-declining kilonova candidates were screened against model grids and archival photometry. DETECT shows how an archival spectroscopic survey can make host redshifts a routine part of transient triage ahead of the Rubin Observatory's LSST.

astro-ph.HE↗

RoboRecover: Benchmarking Robot Policy Recovery under Execution Deviations

Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. During closed-loop execution, actions and contacts can alter object relations and task progress, producing off-nominal intermediate states that need recovery. Recovery requires a policy to infer how task progress has changed, correct the relevant relations, and continue the original goal. We introduce RoboRecover, a benchmark for robot policy recovery under execution deviations. RoboRecover selects deviation states from trajectories, reconstructs them by replaying action prefixes, and evaluates policies on the original task. RoboRecover contains 2,000 scenarios across RoboTwin and LIBERO, with 1,000 scenarios and a fixed 800/200 train/test split on each platform. Results show that initial-state performance does not determine recovery performance and policies exhibit different recovery strengths across scenarios. Using its training split, RoboRecover further supports study on recovery interventions. RoboRecover establishes recovery from execution-induced intermediate states as a distinct dimension of robot policy evaluation.

cs.RO↗

Constraining the Quadratic-mode Amplitude Coupling in GW250114

Detecting quadratic quasi-normal modes in black hole ringdowns would provide evidence for nonlinear gravitational dynamics, while measuring their properties would enable tests of the corresponding predictions of general relativity. Specifically, second-order black hole perturbation theory predicts that their amplitudes scale with the product of the amplitudes of their parent linear modes, with a coupling coefficient that depends on the spin of the remnant black hole. This mode-specific coupling coefficient has not yet been directly measured from gravitational wave data. Here, we use Bayesian inference on the GW250114 ringdown, modeled with the $220$, $221$, and $220\times220$ modes, to infer the amplitude-coupling coefficient of the $220\times220$ mode. The inferred coefficient is consistent with predictions from numerical relativity fits and second-order perturbation theory; no comparison shows a deviation exceeding $1.3 σ$. Although the current data cannot distinguish among these spin-dependent predictions, this measurement establishes a basis for more stringent observational tests of nonlinear gravitational dynamics in black-hole ringdowns.

gr-qc↗

Convergence analysis of a numerical scheme for the Cahn-Hilliard-Navier-Stokes system with dynamical boundary condition and its application to moving contact line problem

A finite difference numerical scheme is proposed and analyzed for the Cahn-Hilliard-Navier-Stokes system, combined with a dynamical boundary condition. Such a physical system has potential applications in the moving contact line problem. The boundary profile is governed by a lower-dimensional energy potential, coupled with a non-homogeneous boundary condition for the phase variable. In the numerical design, a convex-splitting approach is applied to the chemical potential in both bulk and surface levels, which leads to a highly coupled nonlinear system. A semi-implicit discretization is taken in the nonlinear fluid convection, as well as the coupled terms between the fluid motion and phase variable evolution. A careful finite difference approximation and convexity analysis reveals that such a numerical system could be represented as a non-symmetric and monotone mapping associated with the fluid convection. In turn, the unique solvability is valid based on the monotonicity argument. The total energy stability analysis is obtained through a careful summation-by-part calculation. In particular, an optimal rate convergence analysis is theoretically established in this work. The discrete mass conservation of the exact solution is required to preserve the mean-zero property of the error function, so that the associated discrete H_h^{-1} norm is well-defined. A combination of the Fourier projection and an auxiliary function is applied to overcome this difficulty. Furthermore, an approach of rough and refined error estimates concludes the desired convergence result. Some numerical results are presented in this article, which demonstrate the robustness of the proposed numerical scheme. In our knowledge, this work provides a theoretical proof of convergence analysis and error estimate for a numerical scheme to the moving contact line problem, for the first time in the literature.

math.NA↗

ActGaze: Learning Action-Grounded Gaze through Counterfactual Visual Interventions for High-Precision Manipulation

Current Vision-Language-Action (VLA) models often struggle with high-precision robotic manipulation. We attribute this limitation primarily to their visual attention being dispersed across task-irrelevant regions. To address this issue, we propose ActGaze, a training approach that guides VLA policies to gaze on task-relevant regions, much like humans gaze on critical visual cues while executing precise movements. Unlike prior methods that rely on external labels for gaze supervision, ActGaze derives spatial supervision directly from the VLA's own action objective by using counterfactual visual interventions to identify regions that are critical for action prediction. Extensive real-robot experiments on four high-precision robotic manipulation tasks demonstrate that ActGaze induces more focused visual attention on task-relevant regions and consistently outperforms the base VLA policy and other visual-grounding approaches.

cs.RO↗

MoVISA: Multi-Token Reasoning for Video Object Segmentation

Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.

cs.CV↗

The Tambara-Yamagami calculus and finite semigroups

Tambara-Yamagami categories are among the simplest fusion categories beyond the pointed case: one adds a single noninvertible simple object to a finite group of invertible ones. We ask what remains of their (diagrammatic) calculus when the group is replaced by a finite semigroup, and develop the resulting (diagrammatic) calculus.

math.RT↗

Revolutionizing Diffusion MRI Microstructure Mapping via Global Inversion

Diffusion MRI microstructure mapping (MM) is conventionally solved voxel by voxel, ignoring the fact that tissue microstructure forms a spatially organized field. This isolation leaves each estimation problem ill-posed and nonconvex. We instead cast MM as a single global inverse problem, reconstructing the entire parameter field jointly from all measurements of a subject. An untrained neural representation supplies implicit spatial priors and eases the nonconvex optimization, requiring no training data, while coregistered T1-weighted anatomy contributes structural guidance that is freely available in standard protocols. On both synthetic and in-vivo data, our method compares favorably with established voxel-wise and learning-based baselines, suggesting global inversion is a promising alternative.

eess.IV↗

TactileStep: Sole Tactile Learning for Regulating Foot-Terrain Interaction in Humanoid Locomotion

Humanoid parkour policies can traverse various terrains, but task completion may mask challenges of harsh landings, edge contacts, and unstable stance contacts. Humans naturally regulate foot-terrain interaction through tactile feedback, modulating contact compliance according to terrain stiffness. This highlights a key domain gap between humans and humanoid robots: the absence of rich tactile sensing in most humanoid systems. We address this problem with TactileStep, a deployable tactile learning framework that brings sole pressure sensing into humanoid locomotion control for softer touchdowns and more stable support. TactileStep aligns tactile simulation with the real pressure insole, allowing the policy to learn from the same contact features available on hardware. During training, we use tactile and motion cues to recognize different foot-contact phases and apply phase-aware rewards that encourage safer landing and more stable stance. Evaluated in simulation and on a Unitree G1 humanoid across diverse terrains, TactileStep reduces peak touchdown force by up to 48.8% and peak A-weighted impact noise by up to 30.1 dB over a strong perceptive baseline, while increasing stance contact area by up to 23.8%.

cs.RO↗

Echo in the Steps: Learning Perceptive Humanoid Parkour with Gated Memory

While recent advances in perceptive locomotion have enabled humanoid robots to traverse structured terrains, agile parkour in highly discontinuous environments remains an open challenge. In particular, crossing sparse footholds and narrow support regions requires precise foothold selection, effective use of visual observations, and consistent alternating foot placement during fast transitions. In this paper, we present a perceptive humanoid parkour framework that enables stable traversal across terrains with limited foothold availability using only onboard depth observations. The framework features a saliency-guided temporal perception module that combines a saliency prior with gated memory. It retains informative depth features across frames, enabling reliable foot placement from partial observations. By introducing an alternation loss, our symmetry regularization encourages alternating gait patterns and improves traversal robustness. Extensive experiments show that our method significantly improves success rate and foothold accuracy on challenging terrains in both simulation and the real world.

cs.RO↗

A new second-order consistent splitting scheme for the Natural Convection equations

A novel second-order consistent splitting scheme is proposed for the natural convection equations. The scheme is constructed using Taylor expansions about the time level $t^{n+k}$, where $k$ is a parameter to be determined. It is proved to be stable for all $k>4$, with the present study focusing primarily on the case $k=5$. Compared with the conventional splitting scheme based on Taylor expansions about $t^{n+1}$, the proposed scheme exhibits improved stability. By employing the Sobolev inequality and Gronwall's lemma, we rigorously establish stability of the proposed scheme and derive error estimates in both two and three spatial dimensions. Finally, several numerical experiments are presented to verify the stability and accuracy of the proposed scheme and to demonstrate its effectiveness.

math.NA↗

Blast freezing a black hole

What happens to information carried through an evaporating black-hole horizon? We introduce a solvable model of evaporation built from coupled Sachdev-Ye-Kitaev systems, in which an initially two-sided black hole is coupled at a finite time to a larger, colder bath. Evaporation is rapid in this model, so we refer to the process as ``blast freezing'' of a black hole. In an appropriate large-$N$ and large-$p$ limit, the two-point functions and certain four-point probes can be computed analytically. Using the two-point functions as input to a generalized HKLL reconstruction, we obtain the emergent bulk geometry of the evaporation process. We then track the information carried by an infalling particle using operator size and Renyi-2 mutual information, showing how it is preserved in nonlocal many-body degrees of freedom after the blast-freezing transition.

hep-th↗

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.

cs.AI↗

Resolvent reconstruction of minimal strings beyond KdV

Minimal strings offer two ways to study quantum surfaces: through a spectral curve or through a differential equation. We ask how much of the quantum equation can be recovered once the classical curve and its motion with the background are known. Our starting point is the resolvent of a scalar differential operator. We require its one-boundary integral to have poles only at branch points, with the resulting differential regular at infinity. These conditions give a finite reconstruction procedure at each genus. When the branch points are simple and their energy values move, we prove that the quantum corrections are unique whenever the procedure is consistent. A specified string equation provides a solution if its trace primitive satisfies the same analytic conditions. The scalar construction also leads to amplitudes with several boundaries. We introduce the method through pure gravity and Ising, then explore higher-order, dual and nonunitary models. Higher-Airy examples connect the calculation to $r$-spin theory on a degenerate locus that requires a separate treatment of the inverse problem.

hep-th↗

The Conversation Turns First: Crowd Discussion and Price Reversals in Prediction Markets

Prediction markets combine trading with public discussion of the same events. We examine whether comment-derived signals predict subsequent activity, buying direction, and changes in the leading outcome. A correlation sweep across 79 non-political Polymarket markets guides six classification experiments comparing comment features, trading features, and their combinations. On live blocks containing comments, attention nearly matches trading history in predicting heavy trading within 18 hours (PR-AUC 0.786 versus 0.790, against prevalence 0.606), with its relative advantage concentrated in short markets. Comment content carries directional information: toxicity ranks future buying direction above chance in 43 of 53 scored markets, while adding attention, sentiment, and stance to flow history increases ROC-AUC from 0.780 to 0.788. Leadership changes are predicted primarily by market state. Stance shifts against the leader before reversals in 23 of 27 evaluable markets, but adds no clear improvement in individual-block forecasting. These results distinguish attention from directional support and price uncertainty. They establish predictive associations consistent with discussion and trading responding to shared information, without identifying a causal effect of comments on markets.

cs.SI↗

A Combinatorial Proof of Hilton's Conjecture and Beyond

Using refined absorption, we prove that for every integer $g\ge 1$ and real $γ> 0$, and for sufficiently large $n$, there exists an $n^{-1+γ}$-spread distribution on Latin squares of order $n$ and girth at least $g$ that have no proper subsquares. This implies a combinatorial proof of Hilton's conjecture from the 1970s (recently proved algebraically by Allsop and Wanless) that for all sufficiently large $n$, there exists a subsquare-free Latin square of order $n$; indeed, it implies there exist at least $n^{(1-o(1))n^2}$ subsquare-free squares. Simultaneously it also implies the existence of high girth Latin squares (recently proved by Kwan, Sah, Sawhney and Simkin) and even an $n^{-1+γ}$-spread distribution on high girth Latin squares.

math.CO↗

Passive LWIR Hyperspectral Ranging via Transmittance Extraction and Distance Alignment

Passive long-wave infrared (LWIR) hyperspectral ranging enables distance estimation in low-light and nighttime scenes by exploiting atmospheric absorption features in thermal radiance received through the atmosphere.Joint estimation of temperature, emissivity, and distance is computationally expensive. Reference-range joint inversion also uses a distance-invariant effective attenuation coefficient, which can bias range estimates.We introduce transmittance extraction and distance alignment (TEDA), which decouples range estimation from temperature--emissivity inversion. In the first stage, a baseline estimator with a data-fidelity term invariant to the known absorption direction yields two closed-form smoothing branches for the slowly varying thermal continuum. An observation-derived gate combines the branches, and subtracting the blended baseline in the log domain recovers atmospheric transmittance. The second stage estimates range by matching the recovered transmittance to sensor-domain transmittance models recomputed for each candidate distance. Monte Carlo simulations show that TEDA effectively reduces the ranging bias caused by the distance-invariant attenuation coefficient approximation. In a measured scene, TEDA's mean range estimates are closer to the LiDAR medians than those of reference-range joint inversion in both evaluated patches. TEDA processes a complete $256\times256$ region of interest in 8.19~s versus 159.47~s for reference-range joint inversion, an approximately 20-fold speedup.

cs.CV↗

Beyond the Kagome Layer: Interlayer Origin of the Flat Band in FeSn

Kagome FeSn exhibits an occupied flat band at its terminated surface, whereas bulk FeSn adopts A-type antiferromagnetic order and does not display the same feature. Here, we combine density functional theory with primitive- and doubled-cell analysis and a double-layer tight-binding model to determine how interlayer electronic coupling and magnetic stacking control flat-band formation in FeSn. We identify an occupied Fe-$d_{z^2}$-derived flat-band manifold in the ferromagnetic state that is consistent with the experimentally observed surface feature. Brillouin-zone folding reveals that its flat branch originates from the $k_z=\fracπ{c}$ sector of the primitive cell, demonstrating that it cannot be understood as an isolated kagome-layer state. Instead, the state depends on coupling between neighboring kagome layers and can be seen as a result of antibonding coupling between nearest neighbor kagome layers. A double-layer tight-binding analysis identifies interlayer Fe-Fe hopping as the dominant microscopic coupling responsible for this behavior, while a contrasting unoccupied flat-band-related manifold is governed primarily by intralayer Fe-Sn hybridization. These results establish interlayer coupling and magnetic stacking as key control parameters for kagome flat bands and highlight how coupling between layers can generate and tune extended correlated-electron states in quantum materials.

cond-mat.mtrl-sci↗