arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

CODEBLOCK: Learning to Supervise Code at the Right Granularity

Supervised fine-tuning of code LLMs typically applies uniform cross-entropy loss to all response tokens, implicitly assuming that every token provides equally useful learning signals. Recent token-level selection methods challenge this assumption in natural-language SFT by supervising only high-value tokens. However, such pointwise selection can fragment the syntactic structures and program dependencies of code, leaving supervision scattered across incomplete code units. In experiments on 30K code instruction-response pairs, under the same 10% token budget, supervising complete coding blocks improves average performance by 11.7 points over prior approaches that supervise isolated tokens. Motivated by this observation, we propose CodeBlock, a structure-aware sparse supervision framework that uses complete, parser-aligned coding items as the basic units of supervision. CodeBlock constructs coding items from high-quality code instruction data, estimates their supervision utility using GCE, which is more robust to low-probability tokens, and further adjusts their supervision priority using data-flow reach and bridge signals. During training, the full response is retained as context, while loss is applied only to the selected supervision units. Experiments show that partial supervision over only about 6.9% of response tokens consistently outperforms full-token SFT across all five model settings, suggesting that effective code SFT depends not only on identifying high-value supervision, but also on allocating it at the appropriate structural granularity.

cs.LG↗

Guava: Distilling Frontier VLM Agents into a Compact Model with a Manipulation Harness

Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-language-action systems by combining high-level reasoning with external modules for perception, planning, and control. However, it remains unclear what makes an effective harness for embodied manipulation, and to what extent such a harness can unlock embodied capabilities in a wide range of reasoning models. In this work, we present Guava, a harness framework for embodied tool use developed through systematic exploration of the design space of agent workflows, action spaces, and observation spaces. Our study identifies three key ingredients for effective embodied agents: iterative perception-reasoning-action loops, semantic action abstractions, and multimodal observations. To understand whether these design principles are universal even to small models, we develop an end-to-end training pipeline that distills embodied manipulation capabilities into a 4B open-source model using fewer than 2K trajectories collected entirely in simulation. Experimental results in both simulation and real-world environments show performance comparable to frontier proprietary models while exhibiting strong generalization to unseen objects, novel instructions, and long-horizon tasks. Results suggest that a well-designed harness can serve as a scalable, model-agnostic interface for embodied manipulation, enabling strong emergent embodied capabilities in compact open-source models with minimal training data.

cs.RO↗

Simulating Robotic Locomotion in Sand: Resistive Force Theory in an Open-Source Physics Engine

Recent advancements in Resistive Force Theory (RFT) enable approximation of ground reaction forces for locomotion in sand without the computational expense of modeling interactions with individual grains. However, these tools have been absent in 3D physics engines commonly used for robot simulation. We explore if resistive force approximations are sufficient, when integrated with standard dynamics calculations, to provide a stable substrate for a freely walking robot. To determine this, we implement 3D Granular Resistive Force Theory (3D RFT) in a physics simulation engine, MuJoCo. We verify simulations in multiple scenarios to demonstrate that key trends due to end effector shape, speed, and loading are preserved. Our implementation predicts both walking distance and foot sinkage of a 12-Degree of Freedom hexapod robot within 7\% of experiments in sand. While RFT has inherent approximations, the open source tool described here has potential to help develop new and improved robot designs to traverse granular media substrates.

cs.RO↗

Capability Provenance in Language Models: A Case Study in Social Reasoning

We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social reasoning versus STEM reasoning in OLMo3-7B. Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities. We compute gradient-based attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, aggregate influence across WebOrganizer's 24-format x 24-topic taxonomy (576 bins), and contrast benchmark pairs in a 2x2 design that varies domain (social vs. STEM) and capability type (reasoning vs. knowledge): SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM. Social and STEM reasoning draw on qualitatively distinct corpus regions, and the contrast is sharper at the reasoning level than at the knowledge level. Targeted machine unlearning provides partial causal validation: forgetting high-attribution topics (e.g., Literature for SocialIQA) degrades the aligned benchmark more than within-topic random baselines. We open-source all code, data artifacts, influence scores, and checkpoints at https://github.com/HCAI-Lab-GT/capabilibara and https://huggingface.co/HCAI-Lab-GT.

cs.CL↗

Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale

Enterprise agent language models must coordinate specialist agents over continuous business event streams, yet most multi-agent evaluations assume discrete request-response workflows. We present a controlled simulation based on a sanitized enterprise research snapshot: 208 scenarios, 393 events, and 1,051 expected agent calls across Persona (<10 agents), Department (20-80), and Enterprise (200) registries. We compare DAG Plan & Execute with ReAct and evaluate a minimal Task Manager for priority inference, related-event merging, and preemption. Both systems perform well on the Persona cohort but score lower on the Enterprise cohort, especially for Simple scenarios; a larger, more confusable discovery space may contribute. The tested DAG Plan & Execute implementation generally maintains higher agent-call precision, whereas ReAct shows stronger recall and Expected Answer Completeness in several Enterprise conditions. The Task Manager reduces high-priority queue latency by 14-75% and improves related-event Expected Answer Completeness by over 20 percentage points at Enterprise scale. These controlled-simulation results characterize the tested systems but do not establish production readiness.

cs.AI↗

On the Emergence of Discrete Spectrum for Weakly Disordered Schrödinger Operators

We investigate the spectral properties of the Anderson operator perturbed by a localized negative potential, \(-V\). Specifically, we analyze the random Schrödinger operator defined by \(H = -Δ+\ve \sum_{n} ω_n χ_n - V\), where the unperturbed operator exhibits a disordered energy landscape. Our primary focus is to establish precise estimates on the number of negative eigenvalues (bound states) induced by the attractive perturbation. By analyzing the competition between Anderson localization and the binding capacity of the potential, we provide quantitative bounds on the discrete spectrum. These results offer new insights into how randomness enhances the eigenvalue bounds.

math-ph↗

Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence

Embodied artificial intelligence (AI) couples perception and learned decision making to actions that change the physical world. This coupling distinguishes an embodied agent from a conventional connected controller: the agent maintains task state and uncertainty, reasons about the consequences of actions, and adapts from subsequent observations. Wireless networking becomes relevant when perception, inference, or coordination is distributed, but it should not replace local safety control. This article develops a tutorial perception--communication--action (PCA) architecture that exposes task state, action deadlines, uncertainty, agent intent, and safety envelopes to a 6G orchestration plane. It separates capabilities already addressed by 5G and 5G-Advanced from functions that motivate 6G, including task-state interfaces, semantic freshness, predictive digital twins, and safety-aware coordination across agents. A multi-robot simulation study is retained to illustrate joint sensing, communication, and computation control. The results show where network orchestration improves task utility and where local autonomy remains essential.

cs.NI↗

Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection

Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent studies have revealed its vulnerability to backdoor attacks. Existing attacks typically require white-box access or label modification and focus on geometric attacks such as object disappearance or bounding-box manipulation. In this paper, we present Mirage, a black-box and clean-label backdoor attack against deep neural network-based LiDAR 3DOD. Mirage injects a small number of label-consistent poisoning samples into the training set, causing the model to learn a malicious association between a trigger pattern and an attacker-chosen target class while preserving normal training semantics. As a result, the compromised model behaves normally on benign inputs yet systematically misclassifies triggered objects as the target class during deployment. We evaluate Mirage on multiple state-of-the-art LiDAR 3DOD models and benchmark datasets. Experimental results show that Mirage achieves a 73% misclassification success rate with a poisoning rate of only 0.5%, while maintaining detection performance close to that of benign models.

cs.CV↗

On Fixed-Time Stability of Continuous Dynamics for Non-Monotone Variational Inequalities

Non-monotone variational inequalities (NMVI) are an important class of problems that generalize non-convex optimization and have various applications in optimization theory, machine learning, game theory, and economics, among others. Most existing work on NMVIs focuses on the asymptotic convergence of algorithms proposed to solve these problems. In this paper, we tackle the problems of exponential and fixed-time stability of the solution set of a class of NMVIs for both unconstrained and constrained problems. We first present novel conditions guaranteeing exponential stability of solutions to unconstrained NMVIs for a uniquely constructed dynamical system under mild assumptions on the gradient of the non-monotone map. Then, under similar assumptions, we construct another novel dynamical system whose equilibrium point is fixed-time stable, i.e., the trajectories reach the equilibrium within a fixed time, independent of the initial conditions. For the case of constrained NMVIs, we employ a continuous-time variant of the Korpelevich method for exponential stability of the solution set, and provide a novel scaling factor in the dynamics to achieve fixed-time stability. We illustrate the efficacy of the proposed modified dynamical systems through numerical simulations and conclude the paper with a brief note on the behavior of the discretized variant of the proposed dynamics and on further work that remains to be done.

math.OC↗

Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification

Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing methods often produce predictions that are inconsistent across taxonomic levels. For example, a model may predict a fine-grained category whose parent category contradicts its simultaneously predicted higher-level label. By analysis, the issue originates from false negative labels when contrastive comparison involves multiple taxonomic levels. To this end, we propose to restrict contrastive comparisons to categories within the same taxonomic level. In addition, we adopt a group-balanced design, ensuring each taxonomic level receives adequate optimization. As a result, the proposed framework improves both hierarchical consistency and classification accuracy from coarse to fine granularity. We train our model with TreeOfLife-10M based on BioCLIP and evaluate it across multiple hierarchical classification benchmarks, where the model demonstrates significantly improved hierarchical consistency in both Euclidean and hyperbolic spaces. Notably, on iNaturalist 2021 (iNat21), our method improves average accuracy across levels by 30.47% over the baseline, highlighting its effectiveness for hierarchical zero-shot classification.

cs.CV↗

Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation

Infrared small target detection (IRSTD) in high-resolution images is crucial for unmanned aerial vehicle (UAV) surveillance and UAV-based ground monitoring. However, small target size, weak features, and interference from complex dynamic backgrounds make IRSTD challenging. Existing methods incur redundant computation in non-target background regions and insufficiently exploit target context, limiting detection performance. To address these issues, we propose ECFNet, an efficient coarse-to-fine IRSTD framework with attention prior-guided knowledge distillation. In the coarse stage, we design a region binary classification network (RBCN) on grid-based multi-scale feature maps to efficiently identify target-containing context region proposals. A new denoising-assisted training strategy incorporates noisy ground-truth (GT) masks into RBCN feature maps and trains the network to reconstruct the original GT masks. This auxiliary task encourages explicit learning of target-background context to better distinguish target proposals from background regions. In the fine stage, we customize a lightweight target detector to the coarse-stage region proposals to balance accuracy and efficiency. Furthermore, we introduce a knowledge distillation strategy guided by a teacher-student cross-attention prior. This strategy directs the student to focus on critical target regions, enhancing discriminative feature representations for infrared small targets. Extensive experiments on three real infrared datasets demonstrate that ECFNet outperforms existing single-stage and two-stage approaches while maintaining high real-time processing efficiency. Code: https://github.com/IVPLabs/ECFNet.

cs.CV↗

Formation and dynamics of self-bound droplets in dipolar molecular condensate

We study self-bound quantum droplets in the regime dominated by microwave-induced non-axisymmetric dipole-dipole interactions, using the extended Gross-Pitaevskii equation with the Lee-Huang-Yang corrections. We identify the existence region through numerical simulations and employ an anisotropic Gaussian-super-Gaussian variational ansätz to capture the characteristic density profile of the droplets, with a Gaussian profile along the narrow $x$ direction and super-Gaussian profiles in the extended $(y,z)$ plane. Within this variational framework, we characterize the self-binding, spatial localization, and density-compression properties of the droplets and find good agreement between the variational predictions and the numerical results. Collisions between droplets moving along different directions reveal a strong directional dependence, with outcomes ranging from quasi-elastic rebound and merger to fragmentation. In addition, we explore the rotational dynamics of a single self-bound droplet about all three Cartesian axes, revealing rich and controllable three-dimensional rotational dynamics. Together, these results demonstrate how non-axisymmetric dipolar interactions provide versatile means for controlling the translational, collisional, and rotational dynamics of self-bound quantum droplets.

cond-mat.quant-gas↗

Unified theory of oscillons and modes

We show that an oscillon can be understood as a localized discrete resonant (non-normalizable) mode. Specifically, oscillon in the vacuum arises from the threshold mode, which because of nonlinearity gets localized. Following this idea, we find {\it wobblerons} - nonlinear excitations of kinks, that is, oscillons-kink bound state. Now, the oscillon can also originate in an antibound mode, i.e., a discrete, positive energy but non-normalizable mode.

hep-th↗

StreamPPG: Low-Latency rPPG Estimation via Consistent Privileged Learning

Remote photoplethysmography (rPPG) estimates the blood volume pulse (BVP) signal from facial videos, enabling contact-free health monitoring. Conventional clip-wise approaches, which use video clips as input, require capturing over one hundred frames before inference, thus introducing several seconds of delay and hindering real-time use. Meanwhile, frame-wise approaches struggle to capture long-range temporal and periodic features of physiological rhythms, and therefore lead to reduced estimation accuracy. To overcome these issues, we propose StreamPPG, a unified architecture that enables low-latency frame-wise physiological signal estimation while achieving competitive accuracy compared with clip-wise approaches. StreamPPG is trained under a consistent privileged learning (CPL) strategy, which leverages ground-truth rPPG signals as privileged information to enhance the model's representation capability. Extensive experiments demonstrate that StreamPPG achieves state-of-the-art accuracy across multiple datasets while maintaining real-time throughput on edge devices.

cs.CV↗

Rethinking Object-Centric Representations for Video Dynamics Modeling

Learning to decompose videos into persistent objects is a fundamental challenge in unsupervised object-centric representation learning. Despite recent progress, existing methods struggle to simultaneously achieve accurate object segmentation, consistent identities over time, and reliable foreground-background separation. To address these challenges, we introduce UniSlot (Unified Slots), an unsupervised framework for learning robust and disentangled object-centric representations from videos. UniSlot explicitly separates object appearance from its 3D-aware geometric pose in the scene, linking object identity to appearance while leveraging depth to better distinguish objects from their surroundings. UniSlot achieves state-of-the-art performance in unsupervised object-centric video decomposition and tracking across synthetic and real-world benchmarks, yielding substantially tighter object masks and reducing background leakage while preserving object identities. Beyond decomposition and tracking, these improved representations translate directly to downstream tasks such as unsupervised object dynamics prediction, enabling more accurate forecasting of future object trajectories.

cs.CV↗

Universal Dynamical Response to Slow Driving in Chaotic Systems

We propose a unified perspective on classical and quantum chaos based on the sensitivity of a system's stationary states to slow driving. We probe this sensitivity via the system's susceptibility to the average protocol speed, which we call the ``speed-Fisher information," and relate it to irreversible entropy production in the system. We show that chaotic dynamics manifests as a divergence of the speed-Fisher information with the protocol time, and that this response is controlled by the perturbation's low-frequency spectral weight. This approach to chaos applies to both classical and quantum Hamiltonian systems, and naturally extends to non-Hamiltonian classical flows. We illustrate this framework with simple classical and quantum examples, along with a non-Hamiltonian flow that qualitatively exhibits analogous low-frequency spectral behavior.

cond-mat.stat-mech↗

Real vs. Complex Spectral Bases for Neural Operators: The Role of Green's Function Alignment

Fourier Neural Operators (FNO) learn solution operators of partial differential equations by parameterizing global convolutions in the complex Fourier domain. For real-valued PDE solutions, the complex FFT carries representational redundancy through conjugate symmetry. We introduce the Hartley Neural Operator (HNO), the exact real-valued mirror of FNO: it replaces the FFT with the purely real Discrete Hartley Transform and learns a single real multiplier per retained spectral mode, with no complex arithmetic. Because the real Hartley spectrum is not halved by conjugate symmetry, HNO retains twice as many frequency corners as FNO but one real weight where FNO carries a complex pair, so the two operators are iso-parametric at equal width and differ only in spectral basis. Our central thesis is that the best basis is a property of the operator. Self-adjoint elliptic operators (Poisson, biharmonic) have real, symmetric Green's functions that the real Hartley multiplier diagonalizes exactly, and HNO is favored there. Time-dependent operators carry phase, from oscillation in the wave equation to transport in advection, Burgers, and Navier-Stokes, which a real diagonal multiplier cannot represent, so FNO is favored there, and increasingly so with the operator's phase content, leaving the phaseless heat equation as the borderline case. Training both operators identically and benchmarking across PDE classes, initial-condition families, and boundary conditions, we find an elliptic-versus-time-dependent split that is monotone in operator phase content and matches the Green's-function theory we develop. Rather than a universal winner, our findings give a predictive rule: match the spectral basis to the symmetry of the solution operator.

cs.LG↗

A new $H_0$ measurement with SNe Requiem and Encore using $\texttt{Gravity.jl}$

We present a strong-lensing (SL) analysis of the galaxy cluster MACS J0138.0-2155 (z=0.336), the first known lens cluster discovered to host two distinct multiply imaged Type Ia supernovae (SNe): SN Requiem and SN Encore. Both SNe are located in the massive, multiply imaged red galaxy MRG-M0138 at z=1.949. The projected total mass of this cluster has been investigated with several independent lens models (Suyu+26; Pierel+26), using a sample of 23 spectroscopically confirmed multiple images from 8 background sources (0.767 10%). The forthcoming reappearance of SN Requiem offers an immediate opportunity to significantly improve constraints on H0, provided that lens-model systematics are controlled. These results establish M0138 as a premier anchor for high-precision cluster TDC.

astro-ph.CO↗