arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,207 records · Page 67Linked to original sources

Toward Proactive RF Charging Scheduling: Generative AI for Decision Support

Radio frequency wireless power transfer (RF-WPT) is an enabling technology for supporting uninterrupted communications in future Internet of Things systems by reducing the need for battery replacement and mitigating battery-waste-related issues. For large-scale RF-WPT deployment, one of the main challenges is the scheduler-level resource allocation. Specifically, the RF charger must decide how much energy to deliver, when, and to whom, under limited charging resources, incomplete receiver-side information, and uncertain near-future charging conditions. This article positions generative artificial intelligence (GenAI) as a promising tool for this setting because it can foresee multiple plausible charging scenarios conditioned on coarse operational context and receiver-side information. We propose GenAI to act as an uncertainty-aware support layer for the RF-WPT scheduler rather than as a standalone forecasting or decision-making tool. To this end, we first revisit the main challenges of RF-WPT scheduling, and discuss how major GenAI families can support uncertainty-aware charging decisions by generating scenario-based inputs for downstream tasks. We then present a case study showing that distribution-aware prediction can improve robust charging decisions over deterministic, ensemble, and non-learning baselines, particularly under risk-sensitive objectives. Finally, we outline key open challenges and future research directions.

eess.SY↗

Adversarial Vulnerabilities of Learned Telesurgery Policies

While not yet in clinical deployment, learning-based policies are increasingly considered to augment the dexterity of human surgeons in robot-assisted surgery. Can the end-to-end mapping from visual observations to robot actions be vulnerable to adversarial attacks? We present the first study of adversarial vulnerabilities in learning-based policies for surgical robotics, conducted in a laboratory white-box setting where the attacker is assumed to have access to policy information and injects perturbations into the video stream transmitted over the network. Two attack modes are considered: (a) disruptive attacks, where subtle visual perturbations interrupt policy execution without being noticed by a surgeon, and (b) steering attacks, where perturbations steer policy actions toward attacker-specified directions. We study three adversarial attack methods, each with increasing access to policy information, and evaluate their impact on two surgical subtasks: debridement and suturing, performed on phantoms. Our evaluation covers three end-to-end policy architectures: ACT, Diffusion Policy, and pi0. In addition, we identify a vulnerability to photometric perturbations, which mimic natural visual changes such as lighting variation. Results from 620 physical experiments suggest that state-of-the-art policies can be significantly disrupted, resulting in an average 61% reduction in surgical subtask success rates. These findings suggest that adversarial vulnerabilities are important to consider for learned telesurgery policies. Project page: https://surgical-robotics.github.io/adversarial-vulnerability/

cs.RO↗

MoGeFlow: Flowing Through Motion Codebook Geometry for Text-to-Motion Generation

Vector-quantized tokenizers have become the dominant interface for text-to-motion generation, and the priors built on them treat motion codes as unordered categorical labels. We show that this discards structure the tokenizer already learned. A motion code embedding is a decoder-bound movement prototype, and learned motion codebooks carry measurable, decoder-causal geometry: code distances track the distances between the movements those codes decode to, the alignment collapses under shuffled controls, and displacement in code space predicts how far the frozen decoder moves. Geometry of this kind calls for a different generative object---a continuous trajectory through the code space rather than a sequence of decisions over a vocabulary. We introduce MoGeFlow, a text-conditioned flow over part-structured motion-code frames. Its tokenizer quantizes each body-part group separately, so every decoder input is a point on a combinatorial lattice of decoder-bound states, and the flow regresses onto that lattice. Because the decoder responds smoothly to code-space displacement, generated states are decoded as they are, and the detail the flow places between codes reaches the decoder instead of being rounded away. Quantization therefore serves generation entirely at training time, shaping the space rather than gating the samples: under matched conditions the lattice-shaped target beats both the unquantized latents it is built from and a matched variational latent space. MoGeFlow sets a new state of the art in text--motion alignment on HumanML3D, leads Top-3 retrieval and MultiModal Distance on KIT-ML, and on the large-scale MotionMillion benchmark surpasses a 7B-parameter model in retrieval with an order of magnitude fewer parameters. Quantization shapes the space; the flow generates through it.

cs.GR↗

Consensus Time in 3-Majority and 2-Choices Is Determined by the Maximum Initial Opinion Density

We establish the correct parameter governing the convergence time of the 3-Majority and 2-Choices dynamics on the complete graph in the synchronous model. Recent work [Shimizu and Shiraga, PODC'25] provides matching upper and lower bounds on the number of rounds to consensus, but only in a weak sense: the bounds are shown to coincide for some initial opinion configuration. In contrast, we obtain tight bounds in a strong sense, with upper and lower bounds matching up to logarithmic factors for every initial configuration. Let $α^{(0)}$ be the initial opinion-frequency vector, and denote by $\|α^{(0)}\|_\infty$ its maximum entry. We show that 3-Majority reaches consensus in $\tildeΘ(\min\{\|α^{(0)}\|_\infty^{-1},\sqrt n\})$ rounds w.h.p., while 2-Choices reaches consensus in $\tildeΘ(\|α^{(0)}\|_\infty^{-1})$ rounds w.h.p. Our results demonstrate that the convergence time of both dynamics is governed not by global parameters such as the number of opinions $k$ or the squared $\ell_2$ norm of the initial opinion distribution, but rather by the ``local'' parameter $\|α^{(0)}\|_\infty$, the maximum initial opinion density.

cs.DC↗

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding

Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural weaknesses in existing reasoning-based reward models: they do not systematically verify every condition described in the prompt, and the visual evidence supporting each judgment remains implicit in their free-form reasoning. We propose SG-PVR, a video reward model that addresses these limitations through plan-and-verify reasoning grounded in spatio-temporal scene graphs. The verification plan decomposes the prompt into atomic claims, making the set of requirements to be checked explicit. The spatio-temporal scene graph, encoding entities, attributes, and temporally grounded relations, is extracted from the video and maintained as a persistent structured visual reference throughout reasoning. Each claim is verified against both the video and the scene graph, anchoring judgments in explicit visual evidence. SG-PVR achieves strong performance on semantic alignment, including fine-grained temporal semantics. As a test-time reranker, it further enhances compositional alignment in T2V generation.

cs.CV↗

SN 1006: A Cosmic Laboratory for Investigating Shock Acceleration Physics

SN 1006 is a historical Type Ia supernova remnant that exhibits non-thermal emission ranging from radio to multi-TeV $γ$-rays. Most of this emission (particularly X-rays and $γ$-rays) is concentrated in polar caps aligned with the ambient magnetic field, which makes it an ideal laboratory for studying cosmic ray (CR) acceleration at different shock obliquities and the hadronic/leptonic nature of the $γ$-ray emission. We model SN 1006's morphology, multi-wavelength spectrum, and radial profile using a self-consistent multi-zone kinetic model of particle acceleration that accounts for: CR-driven shock modification, magnetic field amplification, drift in magnetic fluctuations, and temporal dynamics including adiabatic and synchrotron losses. Our model can reproduce both the observed spectral and spatial properties, with the exception of the radio profile that we argue requires 3D hydrodynamic effects to replicate. We find that quasi-parallel regions (where the shock normal aligns with the ambient magnetic field) exhibit very prominent CR acceleration ($\sim$20% efficiency), while quasi-perpendicular regions exhibit efficiencies below 1%, consistent with the results of kinetic simulations. We also find that electrons are responsible for the majority of the $γ$-ray emission from SN 1006 (i.e., it is a leptonic source), with the exception of the northwest region due to an encounter with a dense cloud.

astro-ph.HE↗

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

The software engineering capabilities of general-purpose agent harnesses remain underexplored, and existing benchmarks offer limited support for comparing these harnesses under consistent conditions. To address this gap, we introduce Claw-SWE-Bench, a unified benchmark that enables researchers to systematically assess the capabilities and efficiency of general-purpose harnesses on software engineering tasks. The benchmark contains 350 real-world GitHub issue-resolution instances across eight programming languages and 43 repositories and provides a shared adapter protocol to align task inputs, outputs, and execution environments across harnesses. Experiments show that general-purpose harnesses can effectively resolve real-world software issues and that their success rates and resource consumption vary substantially even when the underlying model is held fixed. To lower evaluation costs and support faster debugging and iteration, we also provide Claw-SWE-Bench Lite, an 80-instance subset designed to preserve the key evaluation properties of the full benchmark. We hope this benchmark will help researchers better evaluate and understand the performance of general-purpose harnesses on software engineering tasks and guide the development of more capable and efficient harnesses. The data is available at https://github.com/opensquilla/claw-swe-bench and https://huggingface.co/datasets/TokenRhythm/Claw-SWE-Bench.

cs.LG↗

Laplace--King representations: density and spectral theory

Laplace--King representations combine spherical harmonics with King functions [Wang et al., Chin. Phys. B \textbf{34}, 065201 (2025)], the radial kernels of shifted isotropic Gaussians. The radial parameters may vary across angular modes. We prove that, for every angular degree, fixed-width kernels with positive real shifts have dense complex linear span in a Gaussian-weighted radial \(L^2\) space. Finite Laplace--King representations are dense in the corresponding three-dimensional weighted space; allowing variable widths preserves density in the same reference norm. A generating identity connects the kernels to generalized Laguerre polynomials. The self-adjoint King operator is unitarily equivalent to the free radial Schrödinger operator; its spectral resolution defines a continuous King mixture model (KMM) through imaginary-shift kernels in a distinct weighted Hilbert space.

math-ph↗

Interaction Dynamics MPC for Knee Rehabilitation Exoskeletons: A Closed-Loop SEA Outer-Loop Study

Safe rehabilitation is an interaction-dynamics problem: the controller must regulate a prescribed motion while absorbing involuntary spasm, voluntary effort, actuator compliance, and model mismatch as disturbances. This paper instantiates the predictive interaction-dynamics framework of the base pHRI formulation on a SEA knee joint. SEA feedforward reduces the gravity-compensated knee to the same scalar double integrator as the base framework, while a dynamic-residual measurement from spring deflection supplies an interaction-disturbance observation. A steady-state target converts the estimated disturbance into a cancelling input, and a finite-horizon quadratic program regulates deviations from that target under range-of-motion, torque, and velocity constraints. The evaluation matches stiffness and damping across controllers so gains cannot be attributed to higher impedance. Under a motion-opposing $15\unit{Nm}$ step, classical impedance and MPC without estimation produce about $500\unit{mrad}$ steady-state error, whereas Kalman-augmented interaction MPC reduces this to $1.17\unit{mrad}$ at 100~Hz and $0.70\unit{mrad}$ at 500~Hz; the 500~Hz peak is $7.27\unit{mrad}$. In 30 randomized trials, the 95th-percentile peak is $21.57\unit{mrad}$. Bounded Assist-as-Needed scheduling, a corrective-channel energy tank, constrained OSQP stress cases, direct MuJoCo execution, and a posture-clamped MyoSuite knee slice are implemented. The framework holds on a single-mass, closed-inner-loop SEA approximation; an explicit two-mass plant with a finite-bandwidth, pole-placed inner torque loop (Section~VIII) confirms this for nominal tracking but shows delivered torque can overshoot the commanded bound by 21.7\% near saturation. Scope excludes clinical intent recognition, full-system passivity, safety certification, hardware trials, and multi-joint validation.

eess.SY↗

Oxygen deficiency and valency reconstruction in multiferroic V-doped HfO$_2$

The interplay of oxygen deficiency and vanadium multiple valency in the candidate multiferroic V-doped $Pca2_1$ hafnia HfO$_2$ is studied by first-principles calculations. Low-lying V majority gap states accept electrons from oxygen-vacancy donors, reducing their formation energy, and converting nominal V$^{4+}$ centers into V$^{3+}$. The resulting local magnetization and screening changes are reflected in the calculated V core-level shifts, which are consistent with the experimentally observed XPS signatures. The calculated V$^{3+}$/V$^{4+}$ population ratio determined by oxygen vacancies only matches experiment in reducing conditions, suggesting that additional electron reservoirs may contribute under ALD growth conditions. A similar scenario also seems to apply to the recently observed multiferroicity in Cr-doped hafnia, where oxygen deficiency is intrinsic to the growth technique.

cond-mat.mtrl-sci↗

Running the Gauntlet: Hard Agentic Tasks

As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we introduce GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), across five less-covered professional applications (Video Editor, Workflow Builder, 3D Modeller, Flight Analyser, and Circuit Designer), each with 27 vision-intensive tasks (135 in total). Our benchmark provides a modular pipeline that comprises an environment compatible with both open- and closed-source agent frameworks, a controlled web-based application, a well-structured task suite, and an automated evaluation engine with diverse metrics. Contrary to widespread expectations, our empirical results reveal that frontier agentic systems remain far from achieving human-level performance. Even the state-of-the-art agent achieves only a 28.2% success rate on our GauntletBench, highlighting the limitations in these overlooked capabilities and generalisation. By comparison, non-expert human annotators achieve over 80% success on our challenging yet feasible tasks, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.

cs.LG↗

AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents

Indirect prompt injection (IPI) is a major security threat to LLM-powered agents. Thus, a growing body of work has proposed a variety of defensive approaches against IPI which can be grouped into three broad categories: prompt-level, filter-based, and system-level designs. However, commonly used benchmarks for evaluating defense, such as AgentDojo, are \emph{inherently static}, generating a fixed distribution of IPI attacks. Consequently, a defense can score well on them without being robust to adaptive threats. We introduce \textbf{AutoDojo}, a generative benchmark built on AgentDojo and AgentDyn that generates IPI \emph{adaptively} for a given agent and defense. It supports six task suites across the two benchmarks, covering banking, communication, travel, shopping, coding, and everyday-assistant domains. AutoDojo optimizes injections under a strict black-box setting, observing only whether a candidate injection succeeds, and can readily integrate, and often improve on, any existing black-box attack. Across ten defenses and five target models, AutoDojo demonstrates that standard static benchmarks often significantly overestimate defense efficacy. Moreover, we show that most existing defenses either sacrifice considerable utility or are insecure. Finally, we demonstrate that attack success interacts with task specification, with under-specified tasks particularly vulnerable. AutoDojo is available at https://github.com/xhOwenMa/AutoDojo

cs.CR↗

Uniform integrability of the distance to the nearest leaf in random trees

We study the distance from the root to the nearest leaf, the analogous quantity for a uniformly chosen vertex, and its protection number, in size-conditioned simply generated trees. We prove a uniform exponential tail bound for each of these quantities, valid for arbitrary offspring distributions. As a consequence, these random variables are uniformly integrable of every order. This yields convergence of all moments to those of the corresponding local limit. The argument is probabilistic and unified across the three quantities.

math.PR↗

Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning

Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control. In ECL, feature drift can propagate through sequential decision-making under closed-loop control, turning representation changes into compounding behavioral deviations on previously learned tasks. A key challenge in ECL lies in structured skill reuse across continually evolving tasks, since existing methods primarily focus on skill learning without explicitly organizing them for coherent task execution. To address this issue, we propose SCE, a Skill-Compositional Experts framework for ECL. SCE builds a skill base via Compositional Skill Grounding (CSG), which decomposes task demonstrations into reusable skills. Based on this, Dual Execution-and-Transition Experts (DETE) enable new task learning through skill composition, where one branch ensures skill execution and the other supports transitions between skills for coherent behavior. Experiments on LIBERO benchmarks and real-world manipulation tasks show that SCE improves retention and overall task performance. Further feature drift analyses and ablation studies verify the effectiveness of our method. Project website: https://eqcy.github.io/sce/.

cs.RO↗

RepNN: Tackling spectral bias in deep neural networks for regression and PDE problems via parameter reparameterization

Deep neural networks (DNNs) have achieved remarkable success in scientific computing, yet they often suffer from spectral bias in capturing oscillatory and multiscale behaviors. In this study, we investigate this limitation by examining the failure of shallow ReLU neural networks in fitting high-frequency functions. This observation identifies two important factors in resolving rapid oscillations: the initial slope scale and the distribution of partition points induced by the networks. Motivated by this analysis, we propose RepNN, a reparameterized neural network model with ReLU or tanh activations designed for high-frequency and multiscale problems. The key idea is to reparameterize the weights and biases in the first hidden layer, which enables effective control of the initial slope scale and provides an appropriate distribution of the initial partition points. Furthermore, treating the reparameterized weights and biases as trainable parameters allows the DNN to achieve adaptive frequency scaling during training. In addition, we derive quantitative estimates for the output and slope magnitudes of the reparameterized DNN to guide the initialization of the proposed method. Numerical experiments, including multiscale one-, two-, and four-dimensional function approximations, forward and inverse PDE problems in combination with physics-informed neural networks (PINNs), and operator learning for an earthquake problem using real data, demonstrate that RepNN improves the predicted accuracy of vanilla DNNs in capturing highly oscillatory features. These results indicate that RepNN provides an effective and flexible approach for overcoming spectral bias and applying DNNs to multiscale problems.

cs.LG↗

Depth Projections and the Local Nature of the Cass Criterion

This paper defines the depth projections of the indifference and offer hypersurfaces, and of the transformation frontier. These projections are functions conceived to measure curvature (thus are not numbers, as the Gaussian curvature is) and their second-order behavior is governed by the Hessians of the expenditure and profit functions. In particular, up to second order, the depth projection of the offer hypersurface is always exactly twice that of the indifference hypersurface. In an overlapping generations economy with production, these projections lead to the derivation of general statements for the necessity and sufficiency of the Cass criterion in the form $\sum^{\infty}_{t=1}1/(\Vert p_{t}\Vert\sum_{h\in G_{t}}\Vert c^{h}\Vert)=\infty$, which allows unbounded dynamics for the demography and for per capita endowments, and reconciles the Cass criterion with the Balasko--Cass--Shell relabeling algorithm. Furthermore, the assumptions behind these statements reveal the local nature of the Cass criterion: only the behavior of the indifference hypersurfaces and transformation frontiers in a neighborhood of the equilibrium allocation matters.

econ.TH↗

Fermi-eROSITA cross-correlations suggest annihilation lines of a $\sim70$ GeV WIMP

Galaxy clusters, with their deep potential wells, provide some of the strongest constraints on the velocity-dependent ($p$-wave) annihilation of weakly interacting massive particle (WIMP) dark matter. Even weaker signals can be extracted from sufficiently many aggregated clusters, by cross-correlating $γ$-rays with large-scale structure tracers. Cross-correlating 16.3 years of Fermi-LAT data with the eROSITA map of the western Galactic hemisphere, a blind composite matched-filter scan of the X-ray-correlated spectrum for the $χχ\toγγ$, $γZ$, and $γh$ annihilation-line triad of a single WIMP $χ$ finds $\overline{m}_χc^2\simeq71.9$ and $67.5$ GeV triads, each with a local $Z$-score $\sim3$ corresponding to a (trial-corrected) global $\sim2σ$. The $\sim72$ GeV candidate coincides with the mass implied by interpreting the recently reported $\sim43.2$ GeV $γ$-ray line in three nearby clusters as its $γZ$ channel. At this externally fixed mass, the triad reaches $\sim3σ$ alone, and combines with the cluster line to $3.9$-$5.1σ$ (depending on the significance of the latter). The inferred intrinsic channel cross sections are a few $10^{-19}$ cm$^3$ s$^{-1}$ for WIMPs tracing the X-ray-emitting gas. Future tests of this interpretation are outlined.

astro-ph.HE↗

Recursive Scaling in Masked Diffusion Models

Masked diffusion models (MDMs) generate sequences by iteratively refining a partially masked state and committing tokens in parallel. We introduce recursion in MDMs and propose new Recursive Masked Diffusion Models (R-MDMs), which apply a shared denoising transformer $L$ times within each denoising step, adding recursive depth as an additional compute axis without increasing parameter count. Across structured generation tasks, recursive depth improves quality at fixed parameter budget, matches substantially larger non-recursive models at matched FLOPs, and can reduce the number of denoising steps needed to reach a target quality. We interpret these gains with a dependence--fidelity decomposition of parallel decoding error: recursion refines model marginals at a fixed masked state, whereas denoising steps change that state by committing tokens. Building on this analysis, we propose to treat decoding as a two-axis decision (how many loops to run and which tokens to commit) and show that entropy-guided adaptive rules improve the quality--compute frontier over fixed schedules, transferring across various tasks on Sudoku, Countdown, RNA, and executable math generation. Together, these results establish recursive depth as a practical, complementary test-time scaling mechanism for MDMs.

cs.LG↗