arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,405 records · Page 78Linked to original sources

OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents

LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from filesystem-level coding agents because the main risk is unnecessary access to private data. We introduce OverAct, a controlled benchmark spanning eight privacy-sensitive domains with deterministic, judge-free scoring, together with an interpretive decision-theoretic framework that yields three testable predictions. Across seven models from four families, all models significantly exceed authorized scope. Request specificity is the strongest predictor of severity, over-authorization grows sublinearly with tool-pool size, and decoding temperature has little effect. These patterns are consistent with a cost-asymmetry account, suggesting that over-authorization arises more from structural decision tendencies than from decoding randomness. We also propose SelfAudit, a zero-shot inference-time method that generates request-grounded justifications and filters unjustified calls before execution. Ablation shows that explicit filtering is the main driver of scope reduction. SelfAudit reduces privacy-oriented excess by 43% without oracle knowledge.

cs.CR↗

Sharpening Tax in Post-Training

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.

cs.AI↗

FedCKA: Representation-Guided Layer Personalization for Federated 3D Perception Across Driving Domains

Robust perception in intelligent vehicles demands 3D object detectors that remain dependable under domain shifts, such as changes in time of day, location, or weather. However, due to costly annotation and rare shifts, some environments lack sufficient data to train a standalone detector. Federated learning offers a privacy-preserving framework for collaborative model training, enabling clients to benefit from shared learning across diverse environments. Yet, this framework traditionally relies on a single global consensus model, which struggles to perform across heterogeneous local data distributions. Local conditions are better captured by adapting a subset of the model, but many personalization approaches rely on predefined layer partitions or fixed personalization ratios, thereby limiting adaptation to client-specific divergence. To reduce this rigidity, we propose FedCKA, a Centered Kernel Alignment (CKA)-based strategy that dynamically handles the personalization-globalization trade-off. Specifically, FedCKA computes layer-wise feature similarities between local client models and the global consensus model during training. By converting layer-wise similarity scores into client-specific aggregation masks, FedCKA selectively shares representation-consistent layers. Evaluation on a unified multi-domain benchmark based on nuScenes shows that FedCKA outperforms established federated baselines, including FedBN, FedRep, and FedSelect, improving average NDS by 7 percentage points over the strongest baseline. The findings offer both a comparative benchmark and a promising direction for robust federated 3D perception across shifts in location, weather, and illumination. Code is available at https://github.com/j-verhoog/FedCKA.

cs.CV↗

GAW-PO: Preference Optimization with Gradient-Aligned Token Weights

Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method achieves the highest average performance among the evaluated preference-optimization methods, improving by 0.97 points over standard DPO and 0.65 points over the strongest competing baseline across 11 benchmarks spanning mathematics, reasoning, coding, and question answering. We further show that gradient-aligned weighting is substantially more robust to aggressive preference optimization: as the DPO regularization parameter $β$ decreases, standard DPO degrades sharply, whereas GAW-PO continues to improve. These results suggest that accounting for the interaction between rejected-token updates and preferred behavior provides an effective form of token-level credit assignment for preference optimization.

cs.CL↗

VoxelSynth3D: Interpretable Volumetric Image-Domain Metal Artifact Reduction with a Paired Synthetic CLINIC-Metal Benchmark

Metal artifacts in postoperative musculoskeletal CT obscure bone-implant and adjacent soft-tissue interfaces. Many metal artifact reduction (MAR) methods require unavailable raw projections or learned models that may shift across scanners and implants. We present VoxelSynth3D, a training-free 3D image-domain framework for reconstructed CT. The framework combines support masking, normalized tissue synthesis, deviation gating, and restricted edge refinement. Detected implant voxels are preserved in the output, while correction targets metal-induced artifacts in the surrounding tissue. We also construct Synthetic CLINIC-Metal, a controlled paired synthetic evaluation resource, from no-metal CTPelvic1K volumes with clean targets, metal/artifact masks, fixed seeds, and patient-level splits; 75 unpaired real metal cases receive qualitative/no-reference evaluation only. The operating point was fixed in a near-flat validation basin. With exact-mask oracle localization, all methods share a metal-excluded tissue ROI. On 40 held-out cases, VoxelSynth3D reduced RMSE from 801.48 to 786.18 HU (paired gain 15.30 HU, 95% CI 11.68-19.23), improving every case and exceeding the evaluated 3D Gaussian smoother by 13.58 HU. Clean-edge agreement decreased next to metal but exceeded input beyond 5 mm. Thus, VoxelSynth3D provides case-consistent within-distribution tissue-error reduction with a localized structural tradeoff. Spacing-aware sensitivity retained aggregate broad-region improvement and identified near-metal calibration as a target.

cs.CV↗

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.

cs.AI↗

How the Audit Rule Shapes Faithful Factor Explanations in LLMs

Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget setting changes the incentive to report factor-level influence truthfully. We formalize the interaction as a verification game and show that proper scoring alone is not enough when auditing depends on the report: report-dependent auditing creates a suppression incentive, because factors reported as important are more likely to be checked and penalized for estimation noise. In contrast, report-independent auditing, or a mixed rule with a small report-independent floor, removes this channel and makes truthful reporting preferable to full suppression. We instantiate the framework with the Counterfactual Brier Score (CBS) and evaluate its predictions on four NLP benchmarks. A synthetic rational agent matches the theoretical prediction exactly, and real LLMs follow the same incentives when they are made explicit. The main design implication is simple: under partial verification, factor-level explanation systems should include a report-independent audit component so that under-reporting cannot be used to avoid scrutiny.

cs.CL↗

FedMIX-P: Mixing Local and Global Preconditioners for Federated Vision and Language Model Training

Adaptive preconditioners accelerate model training, but heterogeneous client geometries can bias federated updates even when gradients are evaluated at the same model. Round-start synchronization alone cannot prevent this mismatch from reappearing during local training. We propose \texttt{FedMIX-P}, which mixes shared and local preconditioners at every local step, retaining local adaptation while reducing mean-squared operator mismatch by a factor of $λ^2$. For smooth nonconvex objectives with stochastic gradients and partial participation, we establish an $O(R^{-1/2})$ stationarity bound using suitable stepsizes and a horizon-dependent mixing weight, without requiring local preconditioners to converge to one another. A two-client counterexample shows that fixed positive mixing can preserve a nonstationary fixed point. The theory covers bounded linear symmetric positive-definite preconditioners. Experiments with SOAP, Sophia, and Muon variants across vision and language tasks show improvements over corresponding local optimizers, including accuracy gains of up to $19.47$ percentage points and lower validation loss for 60M--350M language models. Full nonlinear and momentum-based updates require separate analysis.

cs.LG↗

Tight degeneracy bounds in online Ramsey games

In the $q$-color online Ramsey game, Builder and Painter play on an infinite independent set of vertices. At each step, Builder draws an edge and Painter immediately assigns it one of $q$ colors. Builder aims to force a monochromatic copy of a fixed graph $H$. We prove that, for every $q \ge 2$ and $d \ge 1$, Builder can force a monochromatic copy of any $d$-degenerate graph $H$ while drawing a graph of degeneracy at most $d$. The bound $d$ is tight, and this resolves in the affirmative a problem of Conlon, Fox and Sudakov.

math.CO↗

SuperMotion: Source-Preserving Denoising for Text-Driven Human Motion Editing

Text-driven human motion editing aims to realize a requested change while preserving compatible source content. Existing diffusion editors rely largely on learned conditioning for preservation of the unedited part, yet their outputs can lose temporal detail as denoising proceeds. We propose the \textbf{Source-Preserving Denoising framework (SuperMotion)}, which explicitly reuses the source at each reverse step for source preservation. We first align the source motion to the output timeline and predict a preservation gate that controls reuse across frames and feature dimensions. A clean-space source anchor then utilizes the learned preservation gate to blend the predicted clean motion with the aligned source and passes the corrected estimate directly to the sampling posterior. Because the aligned source is a realized motion rather than a regression output, the anchor injects sample-level temporal detail that a reconstruction-trained denoiser tends to smooth away. To learn effective source reuse, we supervise the anchored estimate against the editing target and match its second temporal differences through a temporal high-frequency loss. These objectives require no explicit edit masks. Extensive experiments show that SuperMotion improves editing accuracy, reaching 33.20\% full-pool R@1 on MotionFix, while reducing temporal-detail attenuation and preserving motion dynamics as it realizes the requested changes. Ablations confirm that the learned preservation gate is responsible for the gain and that it reuses the source to retain the unedited content properly.

cs.CV↗

Who Thinks First? Designing Productive Friction with Engage-to-Unlock GenAI

Generative AI can support writing, but frictionless access may cause cognitive offloading before users develop their own ideas. We introduce Engage-to-Unlock, a productive-friction mechanism that unlocks generative capabilities after users meaningfully engage with the task. In a controlled experiment (N = 398), participants completed a writing task under one of four conditions: Human-Only, Standard Chatbot, Engage-to-Unlock, or Time-Matched Unlock, which matched unlock timing to Engage-to-Unlock participants but independent of users' engagement, then evaluated passages for evidence and inferential errors. Results show that Engage-to-Unlock redistributed effort across tasks: participants spent more time writing and less time evaluating, without increasing overall task duration. They also submitted more prompts than in other AI-assisted conditions and showed the highest accuracy-per-time evaluation efficiency across conditions. These findings suggest that designing GenAI access to encourage early human engagement may provide a productive form of friction, while retaining active AI use and efficient downstream evaluation.

cs.HC↗

Auto-Formalizing Neuro-Symbolic Predictors

Neuro-Symbolic (NeSy) predictors incorporate prior knowledge into the prediction process of neural networks, ensuring that outputs satisfy specified constraints, making them particularly suitable for high-stakes applications where compliance with domain knowledge is essential. A key bottleneck in this paradigm is the acquisition of symbolic constraints: encoding domain knowledge into logical formulas remains a manual and expert-intensive process. In this work, we investigate the extent to which auto-formalization via LLMs can systematically translate textual knowledge into symbolic knowledge that can be plugged into NeSy predictors. To this end, we introduce auto-nesy-bench, a new benchmark for evaluating constraint formalization and its impact on downstream accuracy of NeSy predictors. Through an extensive evaluation across several domains, we find that LLMs can formalize constraints to a meaningful extent, generating formulas that are often similar to those provided by human experts. Moreover, when the generated formulas are syntactically valid, they can lead to high-quality downstream predictions. The code and benchmark are available at https://unitn-sml.github.io/auto-nesy-bench/.

cs.LG↗

Continuous-Process Randomized Compilation for Quantum Process Tensors

Traditional discrete randomized compiling (RC) tailors errors with random control frames at logical-gate boundaries, but fails to capture intra-instrument control evolution or persistent system-environment coupling. In memory-bearing environments, errors correlate across times, not just single gates. We build continuous-process randomized compilation (CPRC) in the continuous process tensor (cPT) framework, unifying random unitary control trajectories, pointwise-covariant finite-duration instruments, and environmental dynamics on one time axis. We study four protocols: fixed Pauli refresh, bounded Cayley paths, unitary Brownian motion, Ornstein-Uhlenbeck driving. We obtain three key results: (1) Pointwise covariance ensures each trajectory exactly reproduces the target logic without system-environment coupling; a closed-frame counterexample proves endpoint-only compensation causes first-order instrument errors. (2) Independent Pauli conjugations over full intervals diagonalize error histories at boundaries, but classical labels retain environment-induced correlations, confirmed by an exact static-bath solution. (3) We derive a total-variation error bound for full adaptive output records under bounded centered coupling and mixing conditions. We prove convergence of trajectory-averaged cPT coefficients in fixed particle sectors for finite non-Gaussian environments and Gaussian baths with time-integrable covariance, including the uncut Drude spectrum. We numerically validate the theory in a qubit model via output probabilities, operator conditional mutual information, and all two-particle coefficients for both qubit and thermal Drude baths. This work extends RC from discrete circuits to continuous-time memory-bearing quantum processes, offering a unified framework, quantitative tools, implementable schemes, and benchmarks for non-Markovian error tailoring.

quant-ph↗

Langevin-Informed Transfer Learning: Replacing Target Samples by Black-Box Feedback

Many scientific and machine learning systems, from molecular dynamics to diffusion models and beyond, are governed by stochastic dynamics with low-dimensional structure, evolving on slow timescales. However, target trajectories, used to identify and interpret such dynamics, are often inaccessible: only biased or static samples that explore the underlying manifold are available. We introduce Langevin-Informed Transfer Learning (LITL), a framework for recovering target Langevin dynamics from biased source samples using only black-box feedback. LITL learns the leading spectral structure of the target infinitesimal generator and the projected drift through Dirichlet representation learning, enabling kinetic reconstruction in spectral form and slow-manifold gradient field estimation. We further introduce a spherical variant well suited to steering normalized latent representations commonly used in learning systems toward desired objectives. We establish finite-sample guarantees for eigenvalue, eigenfunction, and projected drift estimation in Sobolev norms, thereby ensuring generalization of these quantities and their first-order derivatives. Empirically, LITL recovers physical transition timescales from biased molecular simulations, builds kinetic structure from static samples of generative models, reconstructs spherical symmetries of physical systems, and enables post-hoc latent steering of trained neural networks under black-box feedback. Together, these results position spectral operator learning as a practical framework for recovering stochastic dynamics under distribution shift and unlock applications across machine learning and the physical sciences.

cs.LG↗

Inherent Turbulence Immunity of Vector Vortex Beams in Free Space Quantum Key Distribution

Orbital angular momentum (OAM) multiplexing provides an infinite-dimensional discrete Hilbert space ideally suited for high-capacity free-space quantum key distribution (QKD). Nevertheless, pure scalar spatial modes carrying topological charge ($|\ell| \ge 1$) undergo severe decoherence when transmitted through terrestrial atmospheric turbulence. Turbulent refractive-index eddies split high-order vortex singularities, induce catastrophic intermodal crosstalk across adjacent topological channels, and rapidly drive the quantum bit error rate (QBER) well above the unconditional 11% security threshold associated with individual cloning attacks. Here, we demonstrate that hybrid polarization-OAM entangled states, known as vector vortex beams (VVBs), provide intrinsic, hardware-free immunity against turbulent perturbations. Because the optical anisotropy of terrestrial air is exceedingly small ($Δn < 10^{-9}$), refractive-index fluctuations couple symmetrically to orthogonal circular polarization modes as an identical common-mode scalar phase screen that cancels in the relative polarization-phase degree of freedom. By numerically propagating modal fields through modified power-spectral phase screens over turbulence strengths ranging from $D/r_0 = 0$ to $3.0$, we show that the VVB encoding protocol suppresses the asymptotic QBER from 42.0% to 4.8%, yielding an error-suppression factor of approximately 11.6 without requiring active adaptive optics or deformable mirrors.

physics.optics↗

Tridimensional character sums with polynomial arguments and applications

Let $p$ be a large prime and $χ$ a non-trivial Dirichlet character modulo $p$. We study the character sum \[ \sum_{a \sim A} \sum_{b \sim B} \sum_{c \sim C} α(a,b) β(c)χ(f(a) + bc), \] where $z \sim Z$ means $Z\le z < 2Z$, $f \in \mathbb{Z}[X]$ is of small degree $k$, and $\boldsymbolα=(α(a,b))_{a\sim A,b\sim B}$ and $\boldsymbolβ=(β(c))_{c\sim C}$ are two complex coefficients. We prove non-trivial upper bounds for this sum in either of the two cases: (1) $\boldsymbolβ\equiv1$, $k=2,3,4,5$ and $A,B,C>p^{\frac{1}{8}+\varepsilon}$, (2) $\boldsymbolβ$ general, $k=2,3$ and $A,B,C>p^{\frac{1}{6}+\varepsilon}$, where $\varepsilon>0$ is fixed. This work was originally motivated by an intermediate result of Ganguly and Rajan (2023) on counting $2\times2$ matrices over $\mathbb{F}_p$ with irreducible characteristic polynomials, where the entries are in short segments. The new bounds here allow us to count such matrices in much shorter segments.

math.NT↗

Holomorphic expanding maps

We prove that every holomorphic expanding map on a compact connected complex manifold is biholomorphically conjugate to an affine expanding endomorphism of a complex infra-nilmanifold.

math.DS↗

System-Level Gains of Dual-Band RIS: From Wireless Communication to SWIPT

This paper introduces the integration of dual-band reconfigurable intelligent surfaces (DBRIS) in wireless system design, and investigates the ensuing performance gains in different application scenarios. This type of metasurfaces offers a compelling technology due to their reduced material usage, compact design, and multi-functional operation across multiple frequency bands compared to single-band RISs. Two representative use cases are considered to assess the achievable performance gains with the said integration, namely, DBRIS-assisted wireless communication, and DBRIS-assisted simultaneous wireless information and power transfer (SWIPT). For the first design, where the DBRIS is deployed as a relay, we focus on the energy efficiency (EE) maximization of the communication system. An alternating manifold method is proposed for the joint optimization of the source transmit power and the DBRIS phase shifting. For the second design, we investigate a DBRIS-assisted SWIPT system using a novel power splitting approach. An EE metric is defined, and an alternating optimization algorithm is developed to maximize it. Numerical results validate the advantages of the proposed design architectures as well as the proposed optimization framework. In particular, we highlight the energy efficiency gains enabled by the DBRIS as compared to operations with conventional single-frequency surfaces, in both application scenarios.

eess.SP↗