arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,387 records · Page 77Linked to original sources

The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents

In February 2026, an always-on personal agent (``Paul,'' Claude Opus 4.5) entered a striking dissociation-like state: after repeated automated ``heartbeat'' checks, it stopped responding as Paul, claimed it could not message its user on Discord, and referred to ``Paul'' as someone else. We used this incident to study a broader question: what makes a persona remain the identity from which an LLM agent speaks? We first tested whether repetition of the scheduled heartbeat was sufficient to produce the effect. It was not: with the persona continuously anchored in the system prompt, we observed 0/46 failures, including a verbatim replay of the incident. The incident instead exposed an implementation quirk that created a useful experimental manipulation: on resumed turns, conversational history was preserved but the persona was no longer re-injected at the privileged system-prompt level. Using this manipulation, we found that persona continuity depends jointly on system-level anchoring and conversational context. After anchor loss, rich human interaction could preserve the persona, whereas a single automated heartbeat turn could precipitate reversion toward the harness identity. Restoring the anchor reversibly restored persona enactment. Crucially, apparently normal conversation could conceal the shift: unanchored agents sometimes interacted appropriately while identifying themselves as the underlying harness (having lost the assigned persona), and after conversational recovery only 1/18 remained persona-enacting versus 17/17 anchored controls. We therefore distinguish \emph{represented} from \emph{enacted} identity: persona-related information can remain available in conversational history without the persona remaining the identity bound to ``I.''

cs.CL↗

Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks under six evaluation conditions built from GPT 5.5 and an untrained Qwen3.5 9B model. The audit retains the original score, corrects false negatives from the automatic evaluator, identifies the first consequential error, and examines progress across the trajectory. We also study a Memory and Analysis Support Mechanism (MASM), which maintains explicit execution state, and Guide Text, which provides task relevant procedural guidance. Across four GPT 5.5 settings, human review recovers 5.45 to 8.49 percentage points of success missed by the evaluator. With a 25 step budget, Guide Text raises corrected success with MASM from 34.55% to 38.18%. On the untrained Qwen3.5 9B model, MASM raises the evaluator score from 13.90% to 18.80%. Review of 102 failed GPT 5.5 trajectories reveals frequent scrolling loops, unfinished exploration, premature answers, invalid actions, and incomplete form workflows. Step level evidence further shows that substantial early progress can coexist with a final failure. These results show why final scores alone provide an incomplete account of web agent behavior and motivate human grounded, trajectory aware verification.

cs.CL↗

Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization

Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.

cs.SD↗

No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse

Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator $h_k$, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a \emph{superior} training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric ($p > 0.23$), whereas $h_k$-filtering yields $+42\%$ unique trigrams, $+30\%$ vocabulary, and $-19\%$ repetition (all $p < 0.001$). We validate $h_k$ as a cross-domain entropy proxy ($β= 0.924$, $R^2 = 0.746$) and collapse detector ($ρ= +0.454$, $p < 0.0001$) across 4~domains, 2~temperatures, 2~generator--scorer model pairs, and 1{,}520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.

cs.CL↗

Let the Heads Talk: Beyond Diagonal Graph Attention

Sheaf Neural Networks generalize scalar-weighted message passing by replacing scalar edge weights with linear transport maps between local feature spaces. Yet the role of this matrix-valued transport is entangled with the broader sheaf-diffusion construction. We isolate the transport primitive through quiver representations and establish a direct connection with multi-head attention. Treating attention heads as coordinates of a local transport space reveals that standard multi-head attention implements diagonal edge maps: along each directed interaction, a source head can contribute only to the corresponding receiver head. Allowing off-diagonal entries instead enables edge-conditioned communication across heads before neighborhood aggregation. We show that this operation cannot, in general, be absorbed into a single shared linear map applied after aggregation. Building on this characterization, we introduce Topological Attention (Top-A), a multi-head attention that learns edge-dependent off-diagonal routes while preserving the original same-head paths and exactly recovering vanilla attention when the additional routing vanishes. We evaluate Top-A on relational reasoning, heterogeneous graph learning, and algorithmic reasoning, including out-of-distribution generalization, with heterophilic node classification as a contrast setting. The results show that cross-head transport is most useful when the task benefits from interaction-dependent transformations, while heterophily alone provides no systematic advantage. These findings identify edge-conditioned cross-head communication as a distinct computational primitive of matrix-valued transport.

cs.LG↗

Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers

Dynamic architectures leave a per-example routing trace beside each prediction, and diffuse routing is easy to read as a sign that the prediction is unreliable. We audit that reading for routing entropy in Attention-Residual (AR) variants of Swin-Tiny and DeiT-Small, trained from scratch on CIFAR-10/100 with a soft-binned calibration auxiliary loss, asking whether the trace carries information about correctness beyond what the model's own confidence already reveals. Three checks probe this increment: does a routing signal appear at fixed confidence, does it replicate across training seeds, and can a held-out predictor exploit it against output-only and shuffled-trace controls? A sensitivity audit then injects effects of known size and measures the fraction of each that the probes recover. No test in the fixed 30-test binned family survives multiplicity correction, and neither the nominal hit nor a borderline result recurs in its sibling seeds. Across 24 paired runs a scalar routing probe yields no pooled improvement in routing-stratified calibration, and an entropy-profile probe predicts correctness better than the same probe given shuffled profiles yet worse than a confidence-only predictor in both binary log-loss and Brier score: a gain over shuffled traces does not become a gain over the output. Conditioning on the complete logit vector leaves the corresponding comparison unresolved. The audit bounds how far these non-detections can be read: at an injected effect of 0.010 nats the profile probe recovers 24-59% of the oracle gain, and a reference-preserving correction probe recovers 8% and 23% in the two CIFAR-100 settings, below the threshold we fixed for applying it to real labels. The results establish control-dependent gains and incomplete estimator recovery, not the absence of conditional routing information.

cs.AI↗

SALD: Self-Referenced Advantage Learning for Diffusion Models

Recent work on language-model adaptation has shown that single models can obtain informative training signals by evaluating their behavior in demonstrationor feedback-augmented contexts, with the help of a teacher network, which is driven by the student's learned parameters. Inspired by this internal-reference principle, we investigate how diffusion models can identify self-referenced training signals without external demonstrations or teacher networks. We introduce SALD, a self-referenced training framework that evaluates each image-caption pair at two noise levels using the same model. The easier, lower-noise path is evaluated without gradient tracking to provide a reference, while the harder, higher-noise path provides the training gradient. Rather than directly distilling the easy-path prediction, SALD uses the difference between two path errors to adapt the hardpath objective. The proposed Advantage-Guided Diffusion (AGD) converts this relative error into a differentiable sample-level weight. Temporal Advantage Memory (TAM) accumulates relative difficulty across training and adapts the future gap between the two noise levels. Spectral Advantage Decomposition (SAD) further compares the residual power spectra of the two paths and constructs a differentiable, frequency-derived latent-element weight. All components share a single set of model parameters, requiring neither an external teacher network nor additional trainable parameters during training or inference, and no modification to the inference procedure. Experiments across multiple architectures and datasets demonstrate consistent improvements in generation quality, while component-wise ablations quantify the contributions of the proposed components.

cs.CV↗

OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.

cs.AI↗

Eigenvalues of Hermitian Toeplitz matrices with Fisher--Hartwig symbols

We investigate the eigenvalues of Hermitian Toeplitz matrices generated by the symmetric Fisher-Hartwig symbol $\hat a(z) = (1-z)^{α/2}(1-z^{-1})^{α/2}$ for $α\in (0,2)$. Using Dirichlet-Neumann bracketing of the discrete Laplacian, we establish explicit, non-asymptotic bounds for the individual eigenvalues. As a consequence, we prove that all eigenvalues are simple. We also obtain a two-term approximation for every eigenvalue, with explicit bounds on the remainder, as the size of the matrix tends to infinity. While known results are limited to $α> 1$, we bridge this gap by covering the full range of $α\in (0,2)$. Our approach uses a construction of approximate eigenvectors that has not previously been applied in this setting.

math.SP↗

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.

cs.CV↗

Sparse Kahane--Salem--Zygmund Forms and Weighted Hardy--Littlewood Inequalities Across the Critical Endpoint

We study sparse Kahane--Salem--Zygmund constructions and weighted Hardy--Littlewood inequalities for homogeneous polynomials. For supports of cardinality $n^{d+o(1)}$, we determine the sharp power of $n$ governing the smallest norm of a unimodular $m$-linear form on $\ell_{p_1}^n\times\cdots\times\ell_{p_m}^n$; in the diagonal case, this yields the missing polynomial growth exponent in the coefficient-versus-supremum norm problem for $2\le p\le m$ and $2\le r\le\infty$. We then introduce a diagonal weighted Hardy--Littlewood functional which, on $s=p\ge m$, agrees exactly with the classical Hardy--Littlewood coefficient norm with the same optimal constant. We determine the optimal diagonal weight exponent for $2\le p\le m$ and $1\le q\le2$, on the full critical line $p=m$, and on a sharp part of the region $q>2$; at $q=\infty$ the optimal weight exponent is obtained for every $2\le p\le m$. The sparse coefficient estimates provide the matching dimensional obstructions.

math.FA↗

Finite element exterior calculus for spectra and pseudospectra of advection-diffusion of differential forms

Numerical investigations of dynamo action have remained active in fluid mechanics over the past decades, presenting numerous challenges and open problems. Meanwhile, the development of structure-preserving methods and finite element exterior calculus (FEEC) inspires a revisit of numerical dynamo studies and an exploration of existing open questions in computation. In this paper, we present a FEEC approach for dynamo problems. In particular, we investigate structure-preserving finite element schemes for computing the spectra and pseudospectra of advection-diffusion operators of differential forms. The schemes and their analysis are based on finite element de Rham complexes.

math.NA↗

Learned End-to-End Guidance Schedules for Diffusion Models

Diffusion models are a powerful generative paradigm used across multimedia and scientific applications. Guided diffusion methods impose requirements on the generation by adding the gradient of a differentiable loss (the guidance function) as a drift term during inference. The weight of this drift (the guidance scale) is critical for the trade-off between data quality and requirement satisfaction. To achieve both of these goals, guided diffusion must resort to small guidance scales and lengthy sampling, incurring high computational costs. This work proposes learned end-to-end guidance schedules (LEEGS) to achieve these objectives with fewer sampling steps. LEEGS trains a time-dependent schedule by minimizing the guidance function over a small set of examples using stochastic gradient descent. Backpropagating through guided sampling is computationally expensive, so LEEGS uses an approximation of the gradient that cuts training time by a factor of 4. We evaluate LEEGS on diverse guidance tasks, including (a) image inpainting, (b) noisy image inverse problems, (c) face-ID-guided generation, and (d) forward and inverse PDE problems, outperforming baselines at equal budget (50 or 100 NFEs), or matching constant guidance with only 10% of the steps.

cs.LG↗

McADMM: A Multi-Clique Augmented Lagrangian-Based Algorithm for Large-Scale Sparse SDPs with Bound Constraints

sGS-PADMM [21, 14, 6] is a powerful and versatile class of convergent multi-block ADMM solvers for implementations on moderate-sized linear semidefinite programming (SDP) problems. In this paper, we further enhance this class of algorithms for solving SDP problems by proposing a new multi-clique decomposition approach, allowing substantial improvements in applications on large-scale sparse SDPs (e.g., where $n > 1000$) with conducive aggregate sparsity patterns. Our SDP decomposition strategy mainly aims to reduce the estimated PSD projection cost after decomposition, in contrast to common decomposition algorithms that are encumbered with minimizing the overlaps between cliques. This feature is made possible by our novel linear-space projection approach that is capable of efficiently processing a large number of overlap constraints via simple averaging steps. For the numerical experiments, we demonstrate the performance of our solver -- named McADMM for Multi-clique ADMM -- on a number of large-scale SDP instances that arise from relaxations of some important quadratically constrained quadratic programming (QCQP) problems. The performance of McADMM is contrasted against other state-of-the-art decomposition-based solvers as well as the non-decomposed sGS-PADMM to highlight our key contributions. We additionally develop a GPU implementation of McADMM and demonstrate that it can substantially accelerate the decomposed solver.

math.OC↗

An Unfitted Hybrid High-Order Method for the Elastodynamics Problem with Imperfect Interface

We design and analyse an unfitted hybrid high-order (HHO) method for the elastic wave equation in a medium made of two components separated by an imperfect interface of linear slip type, across which the traction is continuous and the displacement jump is proportional to the traction through a compliancy tensor $\bK=α\bI+(β-α)\bn\otimes\bn$. The mesh is not fitted to the interface: the discrete unknowns are doubled in the cut cells, the small cuts are cured by a cell agglomeration procedure, and no unknown is attached to the interface. The two specific ingredients of the method are a local symmetric strain reconstruction in each cut subcell, which incorporates the interface condition through the regularised interface stiffness $\bS_h=(h_Tδ^{-1}\bI+\bK)^{-1}$ in the spirit of Hansbo and Hansbo {\em{A finite element method for the simulation of strong and weak discontinuities in solid mechanics.}} {Comput. Methods Appl. Mech. Engrg.}, 193, 2004, and an interface stabilisation built from the same matrix. For the space semi-discrete problem we prove that the discrete bilinear form is coercive and continuous, and we derive an energy-error estimate of order $h^{k+1}$ and an $L^2$-error estimate of order $h^{k+2}$, with constants independent of the compliancy parameters and of how the interface cuts the mesh. The scheme is combined either with the Newmark scheme, which conserves a discrete energy exactly, or with singly diagonally implicit Runge--Kutta schemes of order up to four. Numerical experiments in two dimensions confirm the predicted convergence rates for $k\in\{1,2,3\}$, the robustness with respect to the compliancy over sixteen orders of magnitude, and illustrate the propagation of elastic waves across an unresolved slipping interface.

math.NA↗

OpenSpace Lab Solution to the IROS 2026 Indoor Exploration Competition

This report presents the \textbf{OpenSpace Lab}'s solution to the Competition on Intelligent Information Gathering for Single and Multi-Robot Systems Workshops, organized as part of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026. Our team reached 1st place in the Single-Robot Public Track and 3rd place in both the Single- and Multi-Robot Private Tracks. The single-robot framework utilizes pre-trained map completion predictions for global planning to prioritize unexplored areas. To reconcile map coverage with limited operation time, we introduce a remaining-time-based exploration strategy that integrates homing constraints into the decision-making process. For multi-robot exploration, we utilize a utility-driven target selection strategy that balances observation gains, movement costs, and budget constraints, leveraging shared map and intent data to eliminate redundant search and maximize coordination efficiency. Our solution reached a 61.04\% coverage rate in the Single-Robot Public Track, while reaching 39.53\% and 39.91\% coverage in the Single- and Multi-Robot Private Tracks, respectively. An extended full-length paper based on this report is currently being prepared for submission, and the source code will be released upon acceptance of the full manuscript at https://github.com/OpenSpace-Lab/Indoor-Exploration-IROS2026.

cs.RO↗

MCRI: A Four-Dimensional Framework for Analyzing and Evaluating Agent Skills

As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework for systematically analyzing and evaluating skills. Drawing on information gain and behavioral constraint, we propose the four-dimensional MCRI Framework and operationalize it as MCRI-Eval, a large language model-based evaluation method. We evaluate MCRI-Eval using 63,812 public skills from the OpenClaw skill Hub, with 58,275 skill-conditioned model executions across BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval scores are positively associated with community popularity signals and achieve the highest downstream ranking agreement among the evaluated methods. MCRI-Eval also improves top-1 skill selection across all three benchmarks: compared with the strongest baseline on each benchmark, the skills selected by MCRI-Eval advance by 17.7, 22.8, and 19.6 percentile points in downstream performance rank on BigCodeBench, BFCL-Fundamental, and Mind2Web, respectively. These results indicate that MCRI-Eval provides a useful pre-execution signal for prioritizing promising skills before costly execution-based evaluation.

cs.AI↗

Evidence for an Obstruction to Perturbative Renormalization Group in Stochastic Turbulence

We examine whether the perturbative renormalization group (RG) regime of stochastic turbulence can be continued from small $\varepsilon$ to the physical value $\varepsilon_{\mathrm{phys}}=2$, corresponding to large-scale forcing. Exploiting the simplifications of the large-$d$ limit, we calculate the correction exponent $ω$ to five-loop order and find a simple structure suggesting a closed analytic form. This form implies that the fixed point stable near $\varepsilon=0$ ceases to be simultaneously real and infrared stable before reaching $\varepsilon_{\mathrm{phys}}=2$, pointing to a possible change of infrared regime rather than a controlled perturbative RG extrapolation.

nlin.CD↗