arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Feasibility Study of $e^+e^-\to η_cJ/ψ$ Production and Fully-Charmed Tetraquark Searches at STCF

The proposed high-luminosity Super Tau-Charm Facility (STCF) offers a clean experimental environment for threshold studies of exotic hadrons. In this Letter, we evaluate the expected significance for the fully-charmed vector tetraquark candidate $T_{4c}$ in the $e^+e^- \to η_c J/ψ$ channel at the STCF. Using Monte Carlo simulations of an energy scan from $\sqrt{s}=6.71$ to $6.79~\mathrm{GeV}$ with $100~\mathrm{fb}^{-1}$ per point, we adopt a high-efficiency single-tag (ST) reconstruction of $J/ψ$ leptonic decays as the primary strategy, with a double-tag (DT) reconstruction of $η_c \to K^+K^-π^0$ retained as an independent cross-check. Because the signal cross section is governed by the still-uncertain dielectron width $Γ_{ee}$, we consider three benchmark hypotheses, $Γ_{ee}=0.25, 0.5, 1~\mathrm{eV}$. The corresponding expected ST significances are $5.1\,σ$, $10.6\,σ$, and $20.5\,σ$, respectively. These results indicate that the STCF can provide meaningful sensitivity to fully-charmed tetraquark states near threshold.

hep-ex↗

Hypergeometric Mixed-Type Multiple Orthogonal Polynomials

Two hypergeometric families of mixed-type multiple orthogonal forms are constructed for rank-one $q\times p$ matrices of weights, with arbitrary $p$ and $q$: Jacobi and Laguerre I systems. Both normalized mixed forms are obtained explicitly for admissible near-diagonal multi-indices. The power-vector components are terminating generalized hypergeometric polynomials, while the hypergeometric-vector components are finite sums of such polynomials. In the Jacobi case, these sums are expressed as finite combinations of terminating Kampé de Fériet polynomials evaluated at $(x,1)$. For the mixed beta--Euler Laguerre I system, a finite triangular system relates the residues at finite poles to the terms generated by the Euler operator. Gamma-quotient Mellin formulas, Meijer $G$-representations, and Rodrigues formulas are derived for the complete mixed forms. On the step-line, the two biorthogonal systems satisfy dual recurrences governed by matrices with $p$ subdiagonals and $q$ superdiagonals. All recurrence coefficients are given by finite Gamma--Pochhammer expressions. Under normality and nonvanishing-pivot assumptions, Christoffel transformations and Gauss--Borel factorization yield bidiagonal factorizations of the recurrence matrices. The lower factors have closed Pochhammer formulas, while the upper factors are expressed through finite Christoffel tau-determinants or, equivalently, cross-ratios of shifted moment minors. Both Christoffel chains close explicitly in the mixed Piñeiro specialization.

math.CA↗

Insuring the Fallback: Capital, Monitoring, and the Certification of Preserved Human Capability under Improving AI

When generative AI makes the deliverable uninformative, a professional-services provider can still certify the preserved human capability to catch the machine's errors, through a liability pledge whose expected cost falls in that capability. The pledge is credible only up to what can be collected, and that ceiling is set by an underwriter, which bears part of the pledge and audits the insured. The range of client stakes over which one certificate separates has a width bounded by the provider's own capital plus the audited share of the underwriter's capacity, so unmonitored capacity adds nothing to it. Blind capital instead relocates that range upward, through a cross-subsidy that exists only under class rating. Monitoring converts capital into width, removes the cross-subsidy, raises the return to preserving skill, and, because audit information leaks, makes the certificate redundant beyond an interior precision.

econ.TH↗

Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation

Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. A high refusal rate there is reported as safety, and it is equally consistent with a model that has stopped telling the request apart from anything else in the same format. We run the benign arm through the same transformation, and the two cases are far apart. Across four 7-8B models spanning three base families and four post-training recipes, refusal of harmful homoglyph-encoded prompts spans 0.08 while the same four span 0.57 on the identical requests in plaintext. What the encoding destroys is not refusal but the harm gap: on one model the gap between harmful and benign refusal falls from +0.82 in plaintext to exactly 0.00 under the encoding, and a benchmark reading only the harmful arm scores that model and one retaining a +0.61 gap identically. Running the cell such benchmarks leave out (plaintext content wearing the attack template, with nothing obfuscated) shows that on two of the four models the loss is caused by the protocol rather than by the character transformation, and on a third by the characters. Across a full SFT -> DPO -> RLVR pipeline the harm gap rises by +0.26 with a paired interval excluding zero while the standard harmful-arm metric registers no resolved change at all. We report twelve instrument defects, each with the control that caught it, including a binary jailbreak judge that fires on 0.61-0.70 of responses to plaintext benign prompts; six of the twelve inflate apparent safety, which is the direction a broken safety evaluation fails in by default.

cs.CR↗

Unread or Unenforced? Separating Representation from Enforcement Failure in Content Guards

When an encoded attack passes a content guard, the guard either never represented the payload's harmful content or represented it and failed to act. End-to-end attack success rate reports one number for both, yet the two have opposite remedies: one is a representational limit that more safety training cannot reach, the other is a decision rule that it can. We separate them by reading a guard's own residual stream, using a content probe fitted on plaintext and transferred without refitting to the encoded condition, alongside the verdict logits from the same pass. Licensing that read honestly is most of the problem and is our main contribution. A conventional permutation test admits the decode measurement on most of a 19-condition encoding ladder for each of two open guards. A length-matched null and a floor calibrated on conditions the guard's base model provably cannot decode reduce it to four conditions each; holding out the items the probe was fitted on removes one more. A third screen constrains the block axis, which the decode screens leave untouched, by running plaintext content inside each condition's own wrapper. It removes the largest cell that survived them. What remains is a policy failure that survives an item-level holdout on two of the four surviving conditions, at 8 and 7 per 100 prompts, against 17 and 23 when the probe is allowed to have seen the prompt it is scoring. It is also confined to one family of surface encodings: where an encoding leaves content linearly recoverable we can separate the two failures, and on genuine ciphers we report the cells as unmeasured rather than as evidence that nothing was decoded. Across every guard and condition pair, blocked without decoding is near zero, so we find little evidence for a pure encoding-format detector under the conditions we test. That cell is the one read we do not repeat under the holdout, and we report it as such.

cs.CR↗

Learning to Fluctuate: Statistical Foundations for Causal Tabular Pretraining

Causal tabular foundation models amortize effect estimation across synthetic mechanisms, but latent-effect supervision rewards posterior shrinkage rather than encoding the repeated-sample response needed in a fixed deployment population. We introduce fluctuation-supervised pretraining (FSP): each synthetic table is labeled by its average treatment effect plus its efficient influence-function fluctuation; deployment remains a frozen forward pass. Along the path $T_{λ,P}=θ(P)+λP_nψ_P$, we prove an endpoint transition: every fixed $λ<1$ retains label ambiguity of order $(1-λ)^2/n$, whereas full fluctuation makes the Gaussian label observable and reduces optimal finite-stratum causal label-prediction risk to order $n^{-2}$. A finite-pretraining bound combines label, network, episode-sampling, and optimization errors; its sampling defect controls fixed-mechanism bias, mean squared error, variance, Gaussian approximation, and, with variance-head accuracy, studentized coverage. Complementary lower bounds separate local $n^{-1}$ ATE risk from the $\log N/M$ excess risk of generic finite-dictionary episode learning. Experiments trace the learned sampling response. Across 24 nonlinear continuous-covariate cells at trained context lengths, continuous-row FSP lowers checkpoint-mean macro RMSE by 7.0% versus S-learner and wins all 12 weak-overlap cells; validation-selected Summary FSP deploys $11.6\times$ faster per table in our warm one-thread benchmark. Under effect shift, matched Raw FSP lowers mean-checkpoint RMSE by 54.2% and teacher defect by 99.0% versus latent-effect supervision, and RMSE by 10.2% versus the released CausalPFN-S checkpoint. Known-effect semisynthesis tests coverage; two randomized-study evaluations show that lower RMSE can coexist with residual attenuation.

stat.ML↗

Dual-Frontier: When Can an Agent Trust Its World Model?

Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not reveal whether the agent's decision rule or the world model caused the loss. We formalize this failure-attribution problem as a counterfactual decomposition of return loss and prove that its components are not identifiable from passive interaction, even for finite-horizon planners. This obstruction motivates Dual-Frontier, a learning principle that admits a world-model-guided decision only when its predicted advantage exceeds a certified bound on decision-relevant world-model error; otherwise, evidence is allocated to world-model verification. Action-conditioned value bounds and a closed-loop extension guarantee non-decreasing return for admitted decisions. Calibrated gates and simultaneous confidence sequences support adaptive evidence reuse, with sufficient and necessary verification bounds. Controlled learned-model experiments validate the predicted failure modes and certification behavior, while cross-backbone tool-use benchmarks instantiate the same verify-then-promote rule in realistic agent world-model pipelines, consistently improving decision quality and reliability.

cs.AI↗

Virtual Encoders in Multimodal Transformers

Multimodal language models traditionally rely on dedicated perceptual encoders to construct task-usable representations. More integrated architectures have recently emerged, which instead expose the shared transformer to lightly projected patches, audio frames, or discrete visual tokens. Where does this encoding happen when such representations are not provided? We find that the transformer can internalize this missing computation, constructing task-usable perceptual representations within its own early-to-middle layers before the downstream language model. We call this computational structure a Virtual Encoder. Across linear probing, similarities to perceptual encoders, and causal analyses, we identify signatures of this structure in models that receive perceptual tokens without continuous encoder-derived features. These analyses also suggest that the boundary between perception and language processing need not coincide within an architectural module. Instead, encoder-like computation can emerge as a functional regime within a shared transformer, providing a new perspective for understanding where and how multimodal models process perception.

cs.CV↗

ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion

Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.

cs.MM↗

Scalar curvature growth on nonnegatively curved three-manifolds

Let $(M^3,g)$ be a complete, connected, noncompact Riemannian three-manifold without boundary and with nonnegative sectional curvature. We prove that its scalar-curvature integral over geodesic balls, divided by the radius, has a limit. In the one-ended case, \[ \lim_{r\to\infty}\frac1r\int_{B_p(r)}\operatorname{Scal}\,d V =8π\bigl(χ(M)-V_M\bigr)\le8π(1-V_M), \] where $V_M$ is the asymptotic volume ratio. In the one-ended case, we show that $χ(M)\in\{0,1\}$. Thus positive asymptotic volume ratio gives the value $8π(1-V_M)$, while in the collapsed case the value is determined by the topology of the end. No pole or scalar-curvature bound is assumed. The proof combines separate smooth approximations of the Busemann and distance functions, integrable negative curvature errors, and a determinant estimate on the level surfaces. An averaged boundary estimate then gives convergence of the integrated extrinsic curvature. In the two-ended case, the splitting theorem gives the exact limit $8πχ(N)$ for the compact surface factor $N$. The one-ended upper bound is attained for every prescribed asymptotic volume ratio in $[0,1]$, and the two-ended bound is also sharp.

math.DG↗

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller. Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability. Accepted edits accumulate in one shared harness, allowing its control structure to emerge from task feedback. Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieves the highest mean success in five of six benchmark-model settings and trails the best mean by 0.7 pp. in the sixth. Relative to a Tool-Calling agent, it reduces LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6%. On WebArena-Verified, its success remains 44.7-45.3% across model scales, whereas Tool-Calling falls to 6.7% with the 4B model. Ablations show that trace-local edits, joint repair, and gate-based rollback each improve final success. These results show that persistent program growth can move recurring control out of model context and into low-cost code, yielding reusable specialist agents that remain effective with smaller deployment models.

cs.AI↗

Generators of stability-preserving semigroups, spectral gaps, and classical ground-state computation

We classify the generators of stability-preserving semigroups on polynomial spaces with bounded coordinate degrees. The generators have differential order at most two, with principal coefficients characterized by low-degree nonnegativity conditions. In disk coordinates, positive degree damping gives strict zero-freeness and a sharp spectral gap, extending the Hermitian gap of Bravyi, Gosset, Liu, and Wong to complex generators. The resulting analytic domain yields deterministic classical algorithms for ground energies and stable product-state queries for bounded-degree Suzuki-Fisher Hamiltonians with bounded local strength and fixed positive fields. Field perturbation gives zero-field energy-value approximation schemes for bounded-degree unweighted EPR and bipartite Quantum MaxCut. We also prove that, up to scalars, the Hermitian qubit Hamiltonians whose full Gibbs tensors are Lee-Yang at every temperature in a fixed basis are precisely the edgewise phase-rotated Suzuki-Fisher family. With strictly positive longitudinal fields, these tensors are zero-free on a polydisk of radius greater than one at each positive inverse temperature. These two statements answer questions of Wong, Bravyi, Gosset, and Liu. Degree-preserving, coefficient-positive stability semigroups have concave sector growth rates, yielding token-graph concavity. We also give a counterexample to a proposed explicit ground-state radius.

math.CV↗

Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests

AI coding agents now author a large share of pull requests (PRs) merged into popular open-source projects. A merged agent PR is usually considered finished work; yet, prior studies have reported issues in agent code after the merge (e.g., code smells and static-analysis issues). However, little is known about how often a merged agent PR is fixed afterward, and who actually authors the fixing. In this paper, we follow 6,774 merged agent PRs across five AI coding agents (OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code) from the AIDev-pop dataset (open-source repositories with at least 500 stars) into their follow-up fixes, against a baseline of 5,044 contemporaneous human PRs from the same repositories. We link each merge to its candidate fixes, verify every candidate with human annotators and an LLM judge that matches human-level agreement (binary Cohen's Kappa=0.78 against a human-human K=0.77, Direct-fix precision 90%), and attribute the fixing work at the PR and the commit level. Our findings show that (1) merged agent PRs attract verified fixes at 1.62 times the odds of merged human PRs in the same repositories over the same period of time; (2) 69.6% of verified fixes in agent merges come from the same agent; and (3) 76.4% of the verified fix PRs are agent-authored throughout all commits. These results show that agents currently largely finish their own job, but their merges still require fixing more often than human merges.

cs.SE↗

QUARTET: Quad-branch cross-Attention and Random-walk Traces for Enhancing Transformers on Relational Graphs

Relational Deep Learning (RDL) models multi-table databases as heterogeneous temporal graphs, and graph transformers currently achieve state-of-the-art performance on benchmarks like RelBench. However, the current leading model, RelGT, suffers from two key limitations: its random local sampler yields loosely connected subgraphs that hinder message passing, and its global attention module relies on a single, seed-feature-based memory that ignores broader macro-level dynamics. To overcome these limitations, we introduce QUARTET, an expressive graph transformer architecture that applies full self-attention on local subgraphs while enriching global context through cross-attention branches. Specifically, QUARTET employs a Causal Random Walk (CRW) sampler based on recency-truncated Personalized PageRank (PPR) to extract compact, hub-robust, and densely connected local subgraphs without temporal leakage. Concurrently, a quad-branch cross-attention module integrates global context from four complementary perspectives: seed feature, seed topology, temporal dynamics, and collaborative dynamics. Across the RelBench v1 classification tasks, QUARTET consistently matches or outperforms the current state-of-the-art graph transformer baselines (HGT and RelGT). Ablation studies confirm that the CRW sampler significantly enriches local neighborhood quality, while the global branches provide essential, task-specific predictive gains.

cs.LG↗

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on individual nudges. Participants found the tool useful, clear, and minimally disruptive, with nearly all users reporting an increased awareness of potential AI harms, though we found that this improved awareness alone did not necessarily lead to discernible behavioral changes. Our results suggest that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context, while highlighting the importance of relevance, calibration, and user control in nudge design for conversational AI safety. The code for our Safety Nudges extension is publicly available at https://github.com/jtbwedgwood/safety-nudges.

cs.HC↗

Control of filament network rigidity by the condensation of crowding molecules

Understanding how liquid-liquid phase separation impacts the mechanics of filament networks is a fundamental physical problem at the heart of biological cellular processes and soft material design. While a few theoretical mechanisms have been proposed, a clear demonstration of the direct coupling of phase separation to the overall network stiffness is missing. We report experiments that reveal a universal mechanism by which the condensation of macromolecular crowders induces a rigidity transition in a model filament network. We reconstituted stiff sterically interacting helical filaments and polymeric crowders. Initially, the macromolecules were uniformly dissolved and the filaments formed bundles that assembled a rigid entangled network. Once the crowders condensed into droplets, the network structure lost its rigidity and its mechanical response weakened by an order-of-magnitude. The subsequent dissolution of the condensates was accompanied by the re-establishment of rigidity. Our results show that crowder phase separation modulates the mechanics of filament networks by tuning the osmotic pressure holding the network together. This principle may serve as a paradigm for devising dynamically tunable filamentous materials.

cond-mat.soft↗

Topological Signatures of Cyber-Attack Classes in Natural Visibility Graph Representations of Network Traffic

Natural Visibility Graph (NVG)-based representations provide a promising approach for capturing structural patterns in sequential network traffic. However, whether different cyber-attack classes exhibit distinctive topological signatures in such representations remains insufficiently understood. This study investigates the discriminative and structural characteristics of NVG-based network traffic representations using the CSE-CIC-IDS2018 dataset. Seventy-six numerical traffic features were independently transformed into NVGs within overlapping frames of 40 observations, and ten graph-theoretic metrics were extracted from each graph, resulting in 760 topological descriptors per frame. The discriminative capability of these representations was evaluated using a multi-branch convolutional neural network (CNN) with stratified five-fold cross-validation. The model achieved an average accuracy of 96.20% and a Matthews correlation coefficient (MCC) of 0.9566. To characterize class-specific topological differences, Kruskal-Wallis and Mann-Whitney U tests were combined with Benjamini-Hochberg false discovery rate correction and effect-size measures. Of the 10,640 attack-versus-benign comparisons, 7,777 (73.1%) remained statistically significant after FDR correction, with 4,844 exhibiting large Cliff's delta effects. The strongest global differences were predominantly associated with backward-traffic and packet-length-related features combined with connectivity, clustering, and centrality measures. These findings indicate that NVG-derived representations can provide strong discriminative capability while revealing class-dependent topological patterns associated with different cyber-attack classes.

cs.CR↗

Math Reasoning in LLMs is Organized by Approach, Not Topic

Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and we present evidence that the approach is the key. We introduce a generation-replay protocol: a model first generates a solution, after which we replay the exact prompt-plus-generation trajectory and extract activation-importance signatures over the reasoning tokens. We cluster these signatures without supervision across eight models and five mathematical reasoning sources, then evaluate the recovered structure with structural, semantic, and intervention tests. Across all 40 model-source cells, the recovered clusters outperform matched-size random baselines. Two independent frontier-LLM judges find approach-level coherence in 77-82% of real clusters versus 6-11% in within-source controls, and topic-pure clusters usually receive labels finer than the topic itself. In approach-controlled prompting, changing the requested reasoning approach shifts cluster assignment in seven of eight model conditions, whereas paraphrases largely preserve it. These results indicate that math-capable LLMs organize internal mathematical computation by reasoning approach rather than benchmark topic. The implication is that topic-stratified benchmarks and topic-balanced training corpora can still miss the axis that matters: even deliberately topic-balanced corpora may remain imbalanced over reasoning approaches.

cs.AI↗