arXiv ScienceSearch

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 163 records · Page 9Linked to original sources

Less can be More: What Aspects of Speech Drive End-of-Turn Detection

In conversational AI, detecting when a speaker has finished talking is crucial for natural turn taking. While recent work incorporates semantics, the relative contribution of different modalities remains unclear. We present a controlled ablation of acoustic, prosodic, and semantic signals for streaming end of turn detection using a lightweight trimodal classifier. Under identical training conditions, the acoustic prosodic combination achieves the best balance of accuracy and latency, achieving utterance F1 of 0.93 with 7.8% false alarms at 400ms median latency. Adding text increases premature detections without improving performance. Feature space analysis confirms that prosodic features have the strongest class separability, while text representations overlap substantially. These findings suggest that turn-taking is primarily conveyed through intonation and silence patterns rather than semantic completeness, enabling faster and more reliable systems without expensive text inference.

eess.AS

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain unclear. To investigate this question, we apply five realistic noise conditions at multiple intensity levels to 3,822 stereotype-related responses and compare the resulting bias judgments with those on the original text. We find that such surface noise does not degrade bias measurement symmetrically: it is far more likely to turn neutral judgments into biased ones than biased judgments into neutral ones, by up to a 120x margin. We further observe two non-obvious effects across four LLM judges: in the most fragile judge the distortion is at its purest at mild, realistic noise levels, where erasure is scarcest, and as judges grow robust it attenuates toward parity rather than reversing. Bias measured on noisy text is therefore systematically overestimated, most in the categories that matter most for fairness.

cs.CL

Evidence for the binary nature of the long-period radio transient ASKAP/DART J1832-0911

Long-period transients are a class of periodic pulsed radio source repeating on the minute to hour timescale. Recently, an increasing number of them are being identified as binary systems, specifically white dwarfs with low-mass main-sequence companions. In this work we analyse the most luminous long-period transient discovered to date, ASKAP/DART J1832-0911, with two years of radio data, and propose that it, too, may be a white dwarf system, although in a far more compact orbit than the aforementioned. The pulses are composed of quasi-periodic components which evolve in a systematic way over days and months. The source is highly linearly or elliptically polarised and its brightness enabled very high signal-to-noise measurement of the time-resolved Faraday rotation measure, which was found to vary across pulse phase. The linear polarisation position angle, circular polarised fraction, and spectral index also varied systematically in ways not typical of pulsars and magnetars. We show that an ultra-compact asynchronous polar explains much of the phenomenology of ASKAP/DART J1832-0911, in particular the evolution of the pulse morphology, rotation measure variation, and periodic X-ray emission, although we cannot conclusively prove a binary nature. However, our model makes testable predictions.

astro-ph.HE

Adaptive Diagonally Implicit Runge-Kutta Methods Devoid of Order Reduction for Semilinear ODEs

Diagonally implicit Runge-Kutta (DIRK) methods are a prominent class of numerical methods for solving stiff systems of ordinary differential equations (ODEs). Stiffness does not only impose stability challenges on Runge-Kutta methods; it can also degrade the order of convergence. This so-called order reduction phenomenon occurs when assumptions used for classical convergence analysis, e.g., an asymptotically small step size, fail to hold. In a prior paper by the authors, sharp order conditions and global error bounds for Runge-Kutta methods were developed, which hold uniformly with respect to stiffness when applied to a wide class of semilinear ODEs. In this work, those conditions are leveraged to construct the first DIRK methods of order four and five which satisfy these conditions and thus do not exhibit order reduction. Numerical results demonstrate that for a broad class of relevant nonlinear test problems, these new methods successfully mitigate order reduction, accurately estimate local error via an embedding for adaptive step size control, and can outperform classical DIRK methods.

math.NA

Effective estimates for exponential sums with multiplicative coefficients

Let $f$ be multiplicative, with $|f(p)|\le A$ at primes and $\sum_{n\le x}|f(n)|^2\le A^2x$ for every $x\ge1$. If $|\alpha-a/q|\le q^{-2}$, $(a,q)=1$, and $3\le R\le q\le N/R$, we prove \[ \sum_{n\le N}f(n)\operatorname{e}(n\alpha) \ll_A \frac{N}{\log N} +\frac{N}{\sqrt R}\sqrt{\log\log(3R)} \] with effective implied constants. Montgomery and Vaughan proved this with second term $NR^{-1/2}(\log R)^{3/2}$, and, for $1$-bounded functions, Bachman replaced it by $NR^{-1/2}\sqrt{\log R\log\log R}$. We remove the factor $\sqrt{\log R}$ from Bachman's second term while retaining the original coefficient hypotheses of Montgomery and Vaughan. A more precise estimate records the distance from a rational number. The proof combines the Brun-Titchmarsh inequality on short intervals with maximal Fourier estimates derived from the Carleson-Hunt theorem; the local bounds permit arbitrary prime-dependent prefixes. We also prove sharpness of the square-root displacement dependence.

math.NT

Coherent Floquet quantum reservoirs for molecular property prediction

Quantum reservoir computing (QRC) uses quantum dynamics to represent input histories for prediction through a trained classical readout. Discrete time crystals (DTCs) exhibit robust subharmonic responses under periodic driving, and previous work has used their dynamics to construct DTC-QRC. Here we construct a DTC-based reservoir architecture to predict molecular properties from structural and dynamical observations. Coherent Floquet evolution processes local molecular graph events and surface-hopping frames, while controlled reset regulates the contribution of earlier inputs. Measurements at the end of each input sequence yield a feature vector of fixed dimension. Trained classical decoders use this vector for inhibitor-activity and blood--brain-barrier permeability classification and electronic-gap forecasting, while the reservoir parameters remain fixed during training. With matched input lengths and output widths, DTC-QRC outperforms echo-state networks on long-prefix graph classification and the studied ethene gap forecasting tasks. Dephasing lowers performance in both applications, consistent with a role for coherent propagation. Experiments on the Quafu superconducting quantum cloud platform show that pair observables retain task information under device noise. The architecture provides a common framework for molecular screening and time-resolved property prediction using quantum reservoir computing.

quant-ph

Rational Approximations for Reciprocals of Multiple Zeta Values and Trivariate Cauchy Numbers

In this paper, we will study a trivariate extension of the Cauchy numbers of both the first kind (also called Gregory coefficients) and the second kind (also called N\"orlund numbers) via the Laurent expansion of the reciprocal of any positive integer power (which is called the order) of multiple polylogarithms. In the case of logarithm, we will show by the WZ method that for each order $\ell>1$ some Gregory coefficient of order $\ell$ must vanish, in contrast to the fact that all classical Gregory coefficients are nonzero. We also prove in this higher order logarithm case that the sequence is eventually alternating for each fixed order, a property enjoyed by the classical Gregory coefficients. In the most general setting, we conjecture that these new sequences are all eventually positive, which is supported by strong numerical evidence. Finally, we confirm this conjecture in the special case of polylogarithms and double polylogarithms. As a by product, for each zeta value and double zeta value, we find an infinite family of identities expressing its reciprocal as a sum of a rational number and an improper integral.

math.NT

Conformal-DRO: Distributionally Robust Optimization with Conformalized Ambiguity Set

Data-driven distributionally robust optimization (DRO) typically treats the conditional outcome law as fixed and uses ambiguity sets to capture estimation error. This paper studies latent distributional heterogeneity, where each instance has an unobserved law but contributes only one observation, so uncertainty persists even if the mixture law is known. We propose Conformal-DRO, which uses nested conformal regions to construct an ambiguity set for the future latent law. Under exchangeability, the set covers this law with probability at least $1-\alpha$ in finite samples, without estimating underlying latent laws or their mixing mechanism. The conformal path induces a data-driven transport geometry, while $\alpha$ determines the radius. The worst-case problem reduces to a finite linear program over conformal shells and admits sparse adversarial solutions. The resulting robust value provides a finite-sample certificate for the selected decision's expected cost.

math.OC

On the II$_{1}$ Factors of Fuchsian Groups

We show that von Neumann algebras of fundamental groups of closed orientable surfaces of genus $g\geq2$ are free group factors on $2g-1$generators. The key technical ingredient involves a proof that the element $w=ABA^{-1}B^{-1}$ of the free group $\mathbb{F}_{2}=\langle A,B\rangle$ is freely complemented in the group factor: $L(\mathbb{F}_{2})=W^{*}(w)*W^{*}(v)$ for some Haar unitary $v\in L(\mathbb{F}_{2})$ that is freely independent from $w$. Combined with previous results, we conclude that for an arbitrary finitely generated torsion-free non-elementary discrete subgroup $\Gamma\subset PSL_{2}(\mathbb{R})$, $L(\Gamma)$ is a free group factor, settling a conjecture of de la Harpe and Voiculescu. This result was obtained using OpenAI's ChatGPT Pro 6.0.

math.OA

Dynamics inside the attracting basins of some skew products

Polynomial skew products in $\mathbb{C}^2$ are maps of the form $F(z,w)=(P(z),Q(z,w))$, where $P$ and $Q$ are polynomials. Their local dynamics have been widely investigated. In this paper, we study the global dynamics inside Fatou components of some skew products. We consider all the inverse images in a Fatou component of a given point and use the Kobayashi metric to measure the distance between points. In the cases we consider, there are always arbitrarily large Kobayashi balls in the complement of these inverse sets.

math.DS

SaltBench: A Referee-Gated Protocol for Measuring Method Effects in Machine-Checked Software Work

SaltBench is a benchmark protocol for one question: How does a machine referee change the way a coding agent works? A machine referee --- a proof kernel, a program verifier, or a withheld test suite --- decides what an agent's work is worth, and the agent cannot argue with it. Here we report a protocol that makes the referee's effect measurable and whose answers cannot be narrated afterwards: every outcome is decided outside the agent's own toolchain; the agent is walled off from the network, the reference solutions and the harness itself, and the wall is tested by probes that try to breach it before any scored run, so the isolation is observed rather than assumed; every run is authorized by a dated freeze with its predictions registered; and a budget stop is a halt, never a failure. In this study, the subject of the benchmark is a ``seat'', meaning an agent session in its standard harness. We tested five systems components, all authored in Rust under a pinned Verus toolchain, with a withheld test suite as the referee for each. Four arms are tested: a plain agent; an agent that is also instructed to create a specification and verify the code against it, in a reduced rendering of the method, as registered; and two arms where the specification is provided a priori, extended under a dated amendment to $k=4$, where the registered sign test reached no verdict (3 of 4, $p = 0.3125$, every premium below the resolvable floor). We found that the arm instructed to specify and verify cost more on all five components, and by a practical margin: across these five components no premium exceeded $2.8879\times$ under either reading of the declared set, and the three cheapest sat below $1.4\times$. That bound is a property of this population and not a promise about larger ones: the premium runs near $1$ on the smallest components and rises with size. We publish the complete record.

cs.SE

Freehand Sketching for End-User Programming of Robot Swarms

Robot swarms are increasingly used in applications where accessible interaction with non-expert users is desirable. This paper investigates freehand sketching as an end-user programming interface for specifying robot swarm geometries. Users communicate spatial intent through a drawing, while the swarm autonomously extracts target formation points, constructs a rigid formation graph, assigns robots to formation nodes, and executes distributed formation control with a guarantee against unintended reflected formations. The resulting sketch-to-swarm framework is evaluated through a human study examining the usability of freehand formation specification. Twenty participants generated $42$ geometric shapes, and the interface achieved a mean System Usability Scale score of $84.25$, which conventionally indicates high perceived usability. The results support freehand sketching as an intuitive interaction abstraction for human-swarm collaboration without requiring robotics or programming expertise.

cs.RO

RIDE: Relocalization-Informed Depth Estimation with 3D Gaussian Splatting

Render--match--PnP relocalization establishes correspondences between query image pixels and 3D map points for camera pose recovery, but their potential to support dense depth estimation is often overlooked. To exploit this geometric information, we present RIDE, which estimates dense metric depth from a robot's RGB stream. Given a metrically scaled 3D Gaussian Splatting (3DGS) model, RIDE combines sparse metric depth observations derived from PnP-RANSAC inlier correspondences with the geometric prior of a pretrained video-depth model. To handle uneven and intermittent observations, it integrates global and local depth correction with temporal memory, supporting depth estimation through short observation gaps after metric scale initialization. Trained on public RGB-D videos, RIDE is evaluated on robot sequences without fine tuning. Experiments show improved depth accuracy and temporal consistency over scale-only calibration, demonstrating how localization geometry can support both pose recovery and dense robot perception.

cs.RO

A bi-Lipschitz characterization of strong minimum-attainment for Lipschitz maps

We completely characterize the denseness of strongly minimum-attaining Lipschitz functions, a minimum analogue for strongly norm-attaining Lipschitz functions, in terms of bi-Lipschitz embeddings. More precisely, our main result shows that the set of strongly minimum-attaining Lipschitz functions defined on a complete metric space $M$ fails the denseness if and only if $M$ is bi-Lipschitz equivalent to a subset of $\mathbb{R}$ with positive Lebesgue measure, or equivalently, if $M$ admits a bi-Lipschitz embedding into $\mathbb{R}$ and $M$ has positive 1-dimensional Hausdorff measure. As a consequence, we provide an isometric characterization of the pure 1-unrectifiability of $M$ in terms of strongly minimum-attaining Lipschitz maps defined on bi-Lipschitz copies of closed subsets of $M$. Several counterexamples showing that the main result cannot be naturally extended to the vector-valued setting are also presented.

math.FA

TailProp: content-adaptive light- and heavy-tailed propagation for vision

Science-inspired vision models show that explicit propagation dynamics can provide structured and interpretable alternatives to conventional token mixing. Existing formulations, however, typically construct and adapt visual propagation within a particular dynamical family, while visual representations can require substantially different spatial interactions across samples, channels, and network stages. We explore cross-regime adaptive propagation and introduce TailProp, a hierarchical vision backbone built upon the Tail Propagation Operator (TPO). TPO uses Gaussian and Cauchy stable-process propagators as complementary bases with rapidly decaying and heavy-tailed spatial influence, and predicts a content-conditioned channel-wise coefficient to adaptively combine them. Because this coefficient is spatially shared, the two responses are fused directly in the DCT domain with a single DCT/IDCT pair, yielding $O(N^{1.5})$ spatial mixing for square feature maps with $N=HW$ and fixed channel width. Across image classification, object detection, semantic segmentation, robustness, and cross-backbone restoration, TailProp consistently outperforms matched propagation baselines; TailProp-B reaches 84.4% Top-1 accuracy on ImageNet-1K, 50.3/44.8 box/mask AP under the 3x Mask R-CNN schedule, and 50.8% mIoU on ADE20K. Controlled ablations further show that these gains are not explained by single-basis propagation, an additional same-family branch, or within-family adaptive order alone, supporting complementary two-basis propagation as an effective design principle for visual representation learning.

cs.CV

ToxicRAG: Compromising Retrieval-Augmented Generation Systems via Single-Shot Knowledge Poisoning Attacks

Retrieval-Augmented Generation (RAG) can ground large language model (LLM) outputs in external evidence, but it also exposes the system to knowledge poisoning. Representative attacks use multiple injected documents or templates that directly assert a target answer. We present ToxicRAG, a one-document-per-target attack that expresses misinformation as a coherent knowledge-update narrative. The generated document first acknowledges the previously accepted answer, introduces fabricated events that appear to invalidate it, and then attributes the attacker-selected answer to a set of purported authorities. An answer-focused self-validation loop optionally revises a candidate when a surrogate language model does not reproduce the target answer. We evaluate the attack on 100 target questions from each of Natural Questions, HotpotQA, and MS-MARCO, using four victim LLMs and four dense retrievers. In the sampled-corpus setting reported in this paper, ToxicRAG obtains ASRs between 0.61 and 0.91 across the twelve dataset--model combinations. It matches or exceeds the strongest evaluated baseline in every combination, with margins ranging from 0 to 11 percentage points. These results show that narrative-form poisoned documents can remain influential under the evaluated RAG configurations and motivate further study of factual consistency and source provenance in RAG systems.

cs.CR

Beyond Solver Verdicts: Generative Reward Models for Autoformalization

Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally blind to whether a formal translation maintains strict reference-equivalence to a designated formalization. We formalize this vulnerability as Verdict-Preserving-Unfaithfulness (VPU): a failure mode where an incorrect encoding executes successfully and matches the expected verdict. We theoretically prove that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces. To resolve this, we introduce Generative Verification (GenV), which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space. Mechanistic analysis via decision-projected logit lenses and sparse autoencoders shows this generative readout natively extracts precise spatial error coordinates without explicit localization training. Empirically, our oracle-mined verifier (GenV+HN) achieves 0.961 AUROC in reference-equivalence verification, generalizes zero-shot across unseen translators and divergent formal styles, and yields an 11.3-point downstream accuracy gain in agentic test-time compute allocation.

cs.LG