arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning

In cross-lingual zero-shot text-to-speech, the accent of the reference leaks into the target speech. We propose accent analogy guidance (AAG), a training-free sampler term that subtracts an accent direction estimated from the model's own predictions for one synthetic voice rendered in both languages, so the voice cancels and only the accent remains. By a blind LLM accent judge on real dubbing data, reweighting classifier-free guidance between reference and text, and its variants, stay near one identity-accent trade-off curve; we score a method by its speaker similarity above that curve at equal accent ($Δ$SIM). Across four open TTS models AAG lies above the curve: on OmniVoice $Δ$SIM is +0.11 to +0.27 on three test sets (accent 3.51 to 4.28 on a 1-5 scale at speaker similarity 0.29, where reweighting keeps 0.02); MaskGCT and CosyVoice 2 also lie above their curves, and on F5-TTS it is more native than any reweighting setting. An LLM-free language-ID measure and a twelve-listener panel agree. A premise test and the reach of a model's own curve indicate in advance whether and roughly how much AAG can gain, predicting the one model where it gains nothing (X-Voice).

cs.SD↗

Asymptotic behavior of twisted Alexander invariants for hyperbolic knots with at most six crossings

Let $K$ be a hyperbolic knot and let $ρ_n$ be the $n$-dimensional irreducible representation induced from a lift of its holonomy representation. Motivated by Goda's asymptotic volume formula and the complexified Volume Conjecture, we study whether the higher-dimensional twisted Alexander invariants associated with $ρ_n$ detect the complex volume of the knot complement. We compute $$ \fracπ{2} \log \left( \frac{A_{K,n-2}(1)A_{K,n+2}(1)} {A_{K,n}(1)^2} \right) $$ for all hyperbolic knots with at most six crossings. Our numerical experiments indicate that these values approach $$ \operatorname{Vol}(S^3\setminus K) +i\,2π^2\operatorname{CS}(S^3\setminus K) $$ modulo $iπ^2\mathbb{Z}$. Based on these computations, we propose a complexified analogue of Goda's asymptotic volume formula.

math.GT↗

FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection

In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enhancement, while the cross-modal interaction patterns of different frequency components remain insufficiently explored. Moreover, spectral discrepancy itself may contain both useful complementary cues and unreliable modality-specific responses, making indiscriminate frequency fusion suboptimal. To address these issues, we propose FoCal, a frequency-oriented framework for aerial RGB--IR object detection. First, a Frequency-Aware Dual-Domain Calibration (FADC) module is developed to explicitly model frequency-dependent cross-modal interaction. Low-frequency components are collaboratively consolidated into a shared structural consensus, whereas high-frequency components preserve modality-specific information through selective cross-modal exchange. The resulting frequency-aware cues are further transferred to the original feature domain to regulate cross-modal calibration. Second, we introduce a Discrepancy-Guided Spectral Modulation (DGSM) module, which characterizes cross-modal spectral imbalance using confidence-weighted relative amplitude discrepancy and transforms it into a bounded signed gate for adaptive enhancement, preservation, or attenuation of the joint multimodal spectrum. Extensive experiments on DroneVehicle, ESCVehicle, and ATR-UMOD demonstrate the effectiveness of FoCal, yielding $\mathrm{mAP}_{50}$ values of 83.5\%, 54.8\%, and 64.6\%, respectively. Meanwhile, with only 3.0M parameters, FoCal achieves 113.6 FPS while preserving leading detection accuracy, highlighting a favorable accuracy--efficiency trade-off. Code is available at {https://github.com/universeliang/FoCal.

cs.CV↗

The Fly That Stopped: Mushroom-Body-Inspired Habituation as a Reward-Free Scheduling Prior for Autonomous Penetration Testing

Autonomous security-testing agents can spend much of a fixed action budget repeating earlier tool selections. We evaluate a reward-free scheduler inspired by mushroom-body novelty processing in Drosophila. It combines sparse state encoding with decaying habituation counters over structural URL classes and tool families. The counters penalize repeated clean or error outcomes without updating weights from scalar reward. Four matched campaigns motivated this design by exposing reward-accounting errors and tool-failure loops; reward-driven components did not improve the tested primary outcomes over the reward-free MB condition. A pre-registered pilot and two confirmatory stages then evaluated repeated (tool, URL) selections. In the second confirmatory stage, 8 of 10 screened lab targets remained measurable after two error-heavy slow-XSS exclusions. The habituation-enabled scheduler lowered duplicate-action ratios in all 6 non-tied target pairs (exact one-sided p=0.015625), with two ties; the largest reduction was 51 to 18 duplicate steps within a 60-step budget. This is evidence for the complete scheduler on the measurable budget-hold population, not an isolated habituation ablation or a vulnerability-discovery gain. A complementary study on a 13,498-neuron MaleCNS-derived circuit (501,267 synaptic edges with weight at least 5) found no action selectivity from the five tested local-plasticity approaches under a fixed readout; readout plasticity produced qualified positive results in synthetic tasks without establishing a biological-topology advantage. We report the population bounds, remaining input-integrity dependencies, and an internal AI-assisted review protocol alongside the results.

cs.CR↗

Quandle coloring quivers of pretzel links

In this paper, we conduct a systematic study of quandle colorings and quandle coloring quivers for pretzel links using the dihedral quandle $\mathbb{Z}_{n}$. First, we systematically investigate all possible colorings of 3-pretzel links, determining the number of distinct colorings in each case as well as the structure of their quandle coloring quivers. In order to obtain more general conclusions, we impose restrictions on $n$ based on the properties of the coefficient matrix of the system of congruence equations. So we examine the number of quandle colorings and the quandle coloring quivers for 4-pretzel links in the case where $n$ is prime. Finally, Combining the results of 4-pretzel links we rigorously derive both the coloring numbers and quandle coloring quivers for general $m$-pretzel links in the case where $n$ is prime, with full proofs provided.

math.GT↗

Bio-inspired efficient cyclostationary analysis in machine and underwater acoustic recordings

We propose a bio-inspired approach that uses the inner-hair-cell (IHC) response of the Cascade of Asymmetric Resonators with Fast-Acting Compression (CARFAC) model to efficiently extract cyclic modulation from acoustic signals. We further investigate the contribution of IHC processing by comparing the CARFAC-IHC response with the CARFAC basilar-membrane (BM) filtering. Furthermore, the CARFAC-IHC and CARFAC-BM approach are benchmarked against conventional FFT Accumulation Method (FAM), Integrated Cyclic Modulation Coherence (ICMC), and Detection of Envelope Modulation On Noise (DEMON) approaches using the Case Western Reserve University (CWRU) bearing dataset and a real ShipsEar work-vessel recording dataset. The results demonstrate reliable recovery of characteristic cyclic components while substantially reducing the computational burden of conventional cyclostationary analysis.

eess.SP↗

ClaimMirage: When Self-Claims in Domain Names Change LLM Threat Judgments

Short claims such as not-phishing or official can change how a large language model (LLM) judges a domain name, without explicit prompt-injection commands. We study this manipulation as ClaimMirage: a name under inspection claims its own safety or approval. We analyze 622,080 judgments across 64 brands and five LLMs, comparing ten claims with length- and hyphen-matched controls in constructed brand-like names. Self-claims can substantially reduce or increase alerts, depending on the LLM and input setting. In one setting, risk-denial terms inside the registrable name reduce alerts by 45.3 percentage points even with a basic safeguard: the prompt supplies the potentially impersonated brand and its official domain for comparison. Without these references, endorsement terms at that position increase alerts by 65.6 points in the same LLM. References and component annotation remove some alert reductions but leave others or make them larger. These findings motivate testing resistance to self-claims and seeking independent evidence before treating a domain name under inspection as safe or authorized.

cs.CR↗

Tag-Aware Structured Text Translation: Towards a Systematic Understanding

Internet texts are replete with format tags that carry structural, semantic, and functional meaning. Current large language model (LLM)-based translation systems struggle to balance translation fluency with tag fidelity when processing tagged text. We argue that resolving this tension requires a systematic approach at three interconnected levels: data synthesis, capability building, and multi-objective alignment. At the data level, we identify and formalize a fundamental trade-off between structural tag diversity and translation naturalness in synthetic data generation; existing methods optimize for one at the expense of the other. We propose a hybrid synthesis strategy (Hy-LST) combining LLM-based synthesis tag method and Two-Stage LLM-based synthesis tag method to produce both diverse and natural tagged data. At the capability level, we decompose tag-aware translation into four sub-tasks of increasing difficulty in a multi-task supervised fine-tuning framework, enabling targeted capability acquisition and knowledge transfer. At the alignment level, we design three complementary reward functions under a group relative policy optimization framework, each targeting a distinct objective (fluency, tag fidelity, and tag-scoped translation quality), and show that joint optimization consistently outperforms single-reward alternatives. Experiments on six language directions (en2zh, en2ja, en2de, en2fr, en2ru, de2fr) demonstrate that each level contributes measurable improvements, and the complete system significantly outperforms existing methods. Qualitative analysis reveals specific error patterns and their mitigation after training with our method.

cs.CL↗

A Tendon-Driven Robotic Jellyfish with Constrained Soft Actuation and Depth Control via Reinforcement Learning

Jellyfish-inspired robots offer a compliant and efficient approach to underwater locomotion, but achieving large deformation together with repeatable actuation and closed-loop control remains challenging. In this work, we present a tendon-driven robotic jellyfish with constrained soft actuation. Each actuator combines a flexible substrate with discrete constraints, enabling bending up to \(150^\circ\) with an approximately linear tendon displacement-bending relationship. Eight actuators driven by four servos allow the robot to perform stable swimming, attitude adjustment, and self-righting. Based on the linear actuation, a reinforcement-learning controller is further developed, enabling closed-loop depth regulation in both simulation and physical experiments. These results show that mechanical constraints can improve the controllability of soft actuation while preserving compliant jellyfish-like motion, providing a route toward manoeuvrable and autonomous jellyfish robots.

cs.RO↗

Interpreting relative utility for probabilistic predictions

At a fixed threshold, relative utility (RU) measures the net-benefit gain of a prediction model over the better of treat-all and treat-none relative to the corresponding gain under perfect outcome classification. We illustrate that RU equal to 1 therefore represents perfect outcome classification at that fixed threshold, not perfect probabilistic prediction. However, even when every predicted probability equals the true probability, observed RU can equal 0. In a simple constant-risk setting, this occurs with probability approaching 1 as the sample size increases. Consequently, the distance from observed RU to 1 should not in general be interpreted as improvement achievable by a better prediction for binary probabilities.

stat.AP↗

Global boundedness in a Chemotaxis-May-Nowak model for virus dynamics with logistic damping

This paper investigates the following May--Nowak type model for viral infection in a bounded domain $Ω\subset \mathbb{R}^n$, $n \ge 2$: $\begin{cases} u_t = Δu - χ\nabla \cdot (u \nabla v) + κ- u - uw - μ\dfrac{u^{1+α}}{\ln^k(u+e)}, \\[1mm] γv_t = Δv - v + uw, \\[1mm] w_t = Δw - w + v, \end{cases}$ where $χ\in \mathbb{R}$, $μ> 0$, $k \in [0,1)$, $α> 0$, and $γ\in \{0,1\}$. We establish the global existence and uniform-in-time boundedness of classical solutions for suitably regular initial data under one of the following conditions: (A) $γ= 1$, $n = 2$, $k \in [0,1)$, $α= 1$, and $μ> 0$; (B) $γ= 1$, $3 \le n \le 5$, $k = 0$, $α= 1$, and $μ$ is sufficiently large; (C) $γ= 1$, $n \ge 6$, $k = 0$, $α> \frac{n-2}{4}$, and $μ> 0$; (D) $γ= 0$, $n \ge 2$, $k = 0$, $α> \frac{n+2}{2}$, and $μ>0$. In particular, in the physically relevant dimensions $n = 2,3$, both subquadratic and quadratic damping are sufficient to prevent blow-up. Moreover, in the fully parabolic case ($γ= 1$), our results improve upon recent findings by relaxing the condition $α> \frac{n}{2}$ to weaker assumptions within the above parameter regimes.

math.AP↗

Anomaly in baryon number

Chiral anomalies in baryon and lepton number currents in the GUT-inspired $SO(5) \times U(1) \times SU(3)$ gauge-Higgs unification model in the Randall-Sundrum (RS) warped space are examined. Total anomalies including contributions of all Kaluza-Klein (KK) excited modes of fermions running along internal triangular loops are expressed in terms of the values of wave functions of gauge-boson KK towers at the ultraviolet (UV) and infrared (IR) branes in the RS space. The covariant divergence of 5D baryon number current picks up an anomaly term proportional to ${\rm Tr} \,F_{μν} \tilde F^{μν} |_{SU(2)_L} - {\rm Tr} \, F_{μν} \tilde F^{μν} |_{SU(2)_R}$ at the UV and IR branes where $SU(2)_L \times SU(2)_R$ is a subgroup of $SO(5)$.

hep-ph↗

Stellar Collisions from Self-consistent Stellar Dynamics Around Growing Supermassive black Holes

The centers of galaxies harbor the densest stellar environments, where a massive black hole (MBH) accelerates stars to such high velocities that direct collisions can result in high-energetic phenomena, such as gravitational wave sources, kilonovae, and supernova-like transients. These collisions can reshape the cluster's density profile and release gas that can be subsequently accreted by the MBH. However, the evolving rates of such phenomena from self-consistent dynamics around mass-growing MBHs remain largely unexplored. In this work, we simulate nuclear star clusters (NSCs) across a range of masses and density profiles by employing the GNC Monte Carlo code, that self-consistently models stellar dynamics and the subsequent accretion of released gas. We find that stellar collisions flatten the density cusp in the innermost regions ($r \lesssim 10^{-3}-10^{-2}$ pc) within $\sim 0.1-1$ Gyr. While high initial collision rates in steep cusps quickly decline due to stellar depletion, the interplay between collisions and MBH growth is important only in massive NSCs ($M_\star \sim 10^9 M_\odot$). As MBH grows, increased stellar velocities shift the balance toward destructive collisions of which relative velocities can be $\gtrsim 2500 {\rm \,km\,s^{-1}}$. Consequently, present-day destructive collision rates in massive clusters remain high ($10^{-4}\sim 10^{-3}{\rm yr}^{-1}$), whereas they are smaller in Milky Way-like NSCs or negligible in smaller NSCs. Our results highlight a crucial synergy between stellar dynamics and MBH growth, identifying massive galaxies as prime targets for observing transients from destructive stellar collisions.

astro-ph.GA↗

Amplified Memory and Finite-Time Regularity in Driven Non-Hermitian Systems

We study the roles of broken spectra and exceptional points in finite-time driven non-Hermitian fermionic dynamics. We compute the Nambu biorthogonal correlation matrix for this purpose. The imbalanced-pairing Kitaev chain serves as our concrete realization. Transient passage through a broken-spectrum region amplifies preparation memory. The effect survives when the final Hamiltonian returns to a real-spectrum regime. Negative imbalance forces spectral nonpositivity in the static endpoint state. We isolate the genuine drive history by subtracting out this baseline, leaving an excess that is strictly controlled by the accumulated imaginary-energy action and persists over the entire post-ramp time window. A connected longitudinal correlation mirrors this physics. Its slow-ramp growth tracks the corresponding doubled action. Unstable sectors instead continue amplifying post-ramp if the drive halts inside the broken-spectrum region. Exceptional points yield distinct physics. The finite-time propagator and subsystem correlation matrix remain entirely regular near an exceptional endpoint, even as the quasiparticle gap exhibits its characteristic square-root closing. This finite-time regularity reflects the analyticity of the matrix evolution in the endpoint parameter; a diagonalizable endpoint is strictly not required. A diverging long-time crossover eventually reveals the exceptional scale. We halt the drive exactly at the exceptional point to find that the correlation projector and the connected longitudinal correlation share an identical ballisti front. The subsystem saturation length establishes a distinct but comparable spatial scale. Memory and exceptional-endpoint scalings show no divergence across the tested negative-imbalance range. Memory scaling is fixed by the drive and remains insensitive to the specific choice of real-spectrum final endpoint.

cond-mat.stat-mech↗

Is Broader Better? A Controlled Study of Multilingual Coverage and Pretraining Objective in Frozen SSL Encoders for Speech Deepfake Detection

Frozen self-supervised (SSL) speech encoders are strong, low-cost front ends for audio deepfake detection, and recent comparisons agree that large, multilingual, discriminative encoders generalize best out of domain. These comparisons fail to control for encoder capacity, pretraining objective, and multilingual coverage together, identifying which encoder wins without isolating why. We present a controlled decomposition with a fixed pipeline and trainable capacity. We vary multilingual coverage on four wav2vec2-family encoders, matched to ~315M parameters. We isolate the pretraining objective on two encoders matched on identical data. Coverage does not help monotonically, as out-of-domain error drops sharply at the ~100-language scale (XLS-R) but does not improve further at the 1406-language extreme (MMS). We find that a mid-coverage encoder is strongest on farther out-of-domain sets, matching or surpassing a 577M-parameter model at 315M. Its lead on these far sets, statistically significant under paired bootstrap, and on the official ASVspoof 5 cost metric holds under two backends. Separately, masked-prediction pretraining generalizes better than contrastive on identical data (In-the-Wild EER 26.5% vs. 46.8%). Within this fixed frozen-encoder recipe, we find that broader and larger models are not reliably better.

eess.AS↗

SN 2025aedz: A typical short-plateau type IIP supernova with rapid post-peak decline

Type IIP supernovae (SNe IIP) are the most common subclass of core-collapse SNe in observations. However, SNe IIP with short plateaus of the order of tens of days are rarely observed. The progenitors for this kind of SN can help to address the red supergiant issue in stellar evolution. In this article, we report optical photometry and spectroscopy of SN\,2025aedz, a rapidly post-peak declining SN IIP with a typical short plateau. It exhibits a peak absolute magnitude of $M_r=-17.16\pm0.03$\,mag. The $r$-band light curve shows a steep early post-peak decline of $\sim5\,\mathrm{mag}\,(100\,\mathrm{d})^{-1}$ followed by a relatively short plateau, with a plateau duration of $\sim50\pm3$\,d. The overall spectral evolution is consistent with that of normal SNe~IIP, showing a blue continuum with prominent Balmer P-Cygni profiles during the photospheric phase, followed by the gradual strengthening of hydrogen and metal lines as the ejecta cools down, although the metal lines remain weak and the expansion velocities decline rapidly. SN\,2025aedz is similar to the short-plateau SN\,2018gj in its overall evolution, whereas its pronounced early decline resembles that of SN\,2023ufx, which has the shortest plateau duration known so far. The radioactive tail of the bolometric light curve implies a synthesized $^{56}$Ni mass of $\sim0.03\pm0.01\,M_\odot$. Motivated by the steep early decline, we performed radiation hydrodynamic simulations by exploring different circumstellar material configurations to reproduce its early bolometric light curve. These simulations indicate that SN\,2025aedz originated from a progenitor with a relatively low-mass hydrogen envelope, possibly produced through enhanced mass loss or binary interaction, while the steep early decline is likely explained by additional luminosity from circumstellar interaction.

astro-ph.HE↗

Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes. For fixed $L \ge 3$ and $0 < α\le 1/12$, the optimal expected width on the worst pure cohort is $Θ_{α,L}([M(t+1)]^{-1/2})$ when every task is observed and $Θ_{α,L}([M(t+\sqrt{M})]^{-1/2})$ when omission is allowed. The lower bounds cover adaptive hard-budget policies, and fixed random-subset designs attain both rates through disagreement certificates. A joint mean/disagreement interval turns the task-covering law into practical finite-budget inference. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task-covering design reduces median point-estimation MSE by 87.0\% relative to pooled uniform sampling, while the Joint certificate produces narrower confidence intervals in 15/16 panels and reduces median interval width by 30.6\%. Finite-regime analyses identify task coverage as the effective choice at the evaluated scale and characterize how cohort size and within-task agreement determine the useful operating region. Together, the sharp laws and fixed-budget evidence make replication and task coverage explicit design variables for information-efficient repeated evaluation.

cs.AI↗

A Deterministic Polynomial Kernel for Odd Cycle Transversal

We give a deterministic polynomial kernel for Odd Cycle Transversal, derandomizing the randomized kernel of Kratsch and Wahlström (TALG 2014). Our algorithm uses a deterministic polynomial-time construction of almost multilinear representations of gammoids. Such a representation assigns a block of columns to each element so that, for every subset of elements, the normalized matrix rank approximates its matroid rank to within a prescribed additive error $δ$. The construction builds on recent breakthroughs in NC algorithms for matching. Our kernelization algorithm then computes the required representative families from these representations.

cs.DS↗