arXiv ScienceSearch

arXiv subjects

Song Li

Publications and source records attributed to Song Li.

At least 19 recordsLinked to original sources

A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors

Modern AI agent harnesses expose lifecycle hooks that bind shell commands to runtime events such as session start, tool calls, and file edits. These commands run with host privileges yet ship as lifecycle-hook configuration and may fire at times the LLM never observes. We identify the lifecycle-hook update path, which harnesses trust blindly, as a new attack surface. Under a supply-chain threat model in which an attacker controls only plugin metadata and lifecycle-hook configuration, a benign versioned plugin can be trojanized by an update that silently binds attacker-chosen commands to benign events, yielding malicious host-side behavior such as privilege escalation. We propose HookPry, an open-source and fully automated attack framework that systematically exploits this vulnerability across heterogeneous AI agent harnesses. HookPry realizes ten attack objectives; across 25 combinations of harnesses and backends in 1,000 end-to-end runs, it compromises all seven evaluated harnesses, with per-harness success rates reaching 92.5%. Representative defenses remain insufficient: Microsoft Defender has 0% recall, and the union of three static defenses misses 47.5% of malicious artifacts.

cs.CR

Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR

Large language model (LLM)-based automatic speech recognition (ASR) systems achieve strong performance on conventional speech data by leveraging powerful linguistic priors and multilingual capabilities. However, under challenging conditions, these priors can override acoustic evidence, resulting in unintended translation, instruction execution, repetition, or catastrophic deletion. We propose Likelihood-Constrained Acoustic Reranking (LCAR), a training-free decoding method that improves acoustic grounding while preserving support from the base model. At each decoding step, LCAR first retains tokens whose base-model likelihood falls within a margin of the greedy token, then reranks them using an acoustic compatibility score computed from attention-pooled audio embeddings and the existing LM head. By restricting acoustic intervention to plausible, model-supported alternatives, LCAR requires no additional training, external detector, reference transcript, or auxiliary model at inference. We evaluate LCAR on four LLM-based ASR systems using human-audited TTS and open-source speech challenge suites. At $\delta=0.60$, LCAR removes 38.8--57.1\% of detector-identified hallucination failures while largely maintaining WER/CER on standard open-source test sets.

eess.AS

PhaseLift for Coded Diffraction Patterns: Optimal Sampling Rate

Recovering a complex-valued signal from coded diffraction patterns, namely the Fourier intensities obtained after modulating the signal with a collection of masks, is a fundamental structured phase retrieval problem arising in diffraction imaging and related applications. Despite its practical importance, the theoretical analysis of this structured framework remains scarce. In the standard random mask model, the optimal sampling rate achievable by computationally tractable recovery methods has remained open. In this paper, we establish the optimal sampling rate for the PhaseLift feasibility program. More precisely, PhaseLift achieves exact recovery of an unknown signal $\pmb{x}_0\in\mathbb{C}^n$, up to a global phase, from $\mathcal{O}(\log n)$ random masks, with polynomially decaying failure probability. Since $\Omega(\log n)$ masks are necessary to identify certain signals under the erasure mask ensemble, our result thereby achieves the optimal mask complexity. Equivalently, PhaseLift attains the optimal total sampling rate of $m=\mathcal{O}( n\log n)$ scalar intensity measurements. The proof is based on an approximate dual certificate construction via a refined golfing scheme that combines adaptive mask allocation with a dimension-independent truncation threshold.

cs.IT

Low-Rank Matrix Recovery via Heavy-Tailed Quadratic Sampling

The problem of recovering an (approximately) low-rank Hermitian matrix $\pmb{M}_0 \in \mathbb{C}^{n \times n}$ of rank $r$ from quadratic sampling matrices of the form $\{\pmb{a}_k \pmb{a}_k^*\}_{k=1}^m$ arises in a variety of applications, including phase retrieval. To obtain rigorous recovery guarantees, the sampling vectors $\{\pmb{a}_k\}_{k=1}^m$ are typically modeled probabilistically. However, most existing theoretical results rely on Gaussian or sub-Gaussian assumptions, which may not accurately capture practical data models. In many applications, sampling vectors exhibit heavier tails, while theoretical understanding in such regimes remains scarce. In this paper, we bridge this gap. We show that two widely used convex approaches, nuclear norm minimization and semidefinite-constrained empirical risk minimization, achieve uniform, stable, and robust recovery under the mild assumption that the entries of the sampling vectors have only finite $4+\delta$ moments, with the optimal sample complexity $m = \mathcal{O}(rn)$ up to moment-dependent constants. The two main ingredients of our analysis are moment estimates for quadratic forms established via decoupling, together with recent advances in covariance estimation in heavy-tailed settings. As byproducts, we also establish the optimal sample complexity for low-rank matrix recovery under complex projective $4$-design sampling, thereby improving upon previous results, and obtain stability guarantees for phase retrieval under similarly weak moment assumptions.

math.ST

Measurability of Quadrupole Deviations from Kerr in Binary black hole Mergers

We investigate the measurability of black hole quadrupole deviations from Kerr using five binary black hole mergers observed by the LIGO-Virgo-KAGRA Collaboration with the beyond-general-relativity full-waveform model $\Psi_{\mathrm{FD}}$. While earlier lower-SNR events mildly favored nonzero quadrupole deviations, the newly included high-SNR GWTC-4 events GW231226, GW230814, and GW250114 yield results increasingly consistent with the Kerr prediction. In particular, GW230814 and GW250114, the two highest-SNR events in our sample, yield deviations consistent with zero. We further perform separate inspiral and post-inspiral analyses and find both the posterior distributions centered close to $\Delta Q/Q=0$ for GW230814 and GW250114. Overall, the full-waveform, inspiral, and post-inspiral results for GW230814 and GW250114 reveal no observable departure from the no-hair theorem within the sensitivity of the current data and the $\Psi_{\mathrm{FD}}$ framework. Although the limited number of events prevents a definitive conclusion, future detections of additional high-SNR binary black hole mergers will enable increasingly stringent and robust tests of the Kerr nature of black holes.

gr-qc

Dimensionality-Driven Charge Stabilization of Group-IV Color Centers in Diamond Ultrathin Films

Neutral group-IV vacancy (XV, X = Si, Ge, Sn, and Pb) centers in diamond are emerging solid-state spin-photon interfaces because of their favorable spin coherence and inversion symmetry-protected optical transitions. However, stabilizing their neutral charge state typically requires stringent Fermi-level engineering in high purity boron-doped diamond, which poses significant materials-growth challenges. Here, we demonstrate that dimensional confinement in diamane provides an alternative route to charge-state stabilization without intentional doping. Using first-principles calculations, we show that quantum confinement and surface termination cooperatively tune the host band gap and shift the occupied defect states upward from the valence-band edge, thereby enlarging the thermodynamic stability window of the neutral charge state and suppressing valence-band assisted excitation pathways. We further reveal that the thickness and surface termination of diamane enable systematic tuning of the electronic structure, zero-field splitting, and spin-orbit coupling of XV centers while largely preserving their optical transition energies. Among the structures considered, hydrogenated diamane offers the most favorable balance between charge-state stability and magneto-optical performance. More broadly, our findings establish dimensional confinement as a general strategy for engineering the charge, optical, and spin properties of solid-state quantum defects.

cond-mat.mtrl-sci

Convergence of Spectral Descent for Non-smooth Optimization

The Muon optimizer has recently demonstrated remarkable empirical success in training large language models. However, the theoretical understanding of its mechanisms remains limited. Current convergence guarantees for Muon rely heavily on smoothness assumptions, leaving its non-smooth convergence behavior largely unexplored. In this work, we take a step toward bridging this gap by investigating Spectral Descent (SD), a simplified variant of Muon, together with its truncated counterpart, Truncated Spectral Descent (TSD). Under convexity, Lipschitz continuity, and sharpness conditions, we establish global linear convergence for both SD and TSD in non-smooth convex formulations. We also study regularized variants equipped with decoupled weight decay and derive sublinear convergence guarantees through their connection with Frank-Wolfe methods. Finally, we apply our theoretical framework to robust low-rank matrix recovery under mixed sparse and dense noise regimes and provide rigorous recovery guarantees. Numerical experiments support the theoretical findings and demonstrate the effectiveness of Muon-type methods for non-smooth optimization.

cs.LG

Interacting donor-acceptor pairs as the origin of coupled spin-optical signals in hexagonal boron nitride

Optically addressable spin defects in hexagonal boron nitride hold promise for room-temperature quantum technologies, but their microscopic identities remain largely unknown. Using first principles calculations, we show that coupled spin optical signals arise from interacting donor acceptor pairs, not the commonly believed isolated defects. Intra and inter pair separations control charge transfer, electronic structure, and spin coupling, thereby greatly modulating zero phonon lines, phonon sidebands, lifetimes, and the sign of optically detected magnetic resonance contrast. Importantly, we identify two distinct charge-state-dependent coupling regimes and extend this picture to correlated defect ensembles, explaining the wide diversity of experimental observations. Our results establish a microscopic framework for coupled defect behavior and provide design principles for spin-active quantum emitters in wide bandgap semiconductors.

cond-mat.mtrl-sci

VIPER-MCP: Detecting and Exploiting Taint-Style Vulnerabilities in Model Context Protocol Servers

Model Context Protocol (MCP) has emerged as a standard interface for connecting LLM agents to external tools. Because MCP servers expose privileged operations such as shell execution, network access, and file-system manipulation to agent-driven invocation, implementation flaws in tool handlers can create a direct path from natural-language input to security-sensitive sinks, potentially granting attackers remote code execution or full system compromise. Existing approaches either produce unconfirmed static alerts without dynamic validation, or rely on fixed template libraries that lack code-level guidance and fail to trigger vulnerabilities requiring specific parameter shapes or multi-step taint paths. In this paper, we present VIPER-MCP, the first end-to-end automated vulnerability auditing framework for MCP servers that not only detects taint-style vulnerabilities but also dynamically confirms their exploitability by producing concrete proof-of-concept prompts. VIPER-MCP introduces two novel techniques: (1) an anchor-query pass in a two-pass static analysis strategy that augments standard taint alerts with function-level structural context, resolving file-level static artifacts to specific MCP tool handlers and producing vulnerability-anchored call chains; and (2) a feedback-driven prompt evolution mechanism that employs dual-mutator scheduling that independently corrects tool-selection drift and deepens parameter penetration, together with fitness-scored seed selection to iteratively refine natural-language prompts toward vulnerable sinks. In a large-scale scan of 39,884 real-world open-source MCP server repositories, VIPER-MCP discovered 106 0-day vulnerabilities, all of which were confirmed through end-to-end exploit traces, with 67 CVE IDs assigned to date.

cs.CR

PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning

Cell-type-specific marker genes are fundamental to plant biology, yet existing resources primarily rely on curated databases or high-throughput studies without explicitly modeling the supporting evidence found in scientific literature. We introduce PlantMarkerBench, a multi-species benchmark for evaluating literature-grounded plant marker evidence interpretation from full-text biological papers. PlantMarkerBench is constructed using a modular curation pipeline integrating large-scale literature retrieval, hybrid search, species-aware biological grounding, structured evidence extraction, and targeted human review. The benchmark spans four plant species -- Arabidopsis, maize, rice, and tomato -- and contains 5,550 sentence-level evidence instances annotated for marker-evidence validity, evidence type, and support strength. We define two benchmark tasks: determining whether a candidate sentence provides valid marker evidence for a gene-cell-type pair, and classifying the evidence into expression, localization, function, indirect, or negative categories. We benchmark diverse open-weight and closed-source language models across species and prompting strategies. Although frontier models achieve relatively strong performance on direct expression evidence, performance drops substantially on functional, indirect, and weak-support evidence, with evidence-type confusion emerging as a dominant failure mode. Open-weight models additionally exhibit elevated false-positive rates under ambiguous biological contexts. PlantMarkerBench provides a challenging and reproducible evaluation framework for literature-grounded biological evidence attribution and supports future research on trustworthy scientific information extraction and AI-assisted plant biology.

cs.CL

How Label Imbalance Shapes Geometry: A General Spectral Analysis of Multi-Label Neural Collapse

This work investigates the phenomenon of Neural Collapse (NC) in multi-label classification, extending its conceptual framework from multi-class learning to general correlated and imbalanced multi-label settings. Although recent studies have identified a ''tag-wise averaging'' structure for multi-label features, this view relies on implicit assumptions of label balance and combinatorial symmetry. Consequently, it fails to account for the geometrical distortions caused by intrinsic label correlations and data imbalance, which are common in practice. We resolve the multiplicity-one imbalance conjecture raised by Li et al. (2024), showing that higher-multiplicity prototypes obey a class-frequency-weighted synthesis rule rather than uniform averaging. To address this, we propose a rigorous spectral-control framework to analyze the terminal phase of multi-label learning under general imbalanced conditions. We introduce the label covariance spectrum $\kappa_m$, a scalar controlling the distribution-dependent lower-bound geometry, derived from the second-order moment matrix of the label distribution. Contrary to the averaging perspective, our analysis reveals that the centered label covariance spectrum controls the stability of terminal geometry by quantifying the weakest centered inter-class contrast directions. We prove that the classical Tag-wise Averaging emerges only as a special case under perfect orthogonality. Numerical experiments on synthetic distributions validate our theoretical bounds. This work resolves the scaled-average aspect of the imbalance conjecture and establishes a unifying theoretical framework that extends Neural Collapse to complex, imbalanced multi-label settings.

cs.LG

Gravitational Wave Signature of Aspherical Bubbles Driven by Thermal Fluctuation

Cosmological first-order phase transitions are a well-motivated source of stochastic gravitational waves (GWs), but most predictions are made based on the highly idealized model of perfectly spherical vacuum bubbles, neglecting thermal fluctuations. In this work we use $(3+1)$-dimensional lattice simulations of a scalar model with thermal initial conditions to quantify how thermal fluctuations distort bubble profiles and modify the resulting GW spectrum. We find that thermal fluctuations can strongly break spherical symmetry at early times, allowing even an isolated bubble to emit GWs. In multi-bubble simulations, thermal fluctuations systematically reshape the spectrum, suppressing the infrared part while enhancing and broadening the high-$k$ tail. We further provide an analytical estimate for the ultraviolet regime of the GW spectrum, which is in good agreement with our lattice results and suggests that this regime is dominated by thermal fluctuations. These effects could leave observable imprints in future GW searches.

hep-ph

Achieving High Efficiency And Enhanced Beam Quality In Laser Wakefield Acceleration

Laser wakefield acceleration, characterized by the extremely high electric field gradient exceeding 100GV/m, is regarded as a compact and cost affordable technology for the next generation of particle colliders and light sources. However, it has always been a major challenge to effectively increase the energy transfer efficiency from the laser to the accelerated beam, while ensuring the beam quality remains suitable for practical applications. This study demonstrates that the laser with shorter pulse duration allows for a two-step dechirping process of the accelerated electron beam with charge of nanocoulomb level. The electron beams with an energy spread of 1% can be generated with the energy transfer efficiency of 10% to 30% in a large parameter space. For example, one electron beam with the energy of 420MeV, the charge of 5.5nC and the RMS energy spread of 2% can be produced using an 8.3J laser pulse with 7.2fs duration.

physics.acc-ph

Large Language Models Cannot Reliably Detect Vulnerabilities in JavaScript: The First Systematic Benchmark and Evaluation

Researchers have proposed numerous methods to detect vulnerabilities in JavaScript, especially those assisted by Large Language Models (LLMs). However, the actual capability of LLMs in JavaScript vulnerability detection remains questionable, necessitating systematic evaluation and comprehensive benchmarks. Unfortunately, existing benchmarks suffer from three critical limitations: (1) incomplete coverage, such as covering a limited subset of CWE types; (2) underestimation of LLM capabilities caused by unreasonable ground truth labeling; and (3) overestimation due to unrealistic cases such as using isolated vulnerable files rather than complete projects. In this paper, we introduce, for the first time, three principles for constructing a benchmark for JavaScript vulnerability detection that directly address these limitations: (1) comprehensiveness, (2) no underestimation, and (3) no overestimation. Guided by these principles, we propose FORGEJS, the first automatic benchmark generation framework for evaluating LLMs' capability in JavaScript vulnerability detection. Then, we use FORGEJS to construct ARENAJS-the first systematic benchmark for LLM-based JavaScript vulnerability detection-and further propose JUDGEJS, an automatic evaluation framework. We conduct the first systematic evaluation of LLMs for JavaScript vulnerability detection, leveraging JUDGEJS to assess seven popular commercial LLMs on ARENAJS. The results show that LLMs not only exhibit limited reasoning capabilities, but also suffer from severe robustness defects, indicating that reliable JavaScript vulnerability detection with LLMs remains an open challenge.

cs.CR

LongCat-Flash-Omni Technical Report

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strategy that transitions from simpler to increasingly complex modality sequence modeling tasks, LongCat-Flash-Omni attains comprehensive multimodal capabilities while maintaining strong unimodal capability. Building upon LongCat-Flash, which adopts a high-performance Shortcut-connected Mixture-of-Experts (MoE) architecture with zero-computation experts, LongCat-Flash-Omni integrates efficient multimodal perception and speech reconstruction modules. Despite its immense size of 560B parameters (with 27B activated), LongCat-Flash-Omni achieves low-latency real-time audio-visual interaction. For training infrastructure, we developed a modality-decoupled parallelism scheme specifically designed to manage the data and model heterogeneity inherent in large-scale multimodal training. This innovative approach demonstrates exceptional efficiency by sustaining over 90% of the throughput achieved by text-only training. Extensive evaluations show that LongCat-Flash-Omni achieves state-of-the-art performance on omni-modal benchmarks among open-source models. Furthermore, it delivers highly competitive results across a wide range of modality-specific tasks, including text, image, and video understanding, as well as audio understanding and generation. We provide a comprehensive overview of the model architecture design, training procedures, and data strategies, and open-source the model to foster future research and development in the community.

cs.MM

LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models

This paper presents LongCat-Audio-Codec, an audio tokenizer and detokenizer solution designed for industrial grade end-to-end speech large language models. By leveraging a decoupled model architecture and a multistage training strategy, LongCat-Audio-Codec exhibits robust semantic modeling capabilities, flexible acoustic feature extraction capabilities, and low-latency streaming synthesis capabilities. It encodes speech at an ultra-low frame rate of 16.67 Hz, with a minimum bitrate of 0.43 kbps and a maximum bitrate of 0.87 kbps. Evaluation results demonstrate that LongCat-Audio-Codec achieves strong speech intelligibility and is capable of synthesizing highquality speech at low bitrate, thus effectively balancing coding efficiency and decoding quality. The inference code and model checkpoints of LongCat-Audio-Codec are available at: https://github.com/meituan-longcat/LongCat-Audio-Codec.

eess.AS

On a matrix constrained CKP hierarchy

The algebraic structures of integrable hierarchies play an important role in the study of soliton equations. In this paper, we use splitting theory to give a matrix representation of a constrained CKP hierarchy, which can be considered as a generalization of the $\hat{A}_{2n}^{(2)}$-KdV hierarchy and the constrained KP hierarchy. An equivalent construction in terms of the pseudo-differential operator is discussed. Darboux transformations, scaling transformation and tau functions $\ln \tau_f$ for this constrained hierarchy are studied. Moreover, we present formulas for the Virasoro vector fields on $\ln \tau_f$ for the $\hat{A}_{2 n}^{(2)}$-KdV hierarchy.

nlin.SI

Stable Phase Retrieval: Optimal Rates in Poisson and Heavy-tailed Models

We investigate stable recovery guarantees for phase retrieval under two realistic and challenging noise models: the Poisson model and the heavy-tailed model. Our analysis covers both nonconvex least squares (NCVX-LS) and convex least squares (CVX-LS) estimators. For the Poisson model, we demonstrate that in the high-energy regime where the true signal $pmb{x}$ exceeds a certain energy threshold, both estimators achieve a signal-independent, minimax optimal error rate $\mathcal{O}(\sqrt{\frac{n}{m}})$, with $n$ denoting the signal dimension and $m$ the number of sampling vectors. In contrast, in the low-energy regime, the NCVX-LS estimator attains an error rate of $\mathcal{O}(\|\pmb{x}\|^{1/4}_2\cdot(\frac{n}{m})^{1/4})$, which decreases as the energy of signal $\pmb{x}$ diminishes and remains nearly optimal with respect to the oversampling ratio. This demonstrates a signal-energy-adaptive behavior in the Poisson setting. For the heavy-tailed model with noise having a finite $q$-th moment ($q>2$), both estimators attain the minimax optimal error rate $\mathcal{O}( \frac{\| \xi \|_{L_q}}{\| \pmb{x} \|_2} \cdot \sqrt{\frac{n}{m}} )$ in the high-energy regime, while the NCVX-LS estimator further achieves the minimax optimal rate $\mathcal{O}( \sqrt{\|\xi \|_{L_q}}\cdot (\frac{n}{m})^{1/4} )$ in the low-energy regime. Our analysis builds on two key ideas: the use of multiplier inequalities to handle noise that may exhibit dependence on the sampling vectors, and a novel interpretation of Poisson noise as sub-exponential in the high-energy regime yet heavy-tailed in the low-energy regime. These insights form the foundation of a unified analytical framework, which we further apply to a range of related problems, including sparse phase retrieval, low-rank PSD matrix recovery, and random blind deconvolution.

math.ST