arXiv ScienceSearch

arXiv subjects

Jie Shi

Publications and source records attributed to Jie Shi.

At least 19 recordsLinked to original sources

Human-guided physics-constrained AI agents construct an auditable model of soil-plug evolution

Engineering predictions require physical mechanisms to be translated consistently into equations, discretization, code, and validation, yet errors can propagate despite local checks. Artificial-intelligence (AI) agents automate scientific tasks, but coordinating and independently auditing the theory-to-solver process under physical constraints and human oversight remains unresolved. We introduce a human-in-the-loop, physics-constrained multi-agent workflow where human experts define admissible physics and modeling boundaries, while agents retrieve evidence, derive equations, implement solvers, and audit the theory-to-code chain. Applied to soil-plug evolution during suction-caisson installation, the workflow generated and audited 6 formulations in 2.9 h of agent execution once physical knowledge and inputs were prepared. Among these formulations, adding seepage-driven soil void-ratio evolution to the geometric baseline reduced mean absolute final-heave error from 58.4% to 9.0% across 14 profiles; the selected model further incorporated near-wall dilation and achieved mean absolute percentage errors of 12.4% across 9 final-state cases and 4.2% at the endpoints of 5 process histories. Beyond predictive performance, blinded replay recovered all 9 target problems, while an independent audit uncovered 5 implementation problems after 36 predefined checks had passed. Overall, this work extends multi-agent AI beyond task automation toward human-governed engineering solvers.

cs.CE

When the Agent Becomes the Kernel: A Systematization of Security on the Path to AI-Native Operating Systems

Large language model agents are now privileged principals that take consequential actions: editing code repositories, operating inboxes, completing purchases. Their authority is kernel-grade, but it comes without what classical systems security requires: a trusted mediator interposed on every access. Operating-system vendors are now rebuilding the platform around this de-facto agent kernel, inheriting complete mediation as a design problem. We systematize the security of such systems around a single distinction: a crossing mediated over provenance admits a deterministic check, while one over content semantics does not. A trust-boundary taxonomy locates where mediation must occur and isolates the central mediation gap at two kinds of semantic judgment: distinguishing data from instruction in untrusted input, and an authorized action from an unauthorized one. We argue that this gap leaves an irreducible residual of undetected attacks wherever inputs and actions are not restricted in advance to an enumerated set. The same distinction makes attack-success statistics actionable, placing each number on a spectrum from deployment debt (a sound deterministic mediator left unused) to a structural gap (no such mediator known). We systematize defenses across runtime monitoring, architectural separation, and authorization, and show that current evaluations tend to overstate deployed security through evaluation-validity failures. Finally, we carry that analysis forward beyond the de-facto kernel, to an architecture in which the model itself becomes the arbitration core, and derive the design constraints, open challenges, and research agenda for a security-first AI-native OS.

cs.CR

GenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages

Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain valid after adaptation? We present GenTraceBench, a controlled benchmark spanning five TTS architectures, 16 pre-/post-training variants, and 49,728 utterances generated with fixed texts and speaker prompts. Under a train-on-foundation, test-on-adapted protocol, we evaluate binary detection, closed-set attribution, and open-set verification. DPO and GRPO generally preserve fingerprints, whereas some SFT and pre-training-data changes cause substantial drift; effect sizes vary across three forensic backbones. Repeated training runs confirm the largest W2V-BERT attribution drop, while a data-mixture control with comparable speech quality shows that composition change need not cause drift. In W2V-BERT verification, multi-shot enrollment reduces EER for the SFT condition from 44.4% to 11.0%, whereas the SingNet-only condition remains at or above 45% EER.

cs.SD

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions. Yet aligned models can fail when harmful intent is concealed within seemingly benign contexts. Robust safety therefore requires both knowledge of safety boundaries and \textbf{vigilance}: the ability to detect unusual premises, misleading reasoning, and latent risks beneath surface-level semantics. Vigilance requires models to scrutinize a request's underlying intent and assumptions before acting. To cultivate this capability, we introduce \textbf{cunning questions}, which are not necessarily safety-related but contain misleading premises, atypical reasoning, or subtle inconsistencies. We hypothesize that learning to look beyond such reasoning traps can transfer to safety-critical scenarios. Experiments show that Cunning training improves robustness to out-of-distribution jailbreak attacks and strengthens subsequent safety fine-tuning. Furthermore, augmenting an existing state-of-the-art safety alignment pipeline with Cunning establishes a new state of the art across our evaluated settings, reducing mean ASR across nine backbone--benchmark combinations from 17.40\% to 15.05\%. Trace analysis after matched safety fine-tuning suggests that safety judgments are more likely to govern responses before harmful planning begins. A conditional theoretical analysis further characterizes when invariance learned from cunning data can transfer to safety-related inputs. These findings suggest that cunning data can strengthen model vigilance and complement conventional safety alignment.

cs.AI

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to building effective and trustworthy human-AI interaction. Existing reinforcement learning methods optimize single-turn utterances and sparse outcome rewards, producing short-sighted policies that struggle to manage goal-relationship tensions across multi-turn interactions. We propose SocialRL, a multi-turn reinforcement learning framework addressing both challenges. First, we apply multi-turn reinforcement learning using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning. Second, we design six process reward dimensions capturing the goal-relationship trade-off, including goal advancement, relational attunement, contextual coherence, etc. A reward model dynamically generates fine-grained scoring criteria for each dimension, while a stage-aware weight schedule prioritizes relationship-building in early turns, goal advancement mid-way, and balanced closure late. Across multiple social-dialogue benchmarks, SocialRL improves Goal Achievement by an average of 9.2 percentage points over the corresponding Base models. These results demonstrate the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.

cs.CL

Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchmark organizes heterogeneous trajectories under a unified component schema and provides annotations of the primary attribution component, together with attack and execution chains where applicable. Instantiating the benchmark with trajectories from AgentDojo and the Stage and Canary settings of Agent3Sigma yields more than 1,300 annotated trajectories covering task-aligned actions, unsafe actions, and safety refusals. The benchmark defines two evaluation tasks, primary attribution localization and attribution-chain recovery, and provides reference baselines based on incremental trajectory contribution and component-level leave-one-out perturbation. It captures diverse attribution settings, including local and long-range attribution as well as structured attribution chains. Reference baseline results exhibit substantial performance differences across these settings, providing an initial characterization of the benchmark's attribution challenges. Beyond this initial instantiation, we release a reusable annotation skill that enables trajectories generated by new agent models to be standardized, annotated, and evaluated under the same framework. Project resources and future releases are available at https://github.com/chenjing-2024/agent-trajectory-attribution.

cs.AI

Convexity criterion and radial-profile response for off-shell Kerr geometries: a fuzzy-dark-matter profile as an analytic benchmark

We establish a sufficient one-minimum criterion for the off-shell Kerr family $Δ(r) = r^2 - 2rm(r) + a^2$ with a positive, nondecreasing mass profile $m(r)$, showing that $1 - 2m'(r) - rm''(r) > 0$ ensures strict convexity and determines root counts for $Δ$. Using a fuzzy-dark-matter-inspired benchmark satisfying this bound, we derive first-order responses for the outer horizon, extremal branch, photon sphere, and shadow functional under general deformations $m/M_{\text{ADM}} = 1 + \varepsilon h$. We demonstrate that static horizon and photon responses are profile-controlled, spin-odd shadow displacements are completion-dependent, and scale-consistent weak-field limits render local profile-gradient effects negligible ($\ll 10^{-20}$), confirming the strong-field box as a formal radial-profile benchmark rather than a self-consistent rotating scalar-field solution.

gr-qc

EMRI Dephasing from a Torsion-Inspired Near-Zone Kerr Deformation: Motivated by Spin-Polarized Dark Matter

Extreme-mass-ratio inspirals (EMRIs) are sensitive probes of weak conservative perturbations in the strong-field region of massive black holes. We study a phenomenological EMRI model motivated by Einstein--Cartan gravity in which a spin-polarized dark-matter spike is described by a Weyssenhoff fluid. After torsion is eliminated algebraically, the local spin contribution contains a repulsive exterior source $U_{tt}^{\rm spin}\propto-σ_0^2/r^3$. Solving the corresponding static linearized field equation, however, does not produce a global $1/r^3$ metric perturbation; the response contains a mass renormalization, a logarithmic $r^{-1}$ tail, and an $M/r^2$ term. We therefore introduce $g_{μν}^{\rm eff}=g_{μν}^{\rm Kerr}+αh_{μν}^{\rm eff}$ only as a local near-zone matching ansatz, not as a complete rotating Einstein--Cartan black-hole solution. Within this torsion-inspired deformation we compute circular equatorial inspirals and analytic-kludge waveforms. The fiducial model can produce large phase shifts in an idealized adiabatic calculation, but the forecast is optimistic and does not include a full LISA/Taiji response, Teukolsky/self-force fluxes, eccentricity, inclination, or high-dimensional parameter degeneracies. The results should be read as constraints on an effective near-zone operator rather than as a prediction of minimally coupled Einstein--Cartan dark matter.

gr-qc

Random Fixed Point Theorems for Relaxed Asymptotic Contractions in Random Normed Modules

We introduce the notion of a random relaxed asymptotic contraction in the setting of random normed modules. The contraction condition employs two quasi-metrics that are built directly from the random operator: a lower quasi-metric which adaptively switches between a four-point minimum and the ordinary one-step distance, and an upper quasi-metric which takes the maximum of four fundamental distances. The bounds are allowed to depend on the iteration index and are required to converge locally uniformly almost surely to a Boyd--Wong function. Using the fibre decomposition method based on \(σ\)-stability and the local property, we show that any such mapping defined on an essentially bounded, \(σ\)-stable and \(L^0\)-closed set admits a unique random fixed point, and all iterates converge in the \((ε,λ)\)-topology. Our result strictly generalizes the random analogue of Kirk's asymptotic contraction theorem and unifies several deterministic and random fixed point theorems under a single flexible framework.

math.FA

A Fixed Point Theorem for Random Asymptotically Pointwise Contractions

This paper combines the decomposition technique ($σ$-stability) in random functional analysis with the deterministic theory of asymptotically pointwise contractions to provide a complete self-contained derivation of a fixed point theorem for random asymptotically pointwise contractions. We assume the contraction function is linear $ψ(t)=λt$ ($λ<1$) and focus on the linear case under the assumption that $G$ is bounded. By choosing $p$ sufficiently large so that $5^{1/p}λ<1$, we apply the deterministic theorem in $L^p(E)$. The paper gives detailed explanations of concepts such as random normed modules, the $(ε,λ)$-topology, and $σ$-stability, and reviews the historical development of fixed point theory in the introduction.

math.FA

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than the words alone. Who is speaking, how they sound, and where the conversation takes place can each turn an otherwise benign request into one that is unsafe, unfair, or privacy-violating. Existing benchmarks, however, largely focus on basic audio comprehension, study individual risks in isolation, or conflate content that is inherently harmful with content that only becomes problematic due to its acoustic context. We introduce VoxSafeBench, among the first benchmarks to jointly evaluate social alignment in SLMs across three dimensions: safety, fairness, and privacy. VoxSafeBench adopts a Two-Tier design: Tier1 evaluates content-centric risks using matched text and audio inputs, while Tier2 targets audio-conditioned risks in which the transcript is benign but the appropriate response hinges on the speaker, paralinguistic cues, or the surrounding environment. To validate Tier2, we include intermediate perception probes and confirm that frontier SLMs can successfully detect these acoustic cues yet still fail to act on them appropriately. Across 22 tasks with bilingual coverage, we find that safeguards appearing robust on text often degrade in speech: safety awareness drops for speaker- and scene-conditioned risks, fairness erodes when demographic differences are conveyed vocally, and privacy protections falter when contextual cues arrive acoustically. Together, these results expose a pervasive speech grounding gap: current SLMs frequently recognize the relevant social norm in text but fail to apply it when the decisive cue must be grounded in speech. Code and data are publicly available at: https://amphionteam.github.io/VoxSafeBench_demopage/

cs.SD

Fixed Point Theorems for Relaxed Asymptotic Contractions via Two Quasi-Metrics

We introduce a new class of asymptotic contractions that employs two quasi-metrics defined directly in terms of the underlying mapping. The contraction condition compares these two quantities via a sequence of bounding functions that converge locally uniformly to a Boyd-Wong function. This framework relaxes the hypotheses of Kirk's asymptotic fixed point theorem and strictly contains it as a special case. Assuming only the continuity of the map and the boundedness of some orbit in a complete metric space, we prove both the existence and uniqueness of a fixed point, along with the convergence of all iterates to that point.

math.FA

Fixed Points of Asymptotic Pointwise Contractions under Local Uniform Convergence

We introduce a weak asymptotic version of nonlinear contraction, termed \emph{asymptotic pointwise contraction}. For a mapping on a metric space, this notion requires the existence of a sequence of functions that dominate the distances between the $n$-th iterates of any two points. The sequence is assumed to converge pointwise to a limit function, and the convergence is required to be uniform on every bounded set (i.e., locally uniform). The limit function is then controlled by a Boyd--Wong type condition: there exists a nondecreasing, right upper semicontinuous function strictly below the identity on positive numbers, and the limit function is bounded above by this function evaluated at a maximum term that involves not only the distance between the two points but also distances from each point to its image and mutual distances between each point and the image of the other. By standard analytic arguments we prove that if the mapping is continuous on a complete metric space and possesses a bounded orbit, then its iterates converge to a unique fixed point. This result extends Kirk's asymptotic contraction theorem by replacing global uniform convergence on $[0,\infty)$ with the weaker condition of local uniform convergence.

math.FA

A Diffusion-Contrastive Graph Neural Network with Virtual Nodes for Wind Nowcasting in Unobserved Regions

Accurate weather nowcasting remains one of the central challenges in atmospheric science, with critical implications for climate resilience, energy security, and disaster preparedness. Since it is not feasible to deploy observation stations everywhere, some regions lack dense observational networks, resulting in unreliable short-term wind predictions across those unobserved areas. Here we present a deep graph self-supervised framework that extends nowcasting capability into such unobserved regions without requiring new sensors. Our approach introduces "virtual nodes" into a diffusion and contrastive-based graph neural network, enabling the model to learn wind condition (i.e., speed, direction and gusts) in places with no direct measurements. Using high-temporal resolution weather station data across the Netherlands, we demonstrate that this approach reduces nowcast mean absolute error (MAE) of wind speed, gusts, and direction in unobserved regions by more than 30% - 46% compared with interpolation and regression methods. By enabling localized nowcasts where no measurements exist, this method opens new pathways for renewable energy integration, agricultural planning, and early-warning systems in data-sparse regions.

cs.LG

Detectability and Systematic Bias from First-Order Phase-Transition Dephasing in Kerr EMRIs

We study gravitational-wave dephasing induced by an effective first-order phase transition in a Kerr extreme mass-ratio inspiral (EMRI). The transition is modeled phenomenologically as a finite-width restructuring of the dissipative flux sector, and its observational consequences are quantified with standard LISA matched-filter diagnostics. For a representative system with $M=2\times10^{5}M_\odot$, $μ=1.4M_\odot$, and $\hat a=0.90$, we obtain $ρ_{\rm B}=5.064$, $ρ_{\rm T}=4.073$, $ρ_{\rm R}=1.051$, and a mismatch $\mathcal M=2.986\times10^{-3}$ after maximization over extrinsic time and phase shifts. Although the normalized mismatch remains small, the accumulated phase difference grows to $ΔΦ_{22}^{\rm SF}\sim 5\times10^{3}\,\mathrm{rad}$, indicating that a narrow transition window can generate a large coherent deformation of the inspiral clock while leaving the waveform globally close to the baseline branch in detector-weighted norm. The resulting signal therefore lies in a bias-sensitive regime, characterized by small mismatch, order-unity residual norm, and large cumulative dephasing. Our results suggest that the dominant consequence of the transition sector is not loss of detectability, but loss of faithfulness for precision inference. This motivates future LISA EMRI waveform models that incorporate parameterized transition sectors directly into the waveform manifold.

gr-qc

Integrating Weather Station Data and Radar for Precipitation Nowcasting: SmaAt-fUsion and SmaAt-Krige-GNet

Short-term precipitation nowcasting is essential for flood management, transportation, energy system operations, and emergency response. However, many existing models fail to fully exploit the extensive atmospheric information available, relying primarily on precipitation data alone. This study examines whether integrating multi variable weather-station measurements with radar can enhance nowcasting skill and introduces two complementary architectures that integrate multi variable station data with radar images. The SmaAt-fUsion model extends the SmaAt-UNet framework by incorporating weather station data through a convolutional layer, integrating it into the bottleneck of the network; The SmaAt-Krige-GNet model combines precipitation maps with weather station data processed using Kriging, a geo-statistical interpolation method, to generate variable-specific maps. These maps are then utilized in a dual-encoder architecture based on SmaAt-GNet, allowing multi-level data integration. Experimental evaluations were conducted using four years (2016--2019) of weather station and precipitation radar data from the Netherlands. Results demonstrate that SmaAt-Krige-GNet outperforms the standard SmaAt-UNet, which relies solely on precipitation radar data, in low precipitation scenarios, while SmaAt-fUsion surpasses SmaAt-UNet in both low and high precipitation scenarios. This highlights the potential of incorporating discrete weather station data to enhance the performance of deep learning-based weather nowcasting models.

cs.LG

Analytical derivation of long-term dephasing caused by phase transitions in the context of Kerr black holes

Extreme Mass Ratio Inspirals (EMRIs) constitute a prime target for future space-based gravitational-wave observatories such as LISA. In this paper, we analytically investigate the long-term phase shift (dephasing) in the gravitational wave signal induced by a first-order quantum chromodynamics (QCD) phase transition within a neutron star orbiting a supermassive Kerr black hole. By modeling the transition from a hadronic phase to a quark core phase, we quantify the sudden change in the tidal deformability ($Λ$) of the secondary object. Utilizing the Teukolsky formalism and Post-Newtonian expansions, we derive a strict analytical scaling law for the accumulated dephasing. We demonstrate that the Kerr spin parameter $a$ and the critical phase transition orbital velocity $v_c$ significantly amplify the dephasing effect. Our analytical framework provides a robust tool for probing the non-perturbative QCD equation of state at high baryon densities using gravitational wave astronomy.

gr-qc

Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive Refinement

Large language models (LLMs) aligned for safety often suffer from over-refusal, the tendency to reject seemingly toxic or benign prompts by misclassifying them as toxic. This behavior undermines models' helpfulness and restricts usability in sensitive or nuanced contexts. While prior work has proposed mitigation strategies such as data augmentation and activation steering, these approaches often face a trade-off: reducing over-refusal typically degrades the model's ability to reject genuinely harmful content. We argue that this issue arises from the ambiguous influence of toxic and seemingly toxic prompts on the model's learning dynamics. To address it, we introduce a preceding alignment stage, DCR: Discernment via Contrastive Refinement. Both theoretically and empirically, we demonstrate that contrastive refinement improves an LLM's capacity to distinguish truly toxic prompts from superficially toxic ones. Evaluation across diverse benchmarks shows that our method effectively reduces over-refusal while preserving the safety benefits of alignment. Importantly, it achieves this with minimal degradation of general capabilities, offering a more principled and robust direction for safety alignment.

cs.CL