arXiv ScienceSearch

arXiv subjects

Xicheng Zhang

Publications and source records attributed to Xicheng Zhang.

At least 19 recordsLinked to original sources

Well-Posedness for SDEs with Logarithmical Critical Distributional Drifts

We study the stochastic differential equation $$d X_t=b(t,X_t)d t+\sqrt{2}d W_t$$ on $\mathbb R^d$, where $b$ is a time-dependent, divergence-free distributional drift of critical H\"older--Besov regularity $-1$, strengthened by an iterated-logarithmic correction. For every initial probability law, we construct a weak solution by smooth approximation and realize the singular drift as an additive functional. The main analytic ingredient is the Schauder estimate with a logarithmic smallness factor. Combined with uniform logarithmic Krylov estimates and a stochastic substitution formula for distributional test functions, this estimate allows us to apply a Zvonkin transformation and prove uniqueness in law among weak solutions satisfying the corresponding Krylov bounds. For solutions starting from deterministic points, we further show that their time-marginal distributions admit densities satisfying two-sided Aronson-type Gaussian estimates.

math.PR

DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning

Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limited supervision on how local conclusions should support later decisions. We introduce Dependency-Aware Intermediate QA Supervision (DAIS), a training-time framework that converts filtered teacher rationales into stage-level QA records. Each intermediate record predicts a local answer conditioned on the previous states needed for that decision, while the final-answer record keeps the original task format; evaluation therefore uses only the original input and optional context. Across GDPR, AIACT, MedQA, and FOLIO with multiple Qwen backbones, DAIS improves average final-answer accuracy over answer-only, flat chain-of-thought, and independent-QA baselines. On policy-compliance benchmarks, it achieves a largest gain of 5.6% and an average gain of 4.2% over the strongest non-DAIS baseline. Controlled ablations show that valid previous-state conditioning contributes beyond longer targets or additional intermediate text, supporting dependency-conditioned intermediate QA as a lightweight auxiliary supervision signal for standard final-answer inference.

cs.CL

Quantitative Propagation of Chaos and Fluctuations for Kinetic McKean--Vlasov SDEs with Singular Interaction Kernels

We prove a quantitative propagation of chaos estimate and a central limit theorem for the particle system associated with a class of degenerate kinetic McKean--Vlasov SDEs with external drifts and singular interaction kernels in Kato's class. In particular, the interaction kernel can be in the mixed $L^q_tL^{p_v}_vL^{p_x}_x$-space, where $\frac2q+\frac{3d}{p_x}+\frac d{p_v}<1$. For the associated $N$-particle system, we obtain a path-space relative entropy bound of order $k/N$ for the first $k$ particles, assuming only entropic chaoticity of the initial data. The key ingredients are kinetic Krylov--Khasminskii estimates and a conditional Hilbert-space subgaussian estimate for empirical interaction fields. For the CLT, we also prove a Berry--Esseen-type bound for finite-dimensional projections.

math.PR

Kinetic Fokker-Planck Equations with Nonlinear Diffusion

We study existence, regularity, and uniqueness for the nonlinear kinetic Fokker--Planck equation $$ \partial_t f=\Delta_v\Psi(f)-v\cdot\nabla_x f, \qquad f|_{t=0}=f_0, $$ on $\mathbb R^{2d}$. In the model case $\Psi(r)=r^s$, this equation couples nonlinear fast-diffusion/porous-medium type diffusion with kinetic transport. A distinctive feature is that the diffusion acts only in the velocity variable $v$, so that compactness in the spatial variable $x$ cannot be obtained from standard elliptic estimates and must instead be recovered through the hypoelliptic structure. Under general structural assumptions on $\Psi$, including the fast-diffusion powers $\Psi(r)=r^s$ with $s\in(0,1)$, we construct nonnegative weak solutions and prove quantitative anisotropic Besov regularity estimates. Under an additional mass-critical growth condition on the fast-diffusion side, the constructed weak solution preserves mass, admits a renormalized kinetic formulation, and is unique in the $L^1$-class of mass-preserving renormalized kinetic solutions. In the power-law case $\Psi(r)=r^s$, this condition is precisely $s\ge 1-1/d$ when $d\ge2$, while in dimension $d=1$ the whole fast-diffusion range $s\in(0,1)$ is covered. The main analytic ingredient is a parameter-dependent smoothing estimate for the kinetic semigroup generated by $$ \Psi'(\zeta)\Delta_v - v\cdot\nabla_x , $$ which quantitatively tracks the dependence on the kinetic level $\zeta$. Combined with the kinetic formulation, this estimate yields compactness in both spatial and velocity variables for the nonlinear hypoelliptic problem. As an application, we also obtain martingale-problem solutions to the associated distributional-density dependent stochastic differential equation.

math.AP

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studies have demonstrated the feasibility of backdoor attacks against LLMs. However, existing methods suffer from three key shortcomings: explicit trigger patterns that compromise naturalness, unreliable injection of attacker-specified payloads in long-form generation, and incompletely specified threat models that obscure how backdoors are delivered and activated in practice. To address these gaps, we present BadStyle, a complete backdoor attack framework and pipeline. BadStyle leverages an LLM as a poisoned sample generator to construct natural and stealthy poisoned samples that carry imperceptible style-level triggers while preserving semantics and fluency. To stabilize payload injection during fine-tuning, we design an auxiliary target loss that reinforces the attacker-specified target content in responses to poisoned inputs and penalizes its emergence in benign responses. We further ground the attack in a realistic threat model and systematically evaluate BadStyle under both prompt-induced and PEFT-based injection strategies. Extensive experiments across seven victim LLMs, including LLaMA, Phi, DeepSeek, and GPT series, demonstrate that BadStyle achieves high attack success rates (ASRs) while maintaining strong stealthiness. The proposed auxiliary target loss substantially improves the stability of backdoor activation, yielding an average ASR improvement of around 30% across style-level triggers. Even in downstream deployment scenarios unknown during injection, the implanted backdoor remains effective. Moreover, BadStyle consistently evades representative input-level defenses and bypasses output-level defenses through simple camouflage.

cs.CR

Derivative estimates for SDEs with singular and unbounded coefficients

We develop a unified PDE-probabilistic framework for pointwise gradient and Hessian estimates of Markov semigroups associated with stochastic differential equations with singular and unbounded coefficients. Under mild local structural assumptions on the diffusion matrix and integrability/regularity conditions on the drift, we obtain quantitative sharp short-time regularization estimates as well as long-time decay bounds (including exponential and polynomial rates) for the first and second spatial derivatives of the semigroup. A distinctive feature of our results is the explicit dependence of these estimates on local norms of the coefficients (through scale-invariant quantities), without requiring any global smoothness, boundedness or uniform ellipticity. In particular, our approach allows for degenerate or highly irregular behavior at infinity, subject to suitable local ellipticity and Lyapunov/ergodicity controls. As applications, we establish solvability and regularity results for Poisson equations on the whole space with singular coefficients, and we derive pointwise gradient estimates for SDEs with distributional drifts via a Zvonkin-type transform.

math.PR

Uniform-in-time diffusion approximations for multiscale stochastic systems

This paper establishes a quantitative, uniform-in-time diffusion approximation for the joint law of a broad class of fully coupled multiscale stochastic systems. We derive a precise characterization of the limiting joint distribution as a specific skew-product of the conditional equilibrium of the fast process and the homogenized law of the slow component, thereby providing a rigorous uniform-in-time formulation of the adiabatic elimination principle. The convergence rate explicitly separates the initial relaxation of the fast dynamics from the long-time homogenized evolution and depends only on the regularity of the coefficients in the slow variable. As a consequence, we obtain the first quantitative identification of the limiting stationary distribution of the original multiscale system and prove the commutativity of the limits $\eps\to0$ and $t\to\infty$ for a large class of observables. Our framework accommodates unbounded and irregular coefficients, degenerate structures, and weakly mixing dynamics. We illustrate its scope with three applications: {\it (i)} a uniform-in-time averaging principle for fast-slow systems; {\it (ii)} a uniform Smoluchowski--Kramers approximation for degenerate Langevin systems, yielding convergence of the joint position-scaled velocity law and global-in-time asymptotics of key thermodynamic functionals (e.g., total energy, entropy production, free energy); and {\it (iii)} the first uniform-in-time periodic homogenization result for SDEs with distributional drifts.

math.PR

Harnack inequalities for nonlocal operators with supercritical drifts and their applications

In this paper, we investigate Harnack estimates for weak solutions to the following nonlocal equation: $$ \partial_t u = \Delta^{\alpha/2} u + b \cdot \nabla u + f, $$ where $\Delta^{\alpha/2}$ denotes the fractional Laplacian, $b$ is a divergence-free vector field in a critical or supercritical regularity regime, and $f$ is a distribution in a fractional Sobolev space with negative indices. As applications of the analytical results obtained in this paper, we establish the well-posedness of critical stochastic quasi-geostrophic equations driven by additive Brownian noise, prove the existence of weak solutions to the two-dimensional fractional Navier--Stokes equations with measure-valued initial vorticity, and demonstrate the well-posedness of generalized martingale problems associated with critical stochastic differential equations.

math.AP

Strong approximation for stochastic Volterra equations by compound Poisson processes

We study a compound Poisson (random time-change) approximation for stochastic differential equations (SDEs) and stochastic Volterra equations whose coefficients may be merely measurable in time and may even exhibit integrable singularities. For an SDE driven by Brownian motion, we replace the time variable by the Poisson clock $\mathcal{N}_t^\varepsilon$ and approximate the stochastic integral by $W_{\mathcal{N}_t^\varepsilon}$, which leads to an explicit jump scheme driven by a compensated Poisson random measure. Under standard Lipschitz and linear-growth conditions in the state variable (with no continuity assumed in time for the drift), we prove strong convergence and obtain explicit rates in $\varepsilon$. For Volterra-type equations with singular kernels, we establish strong convergence as well, with a rate that reflects both the temporal regularity of the kernel and the intrinsic $\varepsilon^{1/2}$ fluctuation of the Poisson clock. The compound Poisson scheme differs fundamentally from the Euler-Maruyama method: it does not require pointwise evaluation of time-irregular coefficients on a deterministic grid, and it remains stable in the presence of time singularities. We further illustrate the theory on stochastic Volterra equations driven by fractional Brownian motion and provide numerical experiments showing improved performance over Euler-Maruyama for problems with singular time dependence.

math.PR

Anchored Langevin Algorithms

Standard first-order Langevin algorithms such as the unadjusted Langevin algorithm (ULA) are obtained by discretizing the Langevin diffusion and are widely used for sampling in machine learning because they scale to high dimensions and large datasets. However, they face two key limitations: (i) they require differentiable log-densities, excluding targets with non-differentiable components; and (ii) they generally fail to sample heavy-tailed targets. We propose anchored Langevin dynamics, a unified approach that accommodates non-differentiable targets and certain classes of heavy-tailed distributions. The method replaces the original potential with a smooth reference potential and modifies the Langevin diffusion via multiplicative scaling. We establish non-asymptotic guarantees in the 2-Wasserstein distance to the target distribution and provide an equivalent formulation derived via a random time change of the Langevin diffusion. We provide numerical experiments to illustrate the theory and practical performance of our proposed approach.

stat.ML

Sampling-Based Zero-Order Optimization Algorithms

We propose a novel zeroth-order optimization algorithm based on an efficient sampling strategy. Under mild global regularity conditions on the objective function, we establish non-asymptotic convergence rates for the proposed method. Comprehensive numerical experiments demonstrate the algorithm's effectiveness, highlighting three key attributes: (i) Scalability: consistent performance in high-dimensional settings (exceeding 100 dimensions); (ii) Versatility: robust convergence across a diverse suite of benchmark functions, including Schwefel, Rosenbrock, Ackley, Griewank, L\'evy, Rastrigin, and Weierstrass; and (iii) Robustness to discontinuities: reliable performance on non-smooth and discontinuous landscapes. These results illustrate the method's strong potential for black-box optimization in complex, real-world scenarios.

math.OC

LLM-Driven Self-Refinement for Embodied Drone Task Planning

We introduce SRDrone, a novel system designed for self-refinement task planning in industrial-grade embodied drones. SRDrone incorporates two key technical contributions: First, it employs a continuous state evaluation methodology to robustly and accurately determine task outcomes and provide explanatory feedback. This approach supersedes conventional reliance on single-frame final-state assessment for continuous, dynamic drone operations. Second, SRDrone implements a hierarchical Behavior Tree (BT) modification model. This model integrates multi-level BT plan analysis with a constrained strategy space to enable structured reflective learning from experience. Experimental results demonstrate that SRDrone achieves a 44.87% improvement in Success Rate (SR) over baseline methods. Furthermore, real-world deployment utilizing an experience base optimized through iterative self-refinement attains a 96.25% SR. By embedding adaptive task refinement capabilities within an industrial-grade BT planning framework, SRDrone effectively integrates the general reasoning intelligence of Large Language Models (LLMs) with the stringent physical execution constraints inherent to embodied drones. Code is available at https://github.com/ZXiiiC/SRDrone.

cs.RO

Kinetic SDEs with subcritical distributional drifts

In this paper we study the well-posedness of the kinetic stochastic differential equation (SDE) in $\mathbb R^{2d}(d\geq2)$ driven by Brownian motion: $$\mathord{{\rm d}} X_t=V_t\mathord{{\rm d}} t,\ \mathord{{\rm d}} V_t=b(t,X_t,V_t)\mathord{{\rm d}} t+\sqrt{2}\mathord{{\rm d}} W_t,$$ where the subcritical distribution-valued drift $b$ belongs to the weighted anisotropic H\"{o}lder space $\mathbb L_T^{q_b}\mathbf C_{\boldsymbol{a}}^{\alpha_b}(\rho_\kappa)$ with parameters $\alpha_b\in(-1,0)$, $q_b\in(\frac{2}{1+\alpha_b},\infty]$, $\kappa\in[0,1+\alpha_b)$ and $\div_v b$ is bounded. We establish the well-posedness of weak solutions to the associated integral equation: $$X_t=X_0+\int_0^t V_s\mathord{{\rm d}} s,\ V_t=V_0+\lim_{n\to\infty}\int_0^t b_n(s,X_s,V_s)\mathord{{\rm d}}+\sqrt{2}W_t,$$ where $b_n:=b*\Gamma_n$ denotes the mollification of $b$ and the limit is taken in the $L^2$-sense. As an application, we discuss examples of $b$ involving Gaussian random fields.

math.PR

RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware Reasoning

The integration of external knowledge through Retrieval-Augmented Generation (RAG) has become foundational in enhancing large language models (LLMs) for knowledge-intensive tasks. However, existing RAG paradigms often overlook the cognitive step of applying knowledge, leaving a gap between retrieved facts and task-specific reasoning. In this work, we introduce RAG+, a principled and modular extension that explicitly incorporates application-aware reasoning into the RAG pipeline. RAG+ constructs a dual corpus consisting of knowledge and aligned application examples, created either manually or automatically, and retrieves both jointly during inference. This design enables LLMs not only to access relevant information but also to apply it within structured, goal-oriented reasoning processes. Experiments across mathematical, legal, and medical domains, conducted on multiple models, demonstrate that RAG+ consistently outperforms standard RAG variants, achieving average improvements of 3-5%, and peak gains up to 13.5% in complex scenarios. By bridging retrieval with actionable application, RAG+ advances a more cognitively grounded framework for knowledge integration, representing a step toward more interpretable and capable LLMs.

cs.AI

Stochastic Transport Maps in Diffusion Models and Sampling

In this work, we present a theoretical and computational framework for constructing stochastic transport maps between probability distributions using diffusion processes. We begin by proving that the time-marginal distribution of the sum of two independent diffusion processes satisfies a Fokker-Planck equation. Building on this result and applying Ambrosio-Figalli-Trevisan's superposition principle, we establish the existence and uniqueness of solutions to the associated stochastic differential equation (SDE). Leveraging these theoretical foundations, we develop a method to construct (stochastic) transport maps between arbitrary probability distributions using dynamical ordinary differential equations (ODEs) and SDEs. Furthermore, we introduce a unified framework that generalizes and extends a broad class of diffusion-based generative models and sampling techniques. Finally, we analyze the convergence properties of particle approximations for the SDEs underlying our framework, providing theoretical guarantees for their practical implementation. This work bridges theoretical insights with practical applications, offering new tools for generative modeling and sampling in high-dimensional spaces.

math.PR

Stealthy Backdoor Attack to Real-world Models in Android Apps

Powered by their superior performance, deep neural networks (DNNs) have found widespread applications across various domains. Many deep learning (DL) models are now embedded in mobile apps, making them more accessible to end users through on-device DL. However, deploying on-device DL to users' smartphones simultaneously introduces several security threats. One primary threat is backdoor attacks. Extensive research has explored backdoor attacks for several years and has proposed numerous attack approaches. However, few studies have investigated backdoor attacks on DL models deployed in the real world, or they have shown obvious deficiencies in effectiveness and stealthiness. In this work, we explore more effective and stealthy backdoor attacks on real-world DL models extracted from mobile apps. Our main justification is that imperceptible and sample-specific backdoor triggers generated by DNN-based steganography can enhance the efficacy of backdoor attacks on real-world models. We first confirm the effectiveness of steganography-based backdoor attacks on four state-of-the-art DNN models. Subsequently, we systematically evaluate and analyze the stealthiness of the attacks to ensure they are difficult to perceive. Finally, we implement the backdoor attacks on real-world models and compare our approach with three baseline methods. We collect 38,387 mobile apps, extract 89 DL models from them, and analyze these models to obtain the prerequisite model information for the attacks. After identifying the target models, our approach achieves an average of 12.50% higher attack success rate than DeepPayload while better maintaining the normal performance of the models. Extensive experimental results demonstrate that our method enables more effective, robust, and stealthy backdoor attacks on real-world models.

cs.CR

$W_{\bf d}$-convergence rate of EM schemes for invariant measures of supercritical stable SDEs

By establishing the regularity estimates for nonlocal Stein/Poisson equations under $\gamma$-order H\"older and dissipative conditions on the coefficients, we derive the $W_{\bf d}$-convergence rate for the Euler-Maruyama schemes applied to the invariant measure of SDEs driven by multiplicative $\alpha$-stable noises with $\alpha \in (\frac{1}{2}, 2)$, where $W_{\bf d}$ denotes the Wasserstein metric with ${\bf d}(x,y)=|x-y|^\gamma\wedge 1$ and $\gamma \in ((1-\alpha)_+, 1]$.

math.PR

Heat kernel estimates for nonlocal kinetic operators

In this paper, we employ probabilistic techniques to derive sharp, explicit two-sided estimates for the heat kernel of the nonlocal kinetic operator $$ \Delta^{\alpha/2}_v + v \cdot \nabla_x, \quad \alpha \in (0, 2),\ (x,v)\in {\mathbb R}^{d}\times{\mathbb R}^d,$$ where $ \Delta^{\alpha/2}_v $ represents the fractional Laplacian acting on the velocity variable $v$. Additionally, we establish logarithmic gradient estimates with respect to both the spatial variable $x$ and the velocity variable $v$. In fact, the estimates are developed for more general non-symmetric stable-like operators, demonstrating explicit dependence on the lower and upper bounds of the kernel functions. These results, in particular, provide a solution to a fundamental problem in the study of \emph{nonlocal} kinetic operators.

math.PR