arXiv ScienceSearch

arXiv subjects

Wei Cai

Publications and source records attributed to Wei Cai.

At least 19 recordsLinked to original sources

CaLR: Causal Latent Revision for Robust Diffusion Reasoning

Autoregressive (AR) models suffer from local greediness, while diffusion language models (DLMs) often lack the strict causal structure required for reasoning. To combine the advantages and overcome the drawbacks of the dual, we propose Causal Latent Revision (CaLR), a framework that reformulates reasoning as constrained latent optimization. By adopting a causal topology matrix (CTM) from an expert model and implicit differentiation, CaLR performs gradient-guided ``thought revision" to enforce logical consistency, enabling dynamic self-correction of intermediate steps during parallel generation. Empirically, CaLR achieves SOTA DLM performance on complex benchmarks, surpassing strong AR baselines and demonstrating superior robustness in constrained tasks like Sudoku.

cs.AI

Weak Adversarial Neural Pushforward Method for Boltzmann Equation

In this paper, we extend a weak adversary neural network pushforward method for solving time dependent Boltzmann equation and a weak formulation of the collision operator is proposed where an invertible neural pushforward mapping is used to generating samples given by the distribution governed by the Boltzmann equation. The training of the pushforward mapping is learnt by enforcing the weak form of the Boltzmann equation. Numerical results have demonstrated the effectiveness of the proposed method.

math.NA

Neuro-Symbolic Participation Governance for Verifiable AI Agents in Open Digital Twin Ecosystems

Autonomous AI agents, increasingly empowered by large language models, are becoming important components of human-machine systems for high-stakes decision support in digital twin ecosystems. However, existing multi-agent systems often lack robust verification for identity, capability, and policy compliance, especially in decentralized environments spanning multiple institutions. This paper proposes a neuro-symbolic decentralized governance framework for verifiable agents in collaborative digital twin environments. By representing agents through multi-layer semantic profiles, the framework bridges probabilistic neural reasoning with deterministic institutional governance, thereby supporting trustworthy human-AI collaboration and meaningful human oversight. Capabilities are grounded in formal domain ontologies to enable machine-interpretable, policy-aware, and context-sensitive participation. These credentials, issued by organizational authorities, are validated via blockchain-based smart contracts, ensuring auditable participation without exposing sensitive data. We demonstrate the framework using a decision-support prototype with clinic, digital twin, and wearable provider agents effectively prevents unauthorized interaction and enforces institutional policies with manageable overhead. Our findings suggest that neuro-symbolic decentralized governance provides a scalable and trustworthy pathway for safe human-machine collaboration across institutional boundaries.

cs.CR

Weak Adversarial Neural Pushforward Method for the Wigner Transport Equation

We extend the Weak Adversarial Neural Pushforward Method to the Wigner transport equation governing the phase-space dynamics of quantum systems. The central contribution is a structural observation: integrating the nonlocal pseudo-differential potential operator against plane-wave test functions produces a Dirac delta that exactly inverts the Fourier transform defining the Wigner potential kernel, reducing the operator to a pointwise finite difference of the potential at two shifted arguments. This holds in arbitrary dimension, requires no truncation of the Moyal series, and treats the potential as a black-box function oracle with no derivative information. To handle the negativity of the Wigner quasi-probability distribution, we introduce a signed pushforward architecture that decomposes the solution into two non-negative phase-space distributions mixed with a learnable weight. The resulting method inherits the mesh-free, Jacobian-free, and scalable properties of the original framework while extending it to the quantum setting.

quant-ph

A convolutional neural network surrogate for hierarchical homogenization: fast elastic moduli prediction of digital rocks

Digital rock physics (DRP) aims to estimate effective rock properties (e.g., elastic moduli) directly from 3D micro-CT images. However, direct numerical simulations (DNS) on high-resolution large 3D scans are often computationally prohibitive and severely limit the application of DRP. To address this bottleneck, we combine a lightweight 3D convolutional neural network (CNN) with hierarchical homogenization (HHM) and apply it to determine effective elastic moduli. In this scheme, a large rock image is divided into subcubes. The CNN replaces costly DNS by directly predicting subcube elastic moduli, while HHM upscales subcube-level predictions to the full rock. Using a shared convolutional backbone, we systematically compare three training targets: (i) full anisotropic $6\times6$ stiffness tensors, (ii) isotropic bulk and shear moduli $(K, G)$, and (iii) Hashin--Shtrikman (HS)-normalized factors. Across multiple rock types, all three models agree well with DNS results while substantially reducing the computational cost. Moreover, training from scratch on each rock type is fast enough that transfer learning is unnecessary. Across all three targets, the accuracy is comparable. In our comparative study, the HS-normalized factor offers the best overall speed--accuracy trade-off while guaranteeing physical consistency, making it a convenient default. The isotropic $(K, G)$ target is a slightly more accurate alternative.

physics.comp-ph

Multiscale Fourier Neural Operator for Inverse Wave Scattering in Highly Oscillatory Media

In this paper, we propose an operator learning method based on the multiscale Fourier neural operator (MscaleFNO) for inverse medium problems of Helmholtz equations. The MscaleFNO provides a neural surrogate model with reduced spectral bias for the Helmholtz equations, mapping highly oscillatory medium profiles to scattered wavefields. A plug-and-play inversion using elucidated diffusion model is introduced to regularize the inverse solver based on least squares of data misfits. Numerical results for partial aperture inversion of oscillatory two-dimensional media demonstrate the advantage and effectiveness of MscaleFNO for accurate reconstruction of highly oscillatory medium properties.

math.NA

Crystal Dislocations as Atomic Scale Ratchets

The symmetry of a system's response to external stimuli is a fundamental concept in physics and materials science. At the microscopic scale, breaking this symmetry to achieve a rectified response is exceptionally difficult to engineer and remains rare in nature. Conventional micromechanics models of crystalline solids assume a symmetric response to applied stress, where reversing the load simply inverts the direction of defect velocity without altering its magnitude. In this work, we report an atomic-scale, geometry-rooted mechanism that breaks this symmetry. Molecular dynamics simulations of face-centered cubic nickel reveal that dislocations containing atomic-scale jogs exhibit asymmetric mobility under opposite applied stresses, reversing the loading direction triggers significantly higher drag. This asymmetry arises from an unconventional coupling between an atomic displacement vector and the second-order tensorial eigenstrain of the jog motion mechanism. Because jogs are ubiquitous structures in plastic deformation, this discovery challenges classical descriptions of plastic deformation mechanisms, with direct implications to cyclic creep, and opens new pathways for defect engineering to enhance fatigue resistance.

cond-mat.mtrl-sci

Why is the strength of an elastomeric polymer network so low?

Experiments have long shown that a polymer network of covalent bonds commonly ruptures at a stress that is orders of magnitude lower than the strength of the covalent bonds. Here we investigate this large reduction in strength by coarse-grained molecular dynamics simulations. We show that the network ruptures by sequentially breaking a small fraction of bonds, and that each broken bond lies on the minimum "shortest path". The shortest path is the path of the fewest bonds that connect two monomers at the opposite ends of the network. As the network is stretched, the minimum shortest path straightens and bears high tension set by covalent bonds, while most strands off the path deform by entropic elasticity. After a bond on the minimum shortest path breaks, the process repeats for the next minimum shortest path. As the network is stretched and bonds are broken, the scatter in lengths of the shortest paths first narrows, causing stress to rise, and then broadens, causing stress to decline. This sequential breaking of a small fraction of bonds causes the network to rupture at a stress that is orders of magnitude below the strength of the covalent bonds.

cond-mat.soft

DeRelayL: Sustainable Decentralized Relay Learning

In the era of big data, large-scale machine learning models have revolutionized various fields, driving significant advancements. However, large-scale model training demands high financial and computational resources, which are only affordable by a few technological giants and well-funded institutions. In this case, common users like mobile users, the real creators of valuable data, are often excluded from fully benefiting due to the barriers, while the current methods for accessing large-scale models either limit user ownership or lack sustainability. This growing gap highlights the urgent need for a collaborative model training approach, allowing common users to train and share models. However, existing collaborative model training paradigms, especially federated learning (FL), primarily focus on data privacy and group-based model aggregation. To this end, this paper intends to address this issue by proposing a novel training paradigm named decentralized relay learning (DeRelayL), a sustainable learning system where permissionless participants can contribute to model training in a relay-like manner and share the model. In detail, this paper presents the architecture and workflow of DeRelayL, designs incentive mechanisms to ensure sustainability, and conducts theoretical analysis and numerical simulations to demonstrate its effectiveness.

cs.LG

HiEdit: Lifelong Model Editing with Hierarchical Reinforcement Learning

Lifelong model editing (LME) aims to sequentially rectify outdated or inaccurate knowledge in deployed LLMs while minimizing side effects on unrelated inputs. However, existing approaches typically apply parameter perturbations to a static and dense set of LLM layers for all editing instances. This practice is counter-intuitive, as we hypothesize that different pieces of knowledge are stored in distinct layers of the model. Neglecting this layer-wise specificity can impede adaptability in integrating new knowledge and result in catastrophic forgetting for both general and previously edited knowledge. To address this, we propose HiEdit, a hierarchical reinforcement learning framework that adaptively identifies the most knowledge-relevant layers for each editing instance. By enabling dynamic, instance-aware layer selection and incorporating an intrinsic reward for sparsity, HiEdit achieves precise, localized updates. Experiments on various LLMs show that HiEdit boosts the performance of the competitive RLEdit by an average of 8.48% with perturbing only half of the layers per edit. Our code is available at: https://github.com/yangfanww/hiedit.

cs.CL

Learning Interatomic Force Coefficients from X-ray Thermal Diffuse Scattering Data

We present a fully automated framework for extracting interatomic force constants (IFCs) directly from X-ray thermal diffuse scattering (TDS) data. By formulating scattering intensity as a differentiable function of a symmetry-reduced IFC parameterization, we enable gradient-based optimization via direct, Cholesky-based sampling of correlated atomic displacements at thermal equilibrium. This approach bypasses the computational bottleneck of repeated Hessian matrix diagonalizations, significantly accelerating the inversion process. Benchmark tests demonstrate that the framework accurately recovers ground-truth IFCs and phonon dispersion relations, providing a robust, high-throughput pathway for studying lattice dynamics across diverse crystalline materials. This method bridges the gap between experimental observations and computational modeling, enabling the direct integration of TDS data into the refinement of high-fidelity inter-atomic potentials.

physics.comp-ph

Weak Adversarial Neural Pushforward Method for the McKean-Vlasov / Mean-Field Fokker-Planck Equation

We extend the Weak Adversarial Neural Pushforward Method (WANPM) to the McKean--Vlasov mean-field Fokker--Planck equation, covering both the stationary and time-dependent cases. The key observation is that the mean-field nonlinearity -- an expectation under the solution distribution -- is naturally estimated by Monte Carlo sampling from the pushforward network, requiring no change to the architecture and only minor modifications to the training loop. For the quadratic (granular media) interaction kernel, the interaction term reduces to the batch sample mean, eliminating secondary sampling entirely. We also identify a dimension-dependent frequency initialization rule for the adversarial test functions, necessary to avoid spurious minimizers. Numerical experiments on linear McKean--Vlasov benchmarks in 2, 5, 20, and 100 dimensions confirm accurate recovery of the exact Gaussian stationary and transient distributions, with training times ranging from 27 seconds (2D) to 10 minutes (100D) on a single GPU.

math.NA

Weak Adversarial Neural Pushforward Method for Fractional Fokker-Planck Equations

We extend the Weak Adversarial Neural Pushforward Method (WANPM) to fractional Fokker-Planck equations, in which the classical Laplacian diffusion operator is replaced by the fractional Laplacian of order alpha in (0, 2]. The solution distribution is represented as the pushforward of a simple base distribution through a neural network, and the weak formulation is discretized entirely via Monte Carlo sampling without any temporal mesh. A key computational advantage is that plane-wave test functions are eigenfunctions of the fractional Laplacian, making the operator cost identical to that of classical diffusion for any alpha. We validate the method on seven benchmark problems with alpha = 1.5, spanning one and two spatial dimensions: the steady-state fractional Ornstein--Uhlenbeck (OU) process, a harmonic confining potential, a double-well potential, and a triple-well potential in one dimension, a steady-state 2D double-peak distribution, a time-dependent 2D ring distribution with rotational drift, and a five-dimensional harmonic potential. Each case is benchmarked against particle simulations using symmetric alpha-stable Lévy increments, and robust statistics confirm close agreement throughout. The method is mesh-free, requires no density evaluation or non-local quadrature, and provides a promising foundation for high-dimensional anomalous diffusion solvers.

math.NA

Neural Pushforward Samplers for the Fokker-Planck Equation on Embedded Riemannian Manifolds

In this paper, we extend the Weak Adversarial Neural Pushforward Method to the Fokker--Planck equation on compact embedded Riemannian manifolds. The method represents the solution as a probability distribution via a neural pushforward map that is constrained to the manifold by a retraction layer, enforcing manifold membership and probability conservation by construction. Training is guided by a weak adversarial objective using ambient plane-wave test functions, whose intrinsic differential operators are derived in closed form from the geometry of the embedding, yielding a fully mesh-free and chart-free algorithm. Both steady-state and time-dependent formulations are developed, and numerical results on a double-well problem on the two-sphere demonstrate the capability of the method in capturing multimodal invariant distributions on curved spaces.

math.NA

From Connectivity to Rupture: A Coarse-Grained Stochastic Network Dynamics Approach to Polymer Network Mechanics

We introduce a coarse-grained stochastic network dynamics (CGSND) framework for modeling deformation and rupture in polymer networks. The method replaces explicit molecular dynamics (MD) or coarse-grained molecular dynamics (CGMD) with network-level evolution rules while retaining chain entropic elasticity and force-controlled bond failure. Under uniaxial loading, CGSND reproduces the characteristic nonlinear stress--stretch response of elastomeric networks, including a well-defined ultimate tensile strength and post-peak softening due to progressive bond rupture. Comparison with coarse-grained molecular dynamics (CGMD) simulations shows that CGSND captures the qualitative form of the stress response and the onset of catastrophic damage despite its rate-independent formulation. Analysis of rupture kinetics reveals a pronounced peak in the bond-breaking hazard rate near the ultimate tensile strength in both approaches. In addition, the distribution of broken segment lengths remains statistically indistinguishable from the initial network, indicating that rupture is not biased toward short or long chains. Finally, the evolution of the Gini coefficient of bond force magnitudes reveals strong force localization preceding failure. These results demonstrate that CGSND provides a computationally efficient and physically interpretable framework for connecting force localization and rupture kinetics to macroscopic failure in polymer networks.

cond-mat.soft

Accelerated Markov Chain Monte Carlo Simulation via Neural Network-Driven Importance Sampling

Atomistic simulations provide valuable insights into the physical processes governing material behavior. However, their applicability is fundamentally constrained by the limited time scales accessible to brute-force simulations. This bottleneck often stems from complex energy landscapes where the systems stay trapped in metastable states for long periods of time. Yet, the long-term evolution is controlled by the transitions between the metastable states, which are rare events and difficult to observe. We present an importance sampling method designed to accelerate the time scale of Markov chain Monte Carlo (MCMC) simulations. By employing a bias potential, our approach enhances the sampling of rare transition events while preserving the relative probabilities of distinct transition pathways. The bias potential is represented by a neural network which enables the flexibility needed for high-dimensional systems. We propose a rigorous formulation to obtain the original transition rates between metastable states using transition paths obtained from the biased simulation. We further use a branching random walk (BRW) technique to enhance efficiency and to reduce variance. The proposed methodology is validated on 2-dimensional and 14-dimensional systems, demonstrating its accuracy and scalability.

physics.comp-ph

Link Statistics of Dislocation Network during Strain Hardening

Dislocations are line defects in crystals that multiply and self-organize into a complex network during strain hardening. The length of dislocation links, connecting neighboring nodes within this network, contains crucial information about the evolving dislocation microstructure. By analyzing data from Discrete Dislocation Dynamics (DDD) simulations in face-centered cubic (fcc) Cu, we characterize the statistical distribution of link lengths of dislocation networks during strain hardening on individual slip systems. Our analysis reveals that link lengths on active slip systems follow a double-exponential distribution, while those on inactive slip systems conform to a single-exponential distribution. The distinctive long tail observed in the double-exponential distribution is attributed to the stress-induced bowing out of long links on active slip systems, a feature that disappears upon removal of the applied stress. We further demonstrate that both observed link length distributions can be explained by extending a one-dimensional Poisson process to include different growth functions. Specifically, the double-exponential distribution emerges when the growth rate for links exceeding a critical length becomes super-linear, which aligns with the physical phenomenon of long links bowing out under stress. This work advances our understanding of dislocation microstructure evolution during strain hardening and elucidates the underlying physical mechanisms governing its formation.

cond-mat.mtrl-sci

Visual Attention Reasoning via Hierarchical Search and Self-Verification

Multimodal Large Language Models (MLLMs) frequently hallucinate due to their reliance on fragile, linear reasoning and weak visual grounding. We propose Visual Attention Reasoning (VAR), a reinforcement learning framework that reformulates reasoning as a hierarchical search with self-verification. VAR enforces traceable evidence grounding by generating explicit bounding boxes, guided by a novel reward function combining geometric precision and semantic sufficiency. Furthermore, it replaces linear Chain-of-Thought with a tree-search policy capable of backtracking to correct logical errors. Theoretical analysis validates the framework's reliability, and extensive experiments demonstrate that VAR significantly outperforms state-of-the-art methods on complex hallucination and safety benchmarks.

cs.AI