arXiv ScienceSearch

arXiv subjects

Huan Lei

Publications and source records attributed to Huan Lei.

At least 19 recordsLinked to original sources

Learning a general class of admissible multi-species collision operators from molecular dynamics

We develop a structure-preserving, data-driven collision operator for spatially homogeneous multi-species kinetic systems from molecular dynamics (MD). The operator consists of diagonal self-collision blocks and ordered off-diagonal cross-species blocks to describe intra- and inter-species momentum and energy exchange. Within a local and point-wise identifiable kernel class, we develop the necessary and sufficient condition for the admissible kernel class satisfying the conservation laws, the H-theorem, and the frame indifference. Unlike the classical Landau operator, the off-diagonal kernels are not restricted to be symmetric under permutation of the two velocity variables. This unique structural freedom captures the distinct responses of different species to unresolved correlations and many-body effects arising from micro-scale particle interactions. The equivalent parameterizable kernel formalization enables us to learn a generalized data-driven collision operator directly from MD, where the low-rank tensor representations and random sampling are used to achieve efficient kernel training and numerical simulation. Numerical experiments show that the learned operator accurately predicts transport coefficients and the non-equilibrium relaxation, while retaining discrete conservation and entropy production. In particular, it captures plasma kinetics in the moderately coupled regime, where the predictions of both the Landau and the data-driven model restricted to velocity-permutation symmetry show significant discrepancies.

physics.comp-ph

High-Dimensional Enhanced Sampling via Regularized Path-Dependent McKean--Vlasov Dynamics using Tensor Density Approximation

Sampling from high-dimensional Gibbs measures poses a challenge when the energy landscape consists of multiple metastable states. Enhanced-sampling methods mitigate this difficulty by introducing adaptive biasing potentials to facilitate the exploration along prescribed collective variables (CVs), but their scalability is often limited by the dimension of the CV space. Motivated by the Wasserstein-gradient-flow interpretation of adaptive biasing, we propose a regularized path-dependent McKean--Vlasov formulation for high-dimensional enhanced sampling. The formulation replaces the variational regularization of the Wasserstein functional by a direct regularization of the CV marginal density in the McKean--Vlasov drift, avoiding the outer convolution over the CV domain. Furthermore, it replaces the instantaneous law by a weighted path-history measure to improve statistical stability in the small-replica regime. We establish well-posedness of the resulting regularized and path-dependent stochastic dynamics under suitable assumptions. For numerical realization, the history-averaged CV marginal density is approximated using an optimization-free functional hierarchical tensor representation, leading to a scalable density-based adaptive biasing scheme. Numerical experiments on benchmark potentials and molecular systems demonstrate the effectiveness of the proposed method for sampling problems with CV dimensions up to 64.

math.NA

From molecular dynamics to kinetic models: data-driven generalized collision operators in 1D3V plasmas

We present a data-driven approach for constructing generalized collisional kinetic models for inhomogeneous plasmas in one-dimensional physical space and three-dimensional velocity space (1D-3V). The collision operator is directly learned from micro-scale molecular dynamics (MD) and accurately accounts for the unresolved particle interactions over a broad range of plasma conditions. Unlike the standard Landau operator, the present operator takes an anisotropic, non-stationary form that captures the heterogeneous collisional energy transfer arising from the many-body interactions, which is crucial for plasma kinetics beyond the weakly coupled regime. Efficient numerical evaluation is achieved through a low-rank tensor representation with $O(N \log N)$ computational complexity. The constructed kinetic equation strictly preserves conservation laws and physical constraints and therefore, enables us to develop an explicit second-order, energy-conserving scheme that ensures fully discrete conservation of mass and total energy. Numerical results demonstrate that the present model accurately predicts both transport coefficients and several 1D-3V kinetic processes compared with MD simulations across a broad range of densities and temperatures in spatially inhomogeneous settings. This work provides a systematic pathway for bridging micro-scale MD and inhomogeneous plasma kinetic descriptions where empirical models show limitation.

physics.plasm-ph

Matrix-Free Stabilized BDF Schemes for Semilinear Parabolic Equations with Unconditional Maximum Bound Principle Preservation and Energy Stability

We develop a family of stabilized backward differentiation formula (sBDF) schemes of orders one through four for semilinear parabolic equations. The proposed methods are designed to achieve three properties that are rarely available simultaneously in high-order time discretizations: unconditional preservation of the maximum bound principle (MBP), unconditional discrete energy stability, and practical matrix-free implementation. The construction integrates carefully designed stabilization terms, fixed-point iterations, and a pointwise cut-off strategy. The nonlinear algebraic systems arising from the implicit sBDF discretizations are solved by fixed-point iteration, resulting in fully matrix-free algorithms. This makes the approach particularly attractive for practical computations on general domains and under mixed boundary conditions, where FFT-based exponential time differencing methods are often unavailable or inefficient. We further present a unified analysis for the fully implemented schemes, explicitly incorporating the interplay among time discretization, nonlinear iteration, and cut-off. Unconditional contractivity of the fixed-point iterations and error estimates are established. For the Allen-Cahn equation, we additionally prove an unconditional discrete energy dissipation law. Numerical experiments confirm the theoretical convergence rates and demonstrate the robustness and efficiency of the proposed methods, particularly relative to ETD-based approaches for problems with mixed boundary conditions.

math.NA

A stochastic branching particle method for solving non-conservative reaction-diffusion equations

We propose a stochastic branching particle-based method for solving nonlinear non-conservative advection-diffusion-reaction equations. The method splits the evolution into an advection-diffusion step, based on a linearized Kolmogorov forward equation and approximated by stochastic particle transport, and a reaction step implemented through a branching birth-death process that provides a consistent temporal discretization of the underlying reaction dynamics. This construction yields a mesh-free, nonnegativity-preserving scheme that naturally accommodates non-conservative systems and remains robust in the presence of singularities or blow-up. We validate the method on two representative two-dimensional systems: the Allen-Cahn equation and the Keller-Segel chemotaxis model. In both cases, the present method accurately captures nonlinear behaviors such as phase separation and aggregation, and achieves reliable performance without the need for adaptive mesh refinement.

math.NA

Fast spectral separation method for kinetic equation with anisotropic non-stationary collision operator retaining micro-model fidelity

We present a generalized, data-driven collisional operator for one-component plasmas, learned from molecular dynamics simulations, to extend the collisional kinetic model beyond the weakly coupled regime. The proposed operator features an anisotropic, non-stationary collision kernel that accounts for particle correlations typically neglected in classical Landau formulations. To enable efficient numerical evaluation, we develop a fast spectral separation method that represents the kernel as a low-rank tensor product of univariate basis functions. This formulation admits an $O(N \log N)$ algorithm via fast Fourier transforms and preserves key physical properties, including discrete conservation laws and the H-theorem, through a structure-preserving central difference discretization. Numerical experiments demonstrate that the proposed model accurately captures plasma dynamics in the moderately coupled regime beyond the standard Landau model while maintaining high computational efficiency and structure-preserving properties.

math.NA

A unified framework for data-driven construction of stochastic reduced models with state-dependent memory

We present a unified framework for the data-driven construction of stochastic reduced models with state-dependent memory for high-dimensional Hamiltonian systems. The method addresses two key challenges: (\rmnum{1}) accurately modeling heterogeneous non-Markovian effects where the memory function depends on the coarse-grained (CG) variables beyond the standard homogeneous kernel, and (\rmnum{2}) efficiently exploring the phase space to sample both equilibrium and dynamical observables for reduced model construction. Specifically, we employ a consensus-based sampling method to establish a shared sampling strategy that enables simultaneous construction of the free energy function and collection of conditional two-point correlation functions used to learn the state-dependent memory. The reduced dynamics is formulated as an extended Markovian system, where a set of auxiliary variables, interpreted as non-Markovian features, is jointly learned to systematically approximate the memory function using only two-point statistics. The constructed model yields a generalized Langevin-type formulation with an invariant distribution consistent with the full dynamics. We demonstrate the effectiveness of the proposed framework on a two-dimensional CG model of an alanine dipeptide molecule. Numerical results on the transition dynamics between metastable states show that accurately capturing state-dependent memory is essential for predicting non-equilibrium kinetic properties, whereas the standard generalized Langevin model with a homogeneous kernel exhibits significant discrepancies.

physics.comp-ph

Hybrid Explicit-Implicit Predictor-Corrector Exponential Time-Differencing Multistep Pad\'{e} Schemes for Semilinear Parabolic Equations with Time-Delay

In this paper, we propose and analyze ETD-Multistep-Pad\'{e} (ETD-MS-Pad\'{e}) and ETD Implicit Multistep-Pad\'{e} (ETD-IMS-Pad\'{e}) for semilinear parabolic delay differential equations with smooth solutions. In our previous work [15], we proposed ETD-RK-Pad\'{e} scheme to compute high-order numerical solutions for nonlinear parabolic reaction-diffusion equation with constant time delay. However, the based ETD-RK numerical scheme in [15] is very complex and the corresponding calculation program is also very complicated. We propose in this paper ETD-MS-Pad\'{e} and ETD-IMS-Pad\'{e} schemes for the solution of semilinear parabolic equations with delay. We synergize the ETD-MS-Pad\'{e} with ETD-IMS-Pad\'{e} to construct efficient predictor-corrector scheme. This new predictor-corrector scheme will become an important tool for solving the numerical solutions of parabolic differential equations. Remarkably, we also conducted experiments in Table$10$ to compare the numerical results of the predictor-corrector scheme with the EERK scheme proposed in paper [42]. The predictor-corrector scheme demonstrated better convergence. The main idea is to employ an ETD-based Adams multistep extrapolation for the time integration of the corresponding equation. To overcome the well-known numerical instability associated with computing the exponential operator, we utilize the Pad\'{e} approach to approximate this exponential operator. This methodology leads to the development of the ETD-MS-Pad\'{e} and ETD-IMS-Pad\'{e} schemes, applicable even for arbitrary time orders. We validate the ETD-MS1,2,3,4-Pad\'{e} schemes and ETD-IMS2,3,4 schemes through numerical experiments.

math.NA

Data-driven construction of a generalized kinetic collision operator from molecular dynamics

We introduce a data-driven approach to learn a generalized kinetic collision operator directly from molecular dynamics. Unlike the conventional (e.g., Landau) models, the present operator takes an anisotropic form that accounts for a second energy transfer arising from the collective interactions between the pair of collision particles and the environment. Numerical results show that preserving the broadly overlooked anisotropic nature of the collision energy transfer is crucial for predicting the plasma kinetics with non-negligible correlations, where the Landau model shows limitations.

physics.comp-ph

OffsetOPT: Explicit Surface Reconstruction without Normals

Neural surface reconstruction has been dominated by implicit representations with marching cubes for explicit surface extraction. However, those methods typically require high-quality normals for accurate reconstruction. We propose OffsetOPT, a method that reconstructs explicit surfaces directly from 3D point clouds and eliminates the need for point normals. The approach comprises two stages: first, we train a neural network to predict surface triangles based on local point geometry, given uniformly distributed training point clouds. Next, we apply the frozen network to reconstruct surfaces from unseen point clouds by optimizing a per-point offset to maximize the accuracy of triangle predictions. Compared to state-of-the-art methods, OffsetOPT not only excels at reconstructing overall surfaces but also significantly preserves sharp surface features. We demonstrate its accuracy on popular benchmarks, including small-scale shapes and large-scale open surfaces.

cs.CV

Level-Set Parameters: Novel Representation for 3D Shape Analysis

3D shape analysis has been largely focused on traditional 3D representations of point clouds and meshes, but the discrete nature of these data makes the analysis susceptible to variations in input resolutions. Recent development of neural fields brings in level-set parameters from signed distance functions as a novel, continuous, and numerical representation of 3D shapes, where the shape surfaces are defined as zero-level-sets of those functions. This motivates us to extend shape analysis from the traditional 3D data to these novel parameter data. Since the level-set parameters are not Euclidean like point clouds, we establish correlations across different shapes by formulating them as a pseudo-normal distribution, and learn the distribution prior from the respective dataset. To further explore the level-set parameters with shape transformations, we propose to condition a subset of these parameters on rotations and translations, and generate them with a hypernetwork. This simplifies the pose-related shape analysis compared to using traditional data. We demonstrate the promise of the novel representations through applications in shape classification (arbitrary poses), retrieval, and 6D object pose estimation.

cs.CV

On the generalization ability of coarse-grained molecular dynamics models for non-equilibrium processes

One essential goal of constructing coarse-grained molecular dynamics (CGMD) models is to accurately predict non-equilibrium processes beyond the atomistic scale. While a CG model can be constructed by projecting the full dynamics onto a set of resolved variables, the dynamics of the CG variables can recover the full dynamics only when the conditional distribution of the unresolved variables is close to the one associated with the particular projection operator. In particular, the model's applicability to various non-equilibrium processes is generally unwarranted due to the inconsistency in the conditional distribution. Here, we present a data-driven approach for constructing CGMD models that retain certain generalization ability for non-equilibrium processes. Unlike the conventional CG models based on pre-selected CG variables (e.g., the center of mass), the present CG model seeks a set of auxiliary CG variables based on the time-lagged independent component analysis to minimize the entropy contribution of the unresolved variables. This ensures the distribution of the unresolved variables under a broad range of non-equilibrium conditions approaches the one under equilibrium. Numerical results of a polymer melt system demonstrate the significance of this broadly-overlooked metric for the model's generalization ability, and the effectiveness of the present CG model for predicting the complex viscoelastic responses under various non-equilibrium flows.

physics.comp-ph

AlphaDou: High-Performance End-to-End Doudizhu AI Integrating Bidding

Artificial intelligence for card games has long been a popular topic in AI research. In recent years, complex card games like Mahjong and Texas Hold'em have been solved, with corresponding AI programs reaching the level of human experts. However, the game of Doudizhu presents significant challenges due to its vast state/action space and unique characteristics involving reasoning about competition and cooperation, making the game extremely difficult to solve.The RL model Douzero, trained using the Deep Monte Carlo algorithm framework, has shown excellent performance in Doudizhu. However, there are differences between its simplified game environment and the actual Doudizhu environment, and its performance is still a considerable distance from that of human experts. This paper modifies the Deep Monte Carlo algorithm framework by using reinforcement learning to obtain a neural network that simultaneously estimates win rates and expectations. The action space is pruned using expectations, and strategies are generated based on win rates. The modified algorithm enables the AI to perform the full range of tasks in the Doudizhu game, including bidding and cardplay. The model was trained in a actual Doudizhu environment and achieved state-of-the-art performance among publicly available models. We hope that this new framework will provide valuable insights for AI development in other bidding-based games.

cs.AI

Consensus-based adaptive sampling and approximation for high-dimensional energy landscapes

We present a consensus-based framework that unifies phase space exploration with posterior-residual-based adaptive sampling for surrogate construction in high-dimensional energy landscapes. Unlike standard approximation tasks where sampling points can be freely queried, physical systems with complex energy landscapes such as molecular dynamics (MD) do not have direct access to arbitrary sampling regions due to the physical constraints and energy barriers; the surrogate construction further relies on the dynamical exploration of phase space, posing a significant numerical challenge. We formulate the problem as a minimax optimization that jointly adapts both the surrogate approximation and residual-enhanced sampling. The construction of free energy surfaces (FESs) for high-dimensional collective variables (CVs) of MD systems is used as a motivating example to illustrate the essential idea. Specifically, the maximization step establishes a stochastic interacting particle system to impose adaptive sampling through both exploitation of a Laplace approximation of the max-residual region and exploration of uncharted phase space via temperature control. The minimization step updates the FES surrogate with the new sample set. Numerical results demonstrate the effectiveness of the present approach for biomolecular systems with up to 30 CVs. While we focus on the FES construction, the developed framework is general for efficient surrogate construction for complex systems with high-dimensional energy landscapes.

physics.comp-ph

Data-driven learning of the generalized Langevin equation with state-dependent memory

We present a data-driven method to learn stochastic reduced models of complex systems that retain a state-dependent memory beyond the standard generalized Langevin equation (GLE) with a homogeneous kernel. The constructed model naturally encodes the heterogeneous energy dissipation by jointly learning a set of state features and the non-Markovian coupling among the features. Numerical results demonstrate the limitation of the standard GLE and the essential role of the broadly overlooked state-dependency nature in predicting molecule kinetics related to conformation relaxation and transition.

physics.comp-ph

Training with Product Digital Twins for AutoRetail Checkout

Automating the checkout process is important in smart retail, where users effortlessly pass products by hand through a camera, triggering automatic product detection, tracking, and counting. In this emerging area, due to the lack of annotated training data, we introduce a dataset comprised of product 3D models, which allows for fast, flexible, and large-scale training data generation through graphic engine rendering. Within this context, we discern an intriguing facet, because of the user "hands-on" approach, bias in user behavior leads to distinct patterns in the real checkout process. The existence of such patterns would compromise training effectiveness if training data fail to reflect the same. To address this user bias problem, we propose a training data optimization framework, i.e., training with digital twins (DtTrain). Specifically, we leverage the product 3D models and optimize their rendering viewpoint and illumination to generate "digital twins" that visually resemble representative user images. These digital twins, inherit product labels and, when augmented, form the Digital Twin training set (DT set). Because the digital twins individually mimic user bias, the resulting DT training set better reflects the characteristics of the target scenario and allows us to train more effective product detection and tracking models. In our experiment, we show that DT set outperforms training sets created by existing dataset synthesis methods in terms of counting accuracy. Moreover, by combining DT set with pseudo-labeled real checkout data, further improvement is observed. The code is available at https://github.com/yorkeyao/Automated-Retail-Checkout.

cs.CV

Construction of coarse-grained molecular dynamics with many-body non-Markovian memory

We introduce a machine-learning-based coarse-grained molecular dynamics (CGMD) model that faithfully retains the many-body nature of the inter-molecular dissipative interactions. Unlike common empirical CG models, the present model is constructed based on the Mori-Zwanzig formalism and naturally inherits the heterogeneous state-dependent memory term rather than matching the mean-field metrics such as the velocity auto-correlation function. Numerical results show that preserving the many-body nature of the memory term is crucial for predicting the collective transport and diffusion processes, where empirical forms generally show limitations.

physics.comp-ph

Large-scale Training Data Search for Object Re-identification

We consider a scenario where we have access to the target domain, but cannot afford on-the-fly training data annotation, and instead would like to construct an alternative training set from a large-scale data pool such that a competitive model can be obtained. We propose a search and pruning (SnP) solution to this training data search problem, tailored to object re-identification (re-ID), an application aiming to match the same object captured by different cameras. Specifically, the search stage identifies and merges clusters of source identities which exhibit similar distributions with the target domain. The second stage, subject to a budget, then selects identities and their images from the Stage I output, to control the size of the resulting training set for efficient training. The two steps provide us with training sets 80\% smaller than the source pool while achieving a similar or even higher re-ID accuracy. These training sets are also shown to be superior to a few existing search methods such as random sampling and greedy sampling under the same budget on training data size. If we release the budget, training sets resulting from the first stage alone allow even higher re-ID accuracy. We provide interesting discussions on the specificity of our method to the re-ID problem and particularly its role in bridging the re-ID domain gap. The code is available at https://github.com/yorkeyao/SnP.

cs.CV