arXiv ScienceSearch

arXiv subjects

Zuoqiang Shi

Publications and source records attributed to Zuoqiang Shi.

At least 19 recordsLinked to original sources

Weighted Laplacian Flow: A Deterministic Particle Flow with Provable Convergence

Sampling from a target probability density is a fundamental task in statistics, machine learning, and scientific computing. We introduce weighted Laplacian flow, a deterministic particle-flow method that transports samples from a tractable initial density to a target density known up to normalization. The method evolves the logarithmic density ratio between the target and the transported distribution and constructs the particle velocity by solving a weighted Poisson equation associated with the target density. This design avoids the need to choose a kernel and enables direct control of particle weights along the flow. We establish the global well-posedness of the proposed PDE system and prove that the transported density converges to the target density in both $L^\infty$ distance and Kullback-Leibler divergence. Under a sublinear forcing condition, the method achieves exact convergence in finite time. Numerical experiments on multimodal, heavy-tailed, and ten-dimensional targets demonstrate that weighted Laplacian flow can perform long-range mass transport, overcome energy barriers.

math.NA

A High-Order Surface Finite Element Method Based on Intersections with Background Tetrahedral Meshes

This paper develops a high-order surface finite element method for elliptic equations posed on a smooth closed surface implicitly defined as the zero level set of a function in three dimensions. In contrast to classical surface finite element methods that start from a prescribed triangulation of the surface, the proposed method constructs the discrete surface space from the intersections between the exact surface and an ambient tetrahedral mesh. More precisely, an active tetrahedral shell is generated around the implicit surface, and each cut tetrahedron contributes either a triangular or a quadrilateral surface patch according to its intersection pattern with the zero level set. The finite element space is first defined on the resulting piecewise planar cut surface and is then lifted to the exact surface by local level-set-based parameterizations. The resulting method combines features of surface FEM and unfitted/trace methods. Like surface FEM, it produces a conforming finite element space on the exact surface after lifting; however, the geometry and finite element space are induced by intersections with background tetrahedra like unfitted/trace methods. We prove that the local lifting maps agree pointwise across common faces and assemble into a global homeomorphism. We also derive explicit formulas for tangent vectors, Gram matrices, mass and stiffness integrals, and prove stability and high-order derivative estimates for the lifting maps. The latter estimates are formulated in broken Sobolev norms and lead to a Cea-type energy-norm error bound for the proposed high-order lifted surface finite element method.

math.NA

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling. However, adapting AR checkpoints across corruption kernels remains challenging because existing DLMs use different objectives and prediction parameterizations. We establish connections among SEDD, MDLM/GIDD, M2S, and Neural CTMC by expressing their conditional losses as a single generalized Kullback--Leibler objective over model reverse rates. We further derive conversions from clean-token predictions to concrete-score, posterior-mean, and exit-rate/jump parameterizations, yielding a shared \(x_0\) interface that supports switching between mask and uniform kernels. Building on these connections, we propose \ours{}, a simple continual pre-training approach for directly adapting pretrained GPT2 checkpoints to uniform-noise diffusion. Through systematic evaluation of 124M- and 355M-parameter models, we show that \ours{} steadily improves the trade-off between generative perplexity (GenPPL) and unigram entropy as the sampling budget increases from 16 to 256 steps. At 256 steps, \ours{}-S and \ours{}-M achieve GenPPL/entropy pairs of \(97.783/5.2626\) and \(71.516/5.6669\), respectively; no evaluated model at the same scale simultaneously outperforms \ours{} on both metrics. At both scales, \ours{} also achieves the highest WinoGrande, SIQA, and BBH accuracy among the compared diffusion models.

cs.LG

Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy

Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees nonnegative reverse jump rates, it does not ensure Bayes realizability: ratios at a noisy state need not be jointly induced by any clean-token posterior under the forward kernel. The score-entropy loss has the correct population optimum but does not enforce this constraint away from it. In a trained pure-uniform SEDD checkpoint, roughly one quarter of complete score vectors violate the coordinate box, while more than half lie inside it yet remain materially incompatible with any valid posterior. Such violations can produce negative pre-normalization weights in finite-step sampling. Projecting raw scores onto the bridge polytope removes all observed negative weights and improves external generative PPL from $203.6$ to $175.1$ without changing the sampler. We introduce \emph{mean-to-score} (M2S), which predicts a clean-token posterior mean and converts it to the score through an exact kernel-dependent linear map. The construction applies to any known coordinate-wise continuous-time Markov chain (CTMC) satisfying a mild support condition. For uniform corruption, it maps the probability simplex onto the bridge polytope; for absorbing-mask corruption, the resulting objective recovers MD4 exactly. In a controlled 28.4M-parameter CIFAR-10 comparison, M2S lowers test BPD from $3.173$ to $3.129$ and FID-50k from $\CifarSEDDFID$ to $\CifarMtwoSFID$. A 170M-parameter M2S model trained on about 262B OpenWebText token slots outperforms the evaluated pure-uniform SEDD, GIDD, and Neural CTMC checkpoints at every tested sampling budget, reaching generative PPL $143.3$ at 128 steps versus $183.6$ for the strongest pure-uniform baseline.

cs.LG

OLEDLM: A Unified Language Model for OLED Molecular Design

The development of organic light-emitting diode (OLED) materials faces the compounded challenges of an astronomically large chemical space, stringent quantum-chemical constraints, and a scarcity of labeled data. Although the question of OLED generation is important, few models have been trained effectively for this specific domain. We propose an inverse molecular design framework based on causal language models: given target optoelectronic properties (e.g., excitation energy, oscillator strength), our model directly generates OLED SMILES sequences satisfying the specified constraints. We employ a multi-stage strategy: first, we establish a foundational chemical language model using a LLaMA-style transformer architecture. To the best of our knowledge, this represents the first successful adaptation of LLMs specifically for the OLED domain, bridging the gap between generic molecular generation and the stringent structural requirements of optoelectronic materials. Second, we fine-tune property predictors based on a BERT model pre-trained on our large-scale OLED dataset. Then, we perform Reinforcement Learning on our fine-tuned model, leveraging our property predictor, for better SMILES generation. Finally, through DFT verification, we demonstrate that our framework can efficiently navigate the OLED chemical space, generating novel candidates with high structural validity and optimized optoelectronic properties.

cs.LG

Non-line-of-sight imaging with arbitrary relay surface geometries via 3D Gaussian Transient Rendering

Imaging objects hidden outside the direct line of sight expands the effective field of view and is critical for applications such as autonomous driving and robotic perception. Despite impressive progress in time-of-flight (ToF)-based non-line-of-sight (NLOS) imaging, real-world deployment remains challenging because practical measurements are often collected over spatially limited, arbitrarily shaped relay regions-conditions that violate the planar-wall and dense-sampling assumptions made by most existing methods. To address these limitations, we propose a LOS-guided NLOS imaging pipeline that imposes no geometric assumptions on the relay surface and naturally supports both confocal and non-confocal configurations. Our method represents the hidden scene using 3D Gaussian primitives and couples them with an efficient, differentiable transient rendering model, enabling end-to-end optimization directly from measured transients. We validate our approach on real-world measurements from both a public dataset and a custom-built capture system. Across settings, our method achieves state-of-the-art reconstruction fidelity under spatially limited, sparsely sampled conditions, and significantly outperforms existing methods on complex, arbitrary relay surface geometries.

physics.optics

ND-TNN: Tensor-Neural-Network Approximation for High-Dimensional Nonlocal Diffusion Models

We study a numerical method, built on the tensor neural network (TNN) architecture introduced in \cite{wang2022tensor}, for solving nonlocal diffusion models in high-dimensional spaces. The tensor-product structure of the TNN ansatz, combined with the separability of the Gaussian kernel, reduces the high-dimensional integrals in the nonlocal energy to products of low-dimensional integrals, which are evaluated by Gauss--Legendre quadrature; nonseparable source and boundary data are handled by a TNN-based preconditioning step. For the Dirichlet boundary condition, we establish the asymptotically compatible $L^2$ error estimate \[ \|u_{\mathrm{loc}}-u_{\delta,p}\|_{L^2(\Omega)} \le C\!\left(\frac{\varepsilon_f}{\sqrt\delta} +\frac{\varepsilon_g}{\delta} +\frac{\varepsilon_u}{\sqrt\delta} +\eta_{\mathrm{opt}}\right) +C\sqrt\delta, \] where $\varepsilon_f$, $\varepsilon_g$ and $\varepsilon_u$ are the data and trial-class approximation errors and $\eta_{\mathrm{opt}}$ is the optimization residual. For the Neumann boundary condition, the $L^2$ estimate is improved to $O(\varepsilon_f+\varepsilon_g/\sqrt\delta+\varepsilon_u +\eta_{\mathrm{opt}}+\delta)$, and an $H^1$ gradient estimate is further obtained through a smoothing post-processing step. Numerical experiments on tensor-product domains up to $d=20$ support the theoretical results, and additional tests on two- and three-dimensional $L$-shaped domains demonstrate the practical robustness of the method beyond the smooth-domain setting covered by the analysis.

math.NA

A Nonlocal $p$-Laplacian Interface Model with Sharp Interface

We propose an energy-based nonlocal $p$-Laplacian interface problem. Neumann interface conditions are naturally formulated via the energy, while Dirichlet conditions are enforced through a penalty term. A key feature is that the model retains a sharp interface, which facilitates extension to other interface problems; we illustrate this by developing a nonlocal approximation for the $p$-Laplacian interface problem with membrane conditions. By establishing $\Gamma$-convergence and compactness, we prove that as the nonlocal horizon vanishes, minimizers of the nonlocal functionals converge to those of the local counterparts. Numerical experiments using an efficient finite element method confirm the convergence.

math.AP

Constant-Target Energy Matching: A Unified Framework for Continuous and Discrete Density Estimation

Density estimation is a central primitive in probabilistic modeling, yet continuous, discrete, and mixed-variable domains are often treated by separate objectives, limiting the ability to exploit a common statistical structure across data types. Continuous score-based methods rely on log-density gradients, while discrete extensions typically use concrete score whose unbounded targets become unstable near low-probability states. We introduce Constant-Target Energy Matching (CTEM), a unified energy-based framework for density estimation on general state spaces. CTEM replaces ordinary density-ratio regression with a bounded energy-difference transform and derives from it a sample-only training objective with the constant target 1. The learned scalar potential recovers log p without partition-function estimation or explicit unbounded ratio regression. Across continuous, discrete, and mixed-variable benchmarks, CTEM substantially improves density estimation over competitive baselines and yields higher-quality samples under standard sampling procedures.

cs.AI

Neural Continuous-Time Markov Chain: Discrete Diffusion via Decoupled Jump Timing and Direction

Discrete diffusion models based on continuous-time Markov chains (CTMCs) have shown strong performance on language and discrete data generation, yet existing approaches typically parameterize the reverse rate matrix monolithically -- through proxies such as concrete scores (SEDD) or clean-data predictions (MDLM, GIDD) -- rather than aligning the parameterization with the intrinsic CTMC decomposition into jump timing and jump direction. We propose \textbf{Neural CTMC}, which exploits the underlying Poisson structure of CTMC dynamics by separately parameterizing the reverse process through an \emph{exit rate} (when to jump) and a \emph{jump distribution} (where to jump) via two dedicated network heads. We show that the evidence lower bound (ELBO) reduces to a path-space KL divergence between the true and learned reverse processes that factorizes into a Poisson KL for timing and a categorical KL for direction, and admits a tractable, gradient-equivalent and consistent loss. Experimentally, scored by Gemma2-9B, our pure-uniform Neural CTMC achieves $16.36$ generative perplexity on TinyStories (vs.\ GIDD $37.60$ and MDLM $42.66$). On OpenWebText, it attains the best perplexity at the same training-token budget across 16--128 sampling steps among the methods we compare (e.g., at 128 steps: Neural CTMC $183.6$ vs.\ MDLM $210.5$ and GIDD $249.8$). To facilitate reproducibility, we release our pretrained weights at https://huggingface.co/Jiangxy1117/Neural-CTMC.

cs.LG

BlinDNO: A Distributional Neural Operator for Dynamical System Reconstruction from Time-Label-Free data

We study an inverse problem for stochastic and quantum dynamical systems in a time-label-free setting, where only unordered density snapshots sampled at unknown times drawn from an observation-time distribution are available. These observations induce a distribution over state densities, from which we seek to recover the parameters of the underlying evolution operator. We formulate this as learning a distribution-to-function neural operator and propose BlinDNO, a permutation-invariant architecture that integrates a multiscale U-Net encoder with an attention-based mixer. Numerical experiments on a wide range of stochastic and quantum systems, including a 3D protein-folding mechanism reconstruction problem in a cryo-EM setting, demonstrate that BlinDNO reliably recovers governing parameters and consistently outperforms existing neural inverse operator baselines.

cs.LG

An Efficient Conditional Score-based Filter for High Dimensional Nonlinear Filtering Problems

In many engineering and applied science domains, high-dimensional nonlinear filtering is still a challenging problem. Recent advances in score-based diffusion models offer a promising alternative for posterior sampling but require repeated retraining to track evolving priors, which is impractical in high dimensions. In this work, we propose the Conditional Score-based Filter (CSF), a novel algorithm that leverages a set-transformer encoder and a conditional diffusion model to achieve efficient and accurate posterior sampling without retraining. By decoupling prior modeling and posterior sampling into offline and online stages, CSF enables scalable score-based filtering across diverse nonlinear systems. Extensive experiments on benchmark problems show that CSF achieves superior accuracy, robustness, and efficiency across diverse nonlinear filtering scenarios.

cs.LG

Kernel Variational Inference Flow for Nonlinear Filtering Problem

We present a novel particle flow for sampling called kernel variational inference flow (KVIF). KVIF do not require the explicit formula of the target distribution which is usually unknown in filtering problem. Therefore, it can be applied to construct filters with higher accuracy in the update stage. Such an improvement has theoretical assurance. Some numerical experiments for comparison with other classical filters are also demonstrated.

math.OC

Ultrasound Tomography of Musculoskeletal Tissues with Generative Neural Physics

Ultrasound Tomography (UT) is a radiation-free, high-resolution modality, but remains limited for musculoskeletal imaging due to the high computational cost and instability of full-waveform inversion in strongly scattering media. We propose a generative neural physics framework that couples generative networks with physics-informed neural simulation for fast, high-fidelity 3D UT. By learning a compact surrogate of ultrasonic wave propagation from a limited set of cross-modality images, our method merges the accuracy of wave modeling with the efficiency and stability of deep learning. This enables accurate quantitative imaging of in vivo musculoskeletal tissues, producing spatial maps of acoustic properties beyond reflection-mode images. On synthetic and in vivo data of breasts, arms, and legs, we reconstruct 3D maps of tissue parameters in under ten minutes, with sensitivity to acoustic variations in musculoskeletal tissues and resolution comparable to MRI. By overcoming computational bottlenecks in strongly scattering regimes, this approach demonstrates the feasibility of quantitative UT for musculoskeletal imaging and advances its development toward future routine clinical use.

cs.CV

Fast and Memory-efficient Non-line-of-sight Imaging with Quasi-Fresnel Transform

Non-line-of-sight (NLOS) imaging seeks to reconstruct hidden objects by analyzing reflections from intermediary surfaces. Existing methods typically model both the measurement data and the hidden scene in three dimensions, overlooking the inherently two-dimensional nature of most hidden objects. This oversight leads to high computational costs and substantial memory consumption, limiting practical applications and making real-time, high-resolution NLOS imaging on lightweight devices challenging. In this paper, we introduce a novel approach that represents the hidden scene using two-dimensional functions and employs a Quasi-Fresnel transform to establish a direct inversion formula between the measurement data and the hidden scene. This transformation leverages the two-dimensional characteristics of the problem to significantly reduce computational complexity and memory requirements. Our algorithm efficiently performs fast transformations between these two-dimensional aggregated data, enabling rapid reconstruction of hidden objects with minimal memory usage. Compared to existing methods, our approach reduces runtime and memory demands by several orders of magnitude while maintaining imaging quality. The substantial reduction in memory usage not only enhances computational efficiency but also enables NLOS imaging on lightweight devices such as mobile and embedded systems. We anticipate that this method will facilitate real-time, high-resolution NLOS imaging and broaden its applicability across a wider range of platforms.

cs.CV

An FDM-sFEM scheme on time-space manifolds and its superconvergence analysis

We study superconvergent discretization of the Laplace-Beltrami operator on time-space product manifolds with Neumann temporal boundary values, which arise in the context of dynamic optimal transport on general surfaces. We propose a coupled scheme that combines finite difference methods in time with surface finite element methods in space. By establishing a new summation by parts formula and proving the supercloseness of the semi-discrete solution, we derive superconvergence results for the recovered gradient via post-processing techniques. In addition, our geometric error analysis is implemented within a novel framework based on the approximation of the Riemannian metric. Several numerical examples are provided to validate and illustrate the theoretical results.

math.NA

OpenBreastUS: Benchmarking Neural Operators for Wave Imaging Using Breast Ultrasound Computed Tomography

Accurate and efficient simulation of wave equations is crucial in computational wave imaging applications, such as ultrasound computed tomography (USCT), which reconstructs tissue material properties from observed scattered waves. Traditional numerical solvers for wave equations are computationally intensive and often unstable, limiting their practical applications for quasi-real-time image reconstruction. Neural operators offer an innovative approach by accelerating PDE solving using neural networks; however, their effectiveness in realistic imaging is limited because existing datasets oversimplify real-world complexity. In this paper, we present OpenBreastUS, a large-scale wave equation dataset designed to bridge the gap between theoretical equations and practical imaging applications. OpenBreastUS includes 8,000 anatomically realistic human breast phantoms and over 16 million frequency-domain wave simulations using real USCT configurations. It enables a comprehensive benchmarking of popular neural operators for both forward simulation and inverse imaging tasks, allowing analysis of their performance, scalability, and generalization capabilities. By offering a realistic and extensive dataset, OpenBreastUS not only serves as a platform for developing innovative neural PDE solvers but also facilitates their deployment in real-world medical imaging problems. For the first time, we demonstrate efficient in vivo imaging of the human breast using neural operator solvers.

cs.CV

When surface evolution meets Fokker-Planck equation: a novel tangential velocity model for uniform parametrization

A common issue in simulating geometric evolution of surfaces is unexpected clustering of points that may cause numerical instability. We propose a novel artificial tangential velocity method for this matter. The artificial tangential velocity is generated from a surface density field governed by a Fokker-Planck equation to guide the point distribution. A target distribution matching algorithm is developed leveraging the surface Kullback-Leibler divergence of density functions. The numerical method is formulated within a fully meshless framework using the moving least squares approximation, thereby eliminating the need for mesh generation and allowing flexible treatment of unstructured point cloud data. Extensive numerical experiments are conducted to demonstrate the robustness, accuracy, and effectiveness of the proposed approach across a variety of surface evolution problems, including the mean curvature flow.

math.NA