arXiv ScienceSearch

arXiv subjects

Ruiqing He

Publications and source records attributed to Ruiqing He.

8 recordsLinked to original sources

A HIP-Compatible Accelerator Backend for Fourier-Bessel Particle-in-Cell Simulations on CPU/DCU Heterogeneous Clusters

FBPIC (Fourier-Bessel particle-in-cell) is a high-performance simulation code for relativistic plasma and accelerator physics. Its original accelerator backend relies on Numba CUDA, which limits its direct deployment on accelerators using the HIP (Heterogeneous-Compute Interface for Portability) programming environment, such as DCU (Deep Computing Unit) accelerators. In this work, we develop an accelerator backend compatible with HIP that enables FBPIC to run efficiently on DCU platforms while preserving its Python user interface and high level simulation workflow. For the evaluated LWFA (laser-wakefield acceleration) workloads, the proposed backend achieves 1.32-1.54x speedups over the original FBPIC implementation on an NVIDIA V100 GPU and enables efficient execution on the DCU platform. We also summarize the key lessons learned from porting FBPIC to the DCU platform. Multi-DCU experiments achieve a 1.88x strong-scaling speedup on four accelerators and a 2.72x increase in aggregate throughput at approximately 68\% weak-scaling efficiency, with communication analysis identifying inter-node communication and synchronization as the main scalability limitations. Beyond FBPIC, the proposed approach provides a practical reference for porting and optimizing other scientific computing applications developed with Python on heterogeneous accelerator platforms.

physics.comp-ph

A CPU+DCU Heterogeneous Parallel Framework for Post-Processing Reconstruction in Quantum Circuit Cutting

In the NISQ era, limited qubit resources make it difficult to execute large quantum circuits directly on real hardware. Quantum circuit cutting mitigates this limitation by decomposing a large circuit into smaller subcircuits, but it shifts substantial overhead to classical post-processing. As circuit size, complexity, and cut count increase, reconstruction becomes a major computational and storage bottleneck. This paper presents a CPU+DCU heterogeneous parallel framework for circuit-cutting post-processing reconstruction. Instead of constructing a dense $2^n$-dimensional probability vector or returning only high-probability states, the framework reconstructs the nonzero-probability states in the original output distribution from subcircuit measurement results. It combines heterogeneous CPU+DCU execution with a high/low-word integer representation for global basis-state indices beyond 64 bits and a three-level cooperative storage mechanism spanning device memory, host memory, and out-of-core storage. Experiments on the Songshan supercomputer show that the framework maintains high reconstruction fidelity while achieving up to $259\times$ speedup over an optimized serial baseline on linear-cluster states and up to $4\times$ speedup over a homogeneous CPU-parallel method on random circuits. The framework can also complete reconstruction tasks at the hundred-qubit scale. These results demonstrate that HPC-oriented heterogeneous reconstruction can effectively alleviate the classical post-processing bottleneck and improve reconstruction scalability.

quant-ph

Sparsified Kolmogorov-Arnold Networks for Interpretable Quantum State Tomography

Machine-learning approaches to quantum state tomography can achieve high reconstruction fidelity, but the physical structure used by the trained model often remains implicit. Here we ask whether a sparsified Kolmogorov-Arnold Network (KAN) can be used not only as a regressor, but also as an inspectable reconstruction rule whose internal organization can be checked against known Pauli structure. We study a controlled three-qubit GHZ-family benchmark in which all 63 non-identity Pauli expectation values are used to reconstruct three GHZ-subspace variables: the population imbalance $z$, the real off-diagonal component $c$, and the imaginary off-diagonal component $s$. Under finite-shot sampling and depolarizing noise, external ablation identifies the extended 12-channel GHZ-relevant Pauli set from the 63 measurements, with exact top-12 recovery across the tested shot counts and depolarizing-noise strengths. These support patterns remain stable across multi-seed random-initialization and noise-level analyses, and collapse under random-label controls. The dominant pruned input-hidden-output pathways organize Z-type population observables and X/Y off-diagonal observables in a pattern consistent with the analytic GHZ Pauli grouping, and sparse formula recovery recovers the canonical signed Pauli relations. The contribution of the KAN is therefore pathway-level structural interpretability within a neural reconstruction model, rather than superior sparse regression. Together with negative controls, these probes provide a consistency chain for auditing learned reconstruction rules against known physical structure.

quant-ph

Approximate Hamiltonian Simulation Algorithm for Efficient Fluid Quantum Simulations

This work aims to address the bottleneck issues of hardware resource limitation and decoherence error in the Hamiltonian simulation of quantum fluids, which are caused by the standard quantum Fourier transform and the evolution of momentum operators, resulting in excessively deep circuits and excessive two-qubit gates. We propose an approximate operator optimization scheme aimed at reducing the circuit depth in Hamiltonian evolution. The proposed scheme successfully reduces the depth of analog circuits from $O(n^2)$ to $O(nlogn)$ or even $O(n)$ by eliminating $O(n^2)$ redundant two-qubit entangling gates. In this work, the numerical experiments are implemented on a supercomputing-oriented quantum simulator, simulating two-dimensional unsteady divergent flow. Experimental results demonstrate that although the truncation of high-frequency qubit coupling terms introduces deterministic theoretical errors, scaling at $O(n)$ for AQFT and $O(n^2)$ for momentum truncation, the optimized simulations successfully preserve the inherent macroscopic temporal evolution characteristics of the fluid in a 10-qubit simulation, achieving high correlation coefficients of $r$=0.933, $r$=0.941, and $r$=0.977 for density, X-momentum, and Y-momentum distributions respectively. Furthermore, we also analyzed the relationship between the algorithm truncation error and the hardware cumulative noise when the qubit number is extended to a higher level. This study proves that rationally adjusting truncation thresholds can establish an equilibrium point, preventing the hardware cumulative error from rapidly approaching 100% at the 20-30 qubit scale, providing a feasible engineering pathway for simulating complex fluid systems on real quantum devices in the future.

quant-ph

Scalable Quantum Error Mitigation with Physically Informed Graph Neural Networks

Quantum error mitigation (QEM) provides a practical route for estimating reliable observables on noisy intermediate-scale quantum (NISQ) devices. Traditional QEM strategies, including zero-noise extrapolation (ZNE) and Clifford data regression (CDR), rely on noise scaling or global regression, and their performance is constrained by the exponential growth of the system degrees of freedom. We construct a graph-enhanced mitigation (GEM) framework, which incorporates physical information into the model representation. In this work, quantum circuits are encoded as attributed graphs. Hardware-level physical information is mapped to node and edge features: local noise parameters such as calibration parameters $T_1$, $T_2$, and readout errors are encoded at nodes, while coupling-related information such as two-qubit gate errors is encoded as edge features. Graph neural networks are used to model how errors propagate along the physical coupling structure and build up into non-local correlations. This allows the model to capture local interactions and part of the resulting non-local correlations across qubits. A dual-branch affine correction is applied to maintain consistency with physical constraints. Experiments on 10-qubit and 16-qubit random circuits executed on superconducting quantum processors show that GEM provides a level of accuracy comparable to CDR at small scales, while yielding lower mean absolute error and improved stability in zero-shot transfer to larger systems. Results of the traditional QEM strategy indicate that global regression methods remain effective in low-dimensional settings but become less reliable as system degrees of freedom grow. In contrast, GEM makes use of local physical structures to show better scalability and generalization, while preserving the overall error propagation patterns. This work provides a practical scalable approach to QEM for NISQ devices.

quant-ph

High-resolution borehole earthquake monitoring at San Andreas Fault Observatory at Depth, Parkfield, California

Downhole earthquake monitoring, without the complex effects from the near surface, can record more and better seismic data than monitoring on surface. The San Andreas Fault Observatory at Depth (SAFOD) is a borehole observatory equipped with different instruments inside to study the earthquake mechanism of the San Andreas fault at Parkfield, California. During April to May in 2005, Paulsson deployed an 80-level 3-component geophone array in the SAFOD main hole, and continuously recorded seismic data for about 13 days. We located 125 local earthquakes from the borehole earthquake monitoring data using a homogeneous velocity model and compared it with 35 earthquakes' locations from surface earthquake monitoring by the United State Geological Survey (USGS) during the same monitoring time. The borehole earthquake locating is assumably more accurate in the borehole's vicinity. We also compared the result with 1,074 earthquakes' locations from the surface earthquake monitoring in the last 9 years from 2015 to 2024. The hypocenters from our nearly 2 weeks' borehole earthquake monitoring form similar structures as that from the 9 years' surface earthquake monitoring by the USGS.

physics.geo-ph

Paraxial micro earthquake: a natural effective multi-purpose check shot for downhole earthquake monitoring

Downhole earthquake monitoring, without the effects from the overburden, can record better seismic data than monitoring on surface. However, in order to reasonably use the downhole vector seismic data, a constant challenge is how to accurately orient the downhole radial-component seismometers. A common practice is to use offset check shots on or near the surface. However, in areas with complex geologies, this routine may result in significant orientation errors. A ParAxial Micro Earthquake (PAME) is a micro earthquake at a close distance to the seismometers and near the extended path of the borehole's trajectory. It is rarely recorded during downhole earthquake monitoring unless designed for. If it is recorded, it can be a real treasure not only for P-wave and S-wave velocities' profiling, but for the downhole seismometers' orientation. As an example, during April to May in 2005, Paulsson installed an 80-level 3-component VSP (Vertical Seismic Profiling) array in the SAFOD (San Andreas Fault Observatory at Depth) main hole at Parkfield, California, and continuously recorded seismic data for about 13 days. Large charge offset check shots at 13 different locations near the surface were detonated in order to orient the downhole geophones; the orientation results were unsatisfactory but went unnoticed or unsolved. Besides this, a small charge "zero-offset" check shot was detonated near the wellhead in order to get the P-wave and S-wave velocity profiles, but only the P-wave velocity profiling was successful. Fortunately, we recorded a few PAMEs, through which we not only obtained better P-wave and S-wave velocity profiles, but satisfactorily oriented the downhole geophones.

physics.geo-ph

Indications and implications of a borehole seismic monitoring result of the San Andreas Fault foresee future earthquake prediction

The Parkfield M6 earthquake predicted from 1985 by the USGS to happen by 1993 happened 11 years later in 2004 instead. Till today, satisfactory answers to why this earthquake was mis-predicted have not been found. Seven months after the earthquake, we deployed a seismic array in the SAFOD main hole to monitor the San Andreas Fault. During a 13 days period, we recorded 220 earthquake events. By analyzing the projected hypocenters of 100 earthquake events, we found three main active sub-faults, namely SAF1, SAF2, and SAF3, with SAF1 being the normally regarded SAF. During the monitoring period, there was an earthquake migration trend that shows the move of SAF1 drags SAF2, which in turn drags SAF3, to slip and dip downwards. This could be reflected by the result of a geological trenching across the SAF. On SAF1, the smaller earthquakes tended to occur earlier in southeast and the larger earthquakes tended to occur later in northwest, which could indicate the existence of strong asperities. However, the new earthquakes on SAF2 were overall remarkably smaller, which could indicate that during the 2004 M6 earthquake the main underground rupture took place on SAF2, which could extend with a bend near Parkfield town to the Southwest Fracture Zone, and further to include the seemingly isolated 2004 M6 hypocenter. The possible prevailing of this triple active sub-faults model in the Parkfield area may account for 12, 24 or 36 (instead of a single 22 years as previously believed) recurring years of M6 earthquakes, depending on how many of these active sub-faults have basically absorbed the tectonic stress. In less than 0.1% of the earthquake's last intermission time, and at just one location, our borehole seismic monitoring has revealed so much insightful information that earthquake prediction may be feasible if the fault is monitored in boreholes at more locations and for longer periods of time.

physics.geo-ph