arXiv ScienceSearch

arXiv subjects

Yang Peng

Publications and source records attributed to Yang Peng.

At least 19 recordsLinked to original sources

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular $M$-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes $\alpha_t=c(t+1)^{-a}$ with $a\in(1/2,1)$, the leading last-iterate fluctuation is of order $\widetilde O\bigl(T^{-a/2}/\sqrt{1-\gamma}\bigr)$ and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order $m^{-1}$ in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.

stat.ML

Online Inference in Distributional Temporal-Difference Learning

We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory. For the Polyak--Ruppert averaged estimator, we prove that its root-$T$ error converges weakly to a centered Gaussian random element in Cram\'er space. We also prove that, conditionally on the observed trajectory, the root-$T$ difference between the bootstrap and original averages converges weakly to the same Gaussian limit. These results justify bootstrap inference for smooth statistical functionals, including variance, CVaR, expected shortfall, and expectiles. For nonsmooth statistical functionals, we develop a local asymptotic theory for the estimated return CDF over $T^{-1/2}$-neighborhoods of finitely many thresholds, together with its bootstrap analogue. This theory allows us to conduct inference for nonsmooth statistical functionals characterized by CDF equations, including return quantiles.

stat.ML

Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian motion. We next provide online inference methods. Based on random scaling, the inference procedure constructs an asymptotically pivotal statistic for inference by using the information along the whole QTD path. Meanwhile, the proposed statistic can be computed online without storing the entire trajectory of QTD iterates. This substantially reduces the memory requirement and enables efficient statistical inference in distributional reinforcement learning.

stat.ML

Programmable Heisenberg-limit sensor from a nonlinear quantum energy pump

We introduce a programmable Heisenberg-limited bosonic quantum sensor based on a nonlinear quantum energy pump, implemented with a Kerr-nonlinear resonator coupled to multiple high-Q microwave terminal resonators. For parameter estimation encoded in an arbitrary number-conserving Hamiltonian acting on the terminal modes, we analytically construct optimal sensing protocols that attain the maximal quantum Fisher information, including initial-state loading, probe preparation, and readout. For diagonal multiparameter signals, we further show that the full phase-sensing quantum Fisher information matrix can be obtained from correlations of locally measured physical terminal works, providing a signal-free calibration of the metrological resource. We numerically demonstrate the construction and its robustness using realistic circuit-QED parameters while including experimentally relevant imperfections.

quant-ph

Work Statistics of Autonomous Quantum Energy Pumps

We develop a terminal-resolved theory of work statistics for autonomous quantum energy pumps. By modeling the systems that supply and receive energy as explicit quantum terminals, the energy exchanged with each terminal is defined directly from its Hamiltonian change. When the terminals are modeled by ideal clocks, these full-space observables admit exact representations on the pump Hilbert space and recover the conventional phase-derivative currents of periodically and quasiperiodically driven systems. The framework resolves transported work from energy accumulated in the pump. It also incorporates arbitrary initial pump--terminal correlations and identifies when correlations can enhance directional energy transfer under uncertain driving phases. For periodic pumps, we derive finite-cycle work statistics in Floquet eigenstates, relate their fluctuations to Floquet quantum geometry, and show that the long-time terminal currents become mutually compatible. An exactly solvable two-terminal qubit exhibits noise matching, in which transport becomes sharp through cancellation of common-mode terminal fluctuations even though the individual terminal energies remain noisy. Finally, for physical terminals beyond the ideal-clock limit, we introduce a positive work-variance gap that quantifies the fluctuations missed by a pump-only description. A coherent-cavity benchmark shows systematic convergence toward the ideal-clock regime with increasing occupation while demonstrating that agreement of the mean current alone does not guarantee accurate work statistics.

quant-ph

Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewards under a given policy. To obtain a finite-dimensional representation of the return distribution, we consider the quantile fixed point $\eta_m$ induced by the quantile-projected distributional Bellman equation. Assuming access to a generative model, we construct an estimator $\eta_m^{(n)}$ based on an empirical Markov decision process. For a fixed number of quantiles $m$, we establish a non-asymptotic error bound for $\eta_m^{(n)}$ and $\eta_m$ under the supremum $W_\infty$ metric, showing that the estimation error scales as $\widetilde{O}(\sqrt{m/n})$ with respect to $m$ and $n$. This implies that the quantile-based distributional policy evaluation problem can be solved with sample efficiency, achieving the optimal parametric $\sqrt{n}$ convergence rate. We derive the asymptotic distribution of the quantile parameters $\sqrt{n}(\theta_m^{(n)}-\theta_m)$ and characterize the semiparametric efficiency bound, which is attained by our estimator. Beyond the fixed-dimensional setting, we investigate the asymptotic regime in which the number of quantiles diverges. We characterize the limit covariance structure and show that it matches the semiparametric efficiency bound of the nonparametric model for distributional policy evaluation, showing that quantile-based estimators remain asymptotically efficient in the infinite-dimensional limit. Finally, we establish a Berry--Esseen theorem for smooth functionals $\sqrt{n}(\eta_m^{(n)}(s)-\eta_m(s))f$, thereby providing a foundation for statistically valid inference on functionals of the quantile-projected return distribution.

stat.ML

Universal Neural Propagator: Learning Time Evolution in Many-Body Quantum Systems

Conventional approaches to simulating quantum many-body dynamics produce a single trajectory: if the Hamiltonian or the initial state is changed, the computation must be re-performed. Recent efforts toward foundation models have begun to address this limitation, yet existing methods transfer across either Hamiltonians or initial states, but not both. In this work, we introduce the Universal Neural Propagator (UNP), a single, unified model that learns the functional mapping from driving protocols to time-evolution propagators. Trained in an entirely self-supervised way, a single UNP model predicts dynamics across a function space of driving protocols and an exponentially large Hilbert space of initial states simultaneously. We benchmark on a two-dimensional driven Ising model and demonstrate the UNP's accuracy and transferability across product and entangled initial states, as well as for both in- and out-of-distribution driving protocols. The UNP remains accurate at system sizes beyond exact diagonalization, and can be efficiently fine-tuned across all initial states using observable data. By shifting the object of learning from quantum states to operators, this work opens a route toward transferable simulation of driven quantum matter.

quant-ph

Neural Operator Quantum State: A Foundation Model for Quantum Dynamics

Capturing the dynamics of quantum many-body systems under time-dependent driving protocols is a central challenge for numerical simulations. Existing methods such as tensor networks and time-dependent neural quantum states, however, must be re-run for every protocol. In this work, we introduce the Neural Operator Quantum State (NOQS) as a foundation model for quantum dynamics. Rather than solving the Schr\"odinger equation for individual trajectories, our approach aims to \emph{learn the solution operator} that maps entire driving protocols to time-evolved quantum states. Once trained, the NOQS predicts time evolution under unseen protocols in a single forward pass, requiring no additional optimization. We validate NOQS on the two-dimensional Ising model with time-dependent longitudinal and transverse fields, demonstrating accurate prediction not only for unseen in-distribution protocols, but also for qualitatively different, out-of-distribution functional forms of driving. Further, a single NOQS model can be transferred between different temporal resolutions, and can be efficiently fine-tuned with sparse experimental measurements to improve predictions across all observables at negligible cost. Our work introduces a new paradigm for quantum dynamics simulation and provides a practical computational-experimental interface for driven quantum systems.

quant-ph

Superextensive charging speeds in a correlated quantum charger

We define a quantum charger as an interacting quantum system that transfers energy between two drives. The key figure of merit characterizing a charger is its charging power. Remarkably, the presence of long-range interactions within the charger can induce a collective steady-state charging mode that depends superlinearly on the size of the charger, exceeding the performance of noninteracting, parallel units. Using the driven Lipkin-Meshkov-Glick model and power-law interacting spin chains, we show that this effect persists up to a critical system size set by the breakdown of the high-frequency regime. We discuss optimal work output as well as experimentally accessible initial states. The superlinear charging effect can be probed in trapped-ion experiments, and positions interacting Floquet systems as promising platforms for enhanced energy conversion.

cond-mat.stat-mech

Thermodynamic-Limit Evidence for Chiral Superconductivity Induced by Doping Chiral Topological Phases

The emergence of superconductivity from doping strongly correlated chiral topological phases in purely repulsive two-dimensional fermionic systems is a problem of broad and fundamental interest. However, existing numerical evidence has been limited to finite-size studies, and direct thermodynamic-limit evidence for superconducting long-range order has remained lacking. Here we provide such evidence for chiral superconductivity in the triangular Hofstadter-Hubbard model by advancing a simplex tensor-network approach that simultaneously captures superconducting long-range order and chiral topological order in the presence of intrinsic charge fluctuations, a capability that has remained challenging for previous two-dimensional approaches. We show that a broad intermediate-$U$ chiral spin liquid is separated from the weak-$U$ Chern insulator by a Mott transition, together forming undoped parent chiral topological states. Upon hole doping, we identify a uniform chiral superconducting state in the infinite system, characterized by a finite complex pairing order parameter. The pairing field exhibits an almost universal phase winding over a broad interaction-doping regime, with a distinct pocket of opposite winding near the Mott criticality. In addition, the entanglement spectrum retains the chiral structure of the parent topological phases while developing additional low-energy branches upon doping. These results establish that chiral superconductivity emerges robustly from doped chiral topological phases.

cond-mat.str-el

Adiabatic pumping of topological corner states by coherent tunneling in a 2D SSH model

The active manipulation of topologically protected states represents a pivotal frontier for quantum technologies, offering a unique confluence of topological robustness and precise quantum control. We propose an adiabatic pumping scheme for the long-range transfer of topological corner states in a two-dimensional Su-Schrieffer-Heeger model. The protocol utilizes a modular lattice architecture composed of four topologically distinct subblocks, enabling the modulation of a topological dark state by precise tuning of lattice couplings. This approach is based on coherent tunneling by adiabatic passage among topological corner and interface states. We establish a multi-level model for the adiabatic pumping that provides an accurate description of the underlying mechanism. In comparison with a sequential two-stage Thouless pumping, our protocol offers superior performance in both transfer fidelity and efficiency.

quant-ph

Accelerated Distributional Temporal Difference Learning with Linear Function Approximation

In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The purpose of distributional TD learning is to estimate the return distribution of a discounted Markov decision process for a given policy. Previous works on statistical analysis of distributional TD learning focus mainly on the tabular case. We first consider the linear function approximation setting and conduct a fine-grained analysis of the linear-categorical Bellman equation. Building on this analysis, we further incorporate variance reduction techniques in our new algorithms to establish tight sample complexity bounds independent of the support size $K$ when $K$ is large. Our theoretical results imply that, when employing distributional TD learning with linear function approximation, learning the full distribution of the return function from streaming data is no more difficult than learning its expectation. This work provide new insights into the statistical efficiency of distributional reinforcement learning algorithms.

stat.ML

Non-abelian Geometric Quantum Energy Pump

We introduce a non-abelian geometric quantum energy pump realized by a transitionless geometric quantum drive--a time-dependent Hamiltonian supplemented by a counterdiabatic term generated by a prescribed trajectory on a smooth control manifold--that coherently transports states within a degenerate subspace. When the coordinates of the trajectory are independently addressable by external drives, the net energy transferred between drives is set by the non-abelian Berry-curvature tensor. The trajectory-averaged pumping power is separately controlled by the initial state and by the Hamiltonian topology through the Euler class. We outline an implementation with artificial atoms, which are realizable on various platforms including trapped atoms/ions, superconducting circuits, and semiconductor quantum dots. The resulting energy pump can serve as a quantum transducer or charger, and as a metrological tool for measuring phase coherences in quantum states.

quant-ph

Fourier Neural Operators for Time-Periodic Quantum Systems: Learning Floquet Hamiltonians, Observable Dynamics, and Operator Growth

Time-periodic quantum systems exhibit a rich variety of far-from-equilibrium phenomena and serve as ideal platforms for quantum engineering and control. However, simulating their dynamics with conventional numerical methods remains challenging due to the exponential growth of Hilbert space dimension and rapid spreading of entanglement. In this work, we introduce Fourier neural operators (FNOs) as an efficient, accurate, and scalable framework for nonequilibrium quantum dynamics. Parameterized in Fourier space, FNO naturally captures temporal correlations and remains minimally dependent on discretization of time. We demonstrate the versatility of FNO through three complementary learning paradigms: reconstructing effective Floquet Hamiltonians, predicting expectation values of local observables, and learning quantum information spreading. For each learning task, FNO achieves remarkable accuracy, while attaining a significant speedup, compared to exact numerical methods. Moreover, FNO possesses capabilities beyond that of conventional methods, such as predicting all local observables from a subset of measurements without information about the Hamiltonian, as well as extrapolating beyond the time window provided by training data, enabling access to observables and operator-spreading dynamics that might be beyond the coherence time. By employing a spatially local basis, we argue that the computational cost of FNOs scales only polynomially with the system size. Our results establish FNO as a versatile and scalable computation framework that integrates numerical simulations and experimental data seamlessly, with direct implications for extracting meaningful physics from measurements by near-term quantum computers.

quant-ph

Instability of Nagaoka State and Quantum Phase Transition via Kinetic Frustration Control

We investigate the Nagaoka-Thouless (NT) ferromagnetic instability in the strongly interacting $t$-$t'$ Hubbard model by continuously breaking particle-hole symmetry on a tunable square-triangular lattice geometry. We use an analytic approach to show that the fully spin-polarized state becomes unstable to a metastable spin-polaron when the kinetic frustration $t'/t$ exceeds a critical, dimension-dependent value. Large-scale density matrix renormalization group simulations reveal a quantum phase transition from the NT ferromagnet to a spiral spin-density wave, which evolves continuously into the Haerter-Shastry antiferromagnet in the large-frustration limit. Remarkably, this transition remains robust at low but finite hole density, making it accessible in cold-atom and moir\'e Hubbard platforms under strong interactions. A variational analysis further captures the instability mechanism at finite density via frustration-induced magnon band deformation.

cond-mat.str-el

Matrix Moment and Concentration Inequalities for Martingales and Ergodic Markov Chains with Applications in Statistical Learning

In this paper, we study moment and concentration inequalities for the spectral norm of sums of dependent random matrices. We establish novel Rosenthal-Burkholder inequalities for the spectral norm of discrete-time matrix local martingales, Burkholder-Davis-Gundy inequality for the spectral norm of continuous matrix local martingales, as well as matrix Rosenthal, Hoeffding, and Bernstein inequalities under the spectral norm for ergodic Markov chains. Compared with previous work on matrix concentration inequalities for Markov chains, which assume a non-zero absolute $L^2$-spectral gap or the stronger $\psi$-mixing condition, our results assume geometric ergodicity, a condition commonly used in statistical applications. Furthermore, our results have leading terms that match the Markov chain central limit theorem, rather than relying on suboptimal variance proxies. We also give dimension-free versions of the inequalities, which are independent of the ambient dimension $d$ and relies on the effective rank instead. This enables the generalization of our results to linear operators in infinite-dimensional Hilbert spaces. Our results have extensive applications in statistics and machine learning; in particular, we obtain improved bounds in covariance estimation and principal component analysis on Markovian data.

math.PR

The Observations of Magnetic Reconnection During the Interaction Process of Two Active Region Filaments

We investigate the interaction between two filaments (F1 and F2) and their subsequent magnetic reconnection in active region (AR) NOAA 13296 and AR NOAA 13293 on May 9, 2023, utilizing high spatial and temporal resolution and multi-wavelength observational data from the Solar Dynamics Observatory, the New Vacuum Solar Telescope, and the Chinese H{\alpha} Solar Explorer. The movement of F1 from the southeast toward the northwest, driven by the motion of the positive magnetic polarity (P1), leads to a collision and reconnection with F2. This reconnection exchanges their footpoints, resulting in the formation of two new filaments (F3 and F4) consistent with "slingshot" type filament interaction. During the interaction, the current sheet moving due to the motion of F1 and the reconnection outflows moving along F3 and F4 were both observed. The current sheet is rarely observed in the slingshot type filament interaction, measuring approximately 2.17 Mm in length and 0.84 Mm in width. After the interaction, the F1 disappears whereas a portion of F2 remains, indicating that the interaction involves partial slingshot reconnection, due to the unequal magnetic flux between the filaments. The residual part of F2 will undergo another magnetic reconnection in the same interaction region with the magnetic loops connecting polarities N1 and P1. The material generated by the reconnection is continuously injected into F4, leading to its final morphology. The findings enhance our understanding of slingshot-type filament interactions, indicating that partial slingshot reconnections between filaments may be more common than full slingshot events.

astro-ph.SR

Quantum anomalous Hall effects and emergent $\rm{SU}(2)$ Hall ferromagnets at fractional filling of helical trilayer graphene

Helical trilayer graphene realizes a versatile moir\'e system for exploring correlated topological states emerging from high Chern bands. Motivated by recent experimental observations of anomalous Hall effects at fractional fillings of magic-angle helical trilayers, we focus on the higher Chern number $|C_{band}|=2$ band and explore gapped many-body Hall states beyond the conventional Landau level paradigm. Through extensive exact diagonalization, we predict novel phases unattainable in a single $|C_{band}|=1$ band. At filling $\nu=2/3$ and $\nu=1/3$, a $\sqrt{3}\times \sqrt{3}$ charge-ordered quantum Hall crystal and a Halperin fractional Chern insulator with Hall conductance $|\sigma_{H}|=2e^2/3h$ are predicted respectively, indicating strong particle-hole asymmetry of the system. At half-filling $\nu=1/2$, an extensively degenerate pseudospin Hall ferromagnet featuring emergent $\rm{SU}(2)$ symmetry is found without the band being flat. Inspired by striking robustness of the ferromagnetic degeneracy, we develop a method to unveil and quantify the emergent symmetry via pseudospin operator construction in the presence of band dispersion and Coulomb interaction, and demonstrate persistence of the $\rm{SU}(2)$ quantum numbers even far away from the chiral limit. Incorporating spin-valley degrees of freedom, we identify an optimal filling regime $\nu_{\rm{total}}=3+\nu$ for realizing the above states. Notably, inter-flavor interactions renormalize the bandwidth and stabilize all the gapped phases even in realistic sublattice corrugation parameter regimes.

cond-mat.str-el