arXiv ScienceSearch

arXiv · 2302.10796

Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret

Abstract

While quantum reinforcement learning (RL) has attracted a surge of attention recently, its theoretical understanding is limited. In particular, it remains elusive how to design provably efficient quantum RL algorithms that can address the exploration-exploitation trade-off. To this end, we propose a novel UCRL-style algorithm that takes advantage of quantum computing for tabular Markov decision processes (MDPs) with $S$ states, $A$ actions, and horizon $H$, and establish an $\mathcal{O}(\mathrm{poly}(S, A, H, \log T))$ worst-case regret for it, where $T$ is the number of episodes. Furthermore, we extend our results to quantum RL with linear function approximation, which is capable of handling problems with large state spaces. Specifically, we develop a quantum algorithm based on value target regression (VTR) for linear mixture MDPs with $d$-dimensional linear representation and prove that it enjoys $\mathcal{O}(\mathrm{poly}(d, H, \log T))$ regret. Our algorithms are variants of UCRL/UCRL-VTR algorithms in classical RL, which also leverage a novel combination of lazy updating mechanisms and quantum estimation subroutines. This is the key to breaking the $Ω(\sqrt{T})$-regret barrier in classical RL. To the best of our knowledge, this is the first work studying the online exploration in quantum RL with provable logarithmic worst-case regret.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Han Zhong, Jiachen Hu, Yecheng Xue, Tongyang Li, Liwei Wang. 2024-06-13. Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret. https://arxiv.org/abs/2302.10796

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Designing a Hybrid Digital / Analog Quantum Physics Emulator as Open Hardware

Existing approaches to emulating quantum computing algorithms using classical electronic hardware are limited by exponential scaling limitations in space, such as circuit size, or time, such as runtime or bandwidth. We introduce a scheme for representing quantum information using analog signals that lessens the bandwidth limitation problem (in certain regimes) seen in existing approaches [1, 2] by taking full advantage of the ability of analog signals to encode information using RMS voltage as well as frequency and phase. We introduce the mathematical framework for this representation, which separates the information relevant for measurement in the computational basis from information that is not relevant to it. We introduce circuits that take advantage of this separation of concerns to achieve simplifications, for working with quantum information in this representation. We argue that it is comparatively very inexpensive (as low as ~$5.00 / qubit) to outmatch the computing capabilities of existing FPGA based emulators [3], though scaling beyond tens of qubits is still impractical due to constraints of analog hardware module precision. However, our approach opens the door to a new avenue by which classical emulators can hope to improve: by improving on analog electronic circuit performance.

quant-ph

Measuring a Quantum Measure Exceeding Unity

The history based formalism known as Quantum Measure Theory (QMT) generalizes the concept of probability-measure so as to incorporate quantum interference. The resulting quantum measure $μ$ is defined for arbitrary events (sets of histories), not just for observables at a fixed moment of time. Thanks to interference effects, $μ$ can exceed unity, exhibiting its non-classical nature in a particularly striking manner. Here, in an optical experiment, we illustrate an ancilla based filtering scheme that gives operational meaning to the quantum measure. For a specific photonic event $E$, we report a measured value of $μ(E)=1.172^{+0.013}_{-0.019}$, which within errors agrees with the theoretical value of $5/4$, while exceeding the maximum value permissible for a classical probability (namely $1$) by $13.32$ upper or $8.89$ lower percentile widths. The directly observed quantity is an ordinary detector probability $p_D\le 1$ (or, with laser light, an equivalent power ratio); the value $μ(E)>1$ is inferred via the calibrated relation $μ(E)=2p_D$ for our filter. If an unconventional theoretical concept is to play a role in meeting the foundational challenges of quantum theory, it seems important to bring it into contact with experiment as much as possible. Our experiment does this for the quantum measure.

quant-ph

Complexity Theory for Quantum Promise Problems

We begin by establishing structural results for several fundamental quantum complexity classes: p/mBQP, p/mQ(C)MA, $\text{p/mQSZK}_{\text{hv}}$, p/mQIP, p/mBQP/qpoly, p/mBQP/poly, and p/mPSPACE. This includes identifying complete problems, as well as proving containment and separation results among these classes. Here, p/mC denotes the corresponding quantum promise complexity class with pure (p) or mixed (m) quantum input states for any classical complexity class C. Surprisingly, our findings uncover relationships that diverge from their classical analogues -- specifically, we show unconditionally that p/mQIP$\neq$p/mPSPACE and p/mBQP/qpoly$\neq$p/mBQP/poly. This starkly contrasts the classical setting, where QIP$=$PSPACE and separations such as BQP/qpoly$\neq$BQP/poly are only known relative to oracles. More interestingly, these separation results further connected to the topic of for both quantum property testing and unitary synthesis. This new framework has numerous applications in quantum cryptography, particularly in the contexts of Microcrypt. We provide a better characterization of its primitives; for example, we show that OWSG and PRS can be broken by a p/mQCMA oracle, leading to a natural quantum analogue of Impagliazzo's five worlds by substituting the classical complexity classes in Pessiland, Heuristica, and Algorithmica with mBQP and mQCMA. Moreover, we establish the relativization barrier for proving the existence of EFI, noting that no such barrier currently exists within traditional complexity theory.

quant-ph