arXiv ScienceSearch

arXiv · 2108.04490

Quantum Reinforcement Learning: the Maze problem

Abstract

Quantum Machine Learning (QML) is a young but rapidly growing field where quantum information meets machine learning. Here, we will introduce a new QML model generalizing the classical concept of Reinforcement Learning to the quantum domain, i.e. Quantum Reinforcement Learning (QRL). In particular we apply this idea to the maze problem, where an agent has to learn the optimal set of actions in order to escape from a maze with the highest success probability. To perform the strategy optimization, we consider an hybrid protocol where QRL is combined with classical deep neural networks. In particular, we find that the agent learns the optimal strategy in both the classical and quantum regimes, and we also investigate its behaviour in a noisy environment. It turns out that the quantum speedup does robustly allow the agent to exploit useful actions also at very short time scales, with key roles played by the quantum coherence and the external noise. This new framework has the high potential to be applied to perform different tasks (e.g. high transmission/processing rates and quantum error correction) in the new-generation Noisy Intermediate-Scale Quantum (NISQ) devices whose topology engineering is starting to become a new and crucial control knob for practical applications in real-world problems. This work is dedicated to the memory of Peter Wittek.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nicola Dalla Pozza, Lorenzo Buffoni, Stefano Martina, Filippo Caruso. 2021-08-10. Quantum Reinforcement Learning: the Maze problem. https://doi.org/10.1007/s42484-022-00068-y

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Quantum Agreement Theorem

Intersubjective consistency asks when different observers of a physical system, reasoning from a shared theory, will agree on their probability assessments. Classical probability theory provides a benchmark in the form of the Agreement Theorem (Aumann, 1976). We study the quantum mechanics (QM) analog to the classical Agreement Theorem in a finite-dimensional tripartite setting in which two observers obtain information via local projective measurements. We represent their mutual awareness by common certainty, which is formally an infinite hierarchy of certainty operators concerning the probability of an event of interest. We first prove that if the observers' measurements commute with one another and with the property of interest, common certainty forces their probability assessments to coincide, thereby recovering a QM analog to the classical Agreement Theorem. We then construct a one-parameter family of qutrit-qubit-qubit models in which one observer's measurement does not commute with the property and common certainty of disagreement occurs -- a distinctively QM phenomenon. The mechanism is coherence between branches distinguished by one but not the other observer's measurement. When measurement outcomes are stored in a classical register, the resulting dephasing removes this coherence and restores agreement. Finally, we prove that QM does not permit the extreme 0-1 configuration in which Alice is certain of a property and is also certain that Bob is certain of its negation. These results aim to turn debate in QM about observer-dependent facts into precise mathematical conditions for when intersubjective agreement or disagreement in QM can be sustained.

quant-ph

Exact diffusion coefficients for quantum--classical hybrid walks

One-dimensional quantum--classical hybrid walks are introduced by combining quantum and classical steps under periodic and random protocols. Exact diffusion coefficients are derived for both protocols throughout the diffusive regime. The exact results show how diffusion depends on the average frequency and temporal arrangement of classical steps. As the average frequency approaches zero, both diffusion coefficients diverge inversely with the frequency. At matched average frequencies, the ratio of the diffusion coefficients for the random and periodic protocols approaches two. A finite-mode approximation is also developed and tested against the exact diffusion coefficient for the random protocol. The exact results for the one-dimensional model provide a foundation for future studies of more general quantum--classical hybrid walks.

quant-ph

Error Exponents of Probabilistic Quantum Resource Distillation

In this manuscript, we establish a unified framework for analyzing the conditional error exponents of probabilistic resource distillation under approximately resource-nongenerating instruments. For generic quantum resource theories satisfying suitable structural conditions, we derive general oneshot bounds on the conditional distillation error exponents. By relating probabilistic distillation to postselected composite quantum hypothesis testing, we obtain bounds of the conditional error exponents for coherence distillation under finite blocklength and asymptotic zero-rate scenarios, we furthermore obtain analytical characterizations of the conditional error exponents for entanglement and magic distillation under finite blocklength and asymptotic zero-rate scenarios. For several representative families of states in entanglement and magic resource theories, these characterizations reduce to explicit closed-form formulas. Comparing them with the corresponding deterministic distillation exponents, we identify regimes in which postselection yields a strict improvement in the exponential decay rate of the conditional error. Our results reveal an operational advantage of postselection in quantum resource distillation and establish postselected composite hypothesis testing as a general tool for characterizing probabilistic resource-processing tasks.

quant-ph