arXiv ScienceSearch

arXiv · 2604.13096

Complexity scaling and optimal policy degeneracy in quantum reinforcement learning via analytically solvable unitary-control-then-measure models

Abstract

We propose and analyse a class of analytically solvable models of quantum reinforcement learning (QRL), formulated as finite-horizon Markov decision processes in finite-dimensional Hilbert spaces. The models are built around a `unitary-control-then-measure' protocol, in which a learning agent applies unitary transformations to a quantum state and interleaves each control step with a projective measurement onto a prescribed reference basis. Exact closed-form expressions for trajectory probabilities, rewards, and the expected return are derived for four concrete realisations: a closed-chain and an anti-periodic qubit implementation, a qutrit model with ladder coupling, and a four-level two-qubit system. Two structural features of these QRL protocols are then analysed. First, we identify and quantify the reduction in the computational complexity of the expected return, from the nominally exponential $O(e^N)$ scaling in the trajectory length~$N$ to an explicit power-law $O(N^{\mathcal{I}})$, driven by two rigorously established mechanisms, a trajectory equivalence and a sparsity of the transition graph, besides a third, conjectured one: a spectral concentration of the return, at the optimal policy, onto the polynomially populated trajectory classes. Second, we characterise the degeneracy of optimal policies. The low-dimensional models exhibit unique optima whose asymptotic behaviour with~$N$ is governed by the quantum Zeno effect, while the four-level system displays both plateau-type quasi-degeneracy at large horizons and genuine discrete degeneracy at critical energy parameters -- phenomena with no counterpart in the measurement-free quantum optimal control landscape.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Andrea Cintio, Alessandro Michelangeli, Dmitrii Tsutskov. 2026-07-02. Complexity scaling and optimal policy degeneracy in quantum reinforcement learning via analytically solvable unitary-control-then-measure models. https://arxiv.org/abs/2604.13096

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

On the convergence of a perturbed one dimensional Mann's process

We study a perturbed version of Mann's iterative process (DMP), defined by \[ x_{n+1} = (1 - θ_n)x_n + θ_n f(x_n) + r_n, \] where $f: [0,1] \to [0,1]$ is a continuous function, $\{θ_n\} \subset [0,1]$ is a given sequence, and $\{r_n\} $ represents an error term. We prove that if the sequence $\{θ_n\} $ converges sufficiently slowly to zero and the error term $ r_n $ is suitably small at infinity, then any sequence $\{x_n\} \subset [0,1] $ generated by this process converges to a fixed point of $f$. In addition, we investigate the asymptotic behavior of the trajectories $ x(t) $ as $ t \to \infty$ for a continuous-time version of the process (DMP). We emphasize the parallels between the discrete and the continuous dynamics. Furthermore, through numerical experiments, we analyze the influence of the sequence $\{θ_n\} $ and the error terms on the stability and the convergence rate of the the discrete and the continuous processes. Notably, we observe that the (DMP) algorithm, when affected by stochastic and relatively large error terms, can outperform the bisection method in efficiently identifying the fixed point set of the function $f$.

math.GM

Explicit formula for the discrete Laplace transform of the Möbius function, related special functions, and a criterion for the Riemann hypothesis

In this paper, we assume that all the zeros of the Riemann zeta function are simple. Under this assumption we give an explicit formula for the function $Φ(e^{-t})=\sum_{n=1}^{\infty}μ(n)e^{-nt}$, as a function of the values of $ζ(s)$ and $ζ'(s)$ at the odd integers and as a function of the zeros of $ζ(s)$. A structural feature distinguishes this formula from the classical explicit formula for the Mertens function: the poles of $Γ(s)$ collide with the trivial zeros of $ζ(s)$, producing double poles whose residues contain a logarithmic term. Using this formula, we give a criterion for the Riemann hypothesis: the bound $O(x^{-1/2})$ on the transform implies the Riemann hypothesis unconditionally, while the converse direction requires additional hypotheses on the zeros. We also introduce special entire functions related to $ζ(s)$ and show that they admit absolutely convergent closed forms as Möbius-weighted series of Bessel functions of rotated argument.

math.GM

There exist blow-ups in the incompressible Navier-Stokes equation

In this paper, we solve the Navier-Stokes equation, one of the seven Millennium Prize Problems suggested by Clay Mathematics Institute(CMI). In our paper, the fluid is confined in a solid sphere. We prove that, for any initial velocity u0, there has been a force vector f , such that there exists no smooth solution (p,u) to the respective Navier-Stokes equation. This result also holds for the Euler equation. The paper was first published in 2021, soon after, Professor PG Lemarié-Rieusset contact me and tell that there is a possible flaw, because u may be non-integrable when it decreases not very fast. Hence, I withdraw the paper. Through carefull consideration of several years, we think we can choose a fluid confined in a solid sphere to fix it. So, we update the paper now.

math.GM