Reinforcement learning for post-coronagraphic wavefront control
Direct imaging of exoplanets is limited by the extreme contrast between the star and the planets, which is mitigated using a coronagraph. However, optical aberrations cause starlight leakage through the coronagraph, producing speckles that obscure the planetary signal. Achieving the required contrast levels demands wavefront control with subnanometric precision. Deep reinforcement learning offers a promising alternative to traditional focal-plane wavefront control techniques by enabling adaptive correction strategies learned directly from interaction with the system. In this work, we present a fully data-driven method for post-coronagraphic aberration correction in a simulated high-contrast imaging testbed. The agent controls a deformable mirror using observations consisting of focal-plane measurements (images) and physics-informed wavefront sensing information derived from these images. We evaluate different observation representations and control strategies, and the method is validated on simplified simulations of a high-contrast imaging testbed, where it successfully creates dark holes, i.e., regions of the focal plane in which residual starlight is strongly suppressed, while approaching the performance of conventional wavefront control methods.