On Data-Driven Model Identification for Nonlinear Optimal Control
In this paper, we study the use of nonlinear model identification techniques for the optimal control of nonlinear systems, also known as model-based Reinforcement Learning. We show that the nonlinear model identification problem is equivalent to estimating the generalized moments of an underlying sampling distribution and is bound to suffer from ill-conditioning and variance when approximating a system to high order and over a large domain, requiring samples combinatorial-exponential in the order of the approximation and domain size: a ``Curse of Variance and Ill-Conditioning (COVIC)" that shows up even in very low dimensional problems, quite apart from the usual ``Curse of Dimensionality". We show that the iterative identification of ``local" linear time varying (LTV) models around the current estimate of the optimal trajectory, coupled with a suitable optimal control algorithm such as iterative LQR (ILQR), alleviates these issues and is sufficient to locally accurately solve the underlying optimal control problem.