arXiv ScienceSearch

arXiv subjects

Ryan Self

Publications and source records attributed to Ryan Self.

6 recordsLinked to original sources

Switched Optimal Control and Dwell Time Constraints: A Preliminary Study

Most modern control systems are switched, meaning they have continuous as well as discrete decision variables. Switched systems often have constraints called dwell-time constraints (e.g., cycling constraints in a heat pump) on the switching rate. This paper introduces an embedding-based-method to solve optimal control problems that have both discrete and continuous decision variables. Unlike existing methods, the developed technique can heuristically incorporate dwell-time constraints via an auxiliary cost, while also preserving other state and control constrains of the problem. Simulations results for a switched optimal control problem with and without the auxiliary cost showcase the utility of the developed method.

eess.SY

Online Observer-Based Inverse Reinforcement Learning

In this paper, a novel approach to the output-feedback inverse reinforcement learning (IRL) problem is developed by casting the IRL problem, for linear systems with quadratic cost functions, as a state estimation problem. Two observer-based techniques for IRL are developed, including a novel observer method that re-uses previous state estimates via history stacks. Theoretical guarantees for convergence and robustness are established under appropriate excitation conditions. Simulations demonstrate the performance of the developed observers and filters under noisy and noise-free measurements.

eess.SY

Online inverse reinforcement learning with limited data

This paper addresses the problem of online inverse reinforcement learning for systems with limited data and uncertain dynamics. In the developed approach, the state and control trajectories are recorded online by observing an agent perform a task, and reward function estimation is performed in real-time using a novel inverse reinforcement learning approach. Parameter estimation is performed concurrently to help compensate for uncertainties in the agent's dynamics. Data insufficiency is resolved by developing a data-driven update law to estimate the optimal feedback controller. The estimated controller can then be queried to artificially create additional data to drive reward function estimation.

eess.SY

Online inverse reinforcement learning with unknown disturbances

This paper addresses the problem of online inverse reinforcement learning for nonlinear systems with modeling uncertainties while in the presence of unknown disturbances. The developed approach observes state and input trajectories for an agent and identifies the unknown reward function online. Sub-optimality introduced in the observed trajectories by the unknown external disturbance is compensated for using a novel model-based inverse reinforcement learning approach. The observer estimates the external disturbances and uses the resulting estimates to learn the dynamic model of the demonstrator. The learned demonstrator model along with the observed suboptimal trajectories are used to implement inverse reinforcement learning. Theoretical guarantees are provided using Lyapunov theory and a simulation example is shown to demonstrate the effectiveness of the proposed technique.

eess.SY

Output-feedback online optimal control for a class of nonlinear systems

In this paper an output-feedback model-based reinforcement learning (MBRL) method for a class of second-order nonlinear systems is developed. The control technique uses exact model knowledge and integrates a dynamic state estimator within the model-based reinforcement learning framework to achieve output-feedback MBRL. Simulation results demonstrate the efficacy of the developed method.

eess.SY

Online inverse reinforcement learning for nonlinear systems

This paper focuses on the development of an online inverse reinforcement learning (IRL) technique for a class of nonlinear systems. The developed approach utilizes observed state and input trajectories, and determines the unknown cost function and the unknown value function online. A parameter estimation technique is utilized to allow the developed IRL technique to determine the cost function weights in the presence of unknown dynamics. Simulation results are presented for a nonlinear system showing convergence of both unknown reward function weights and unknown dynamics.

eess.SY