arXiv · 2307.00141
Risk-sensitive Actor-free Policy via Convex Optimization
Abstract
Traditional reinforcement learning methods optimize agents without considering safety, potentially resulting in unintended consequences. In this paper, we propose an optimal actor-free policy that optimizes a risk-sensitive criterion based on the conditional value at risk. The risk-sensitive objective function is modeled using an input-convex neural network ensuring convexity with respect to the actions and enabling the identification of globally optimal actions through simple gradient-following methods. Experimental results demonstrate the efficacy of our approach in maintaining effective risk control.
Explore related subjects
Keep this discovery
Ruoqi Zhang, Jens Sjölund. 2023-06-30. Risk-sensitive Actor-free Policy via Convex Optimization. https://arxiv.org/abs/2307.00141
Cite the original work for its findings. Save a collection to share your selection of sources.