arXiv ScienceSearch

arXiv · 2609.07045

Equivalence Between Continuous-Time Risk-Sensitive Control and Rényi Divergence Minimization

Abstract

In this study, we show that a continuous-time risk-sensitive control problem is equivalent to a Rényi divergence minimization problem over trajectory path measures. Reformulating stochastic optimal control as probabilistic inference via Kullback-Leibler (KL) divergence minimization avoids the computational intractability of the Hamilton-Jacobi-Bellman equation. However, standard KL control is inherently risk-neutral, and recent minimax extensions remain restricted to risk-averse settings. Our equivalence result resolves this limitation by offering a unified probabilistic framework for arbitrary risk attitudes in continuous-time nonlinear systems. Based on Girsanov theorem, we explicitly map the risk sensitivity to the Rényi divergence order, deriving a noise-dependent control penalty scaled by risk preference. This formulation seamlessly modulates tail-weighting behaviors, interpolating between zero-forcing for risk-averse policies and mass-covering for risk-seeking policies. These findings bridge stochastic control and information-theoretic inference, providing a foundation for sampling-based control algorithms.

Explore related subjects

Keep this discovery

BibTeXRIS

Shinji Kataoka, Kaoru Teranishi, Yasumasa Fujisaki. 2026-09-07. Equivalence Between Continuous-Time Risk-Sensitive Control and Rényi Divergence Minimization. https://arxiv.org/abs/2609.07045

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

A simple derivation of the Kalman filter

In this lecture note, we present a concise and self-contained derivation of the discrete-time Kalman filter equations that requires only a basic understanding of least squares estimation. The treatment is designed to minimize mathematical overhead while preserving both rigor and generality.

math.OC

Comment on "Event-Triggered Stabilization of Linear Time-Delay Systems via Halanay-Type Inequality"

This comment revisits Lemma 1 in [1], which plays a central role in the event-triggered stabilization analysis developed therein. We identify technical gaps in the proof of the lemma and provide a corrected argument. In particular, careful treatment of the exponentially decaying term shows that its decay rate must be retained in the resulting convergence estimate. The statement of the original lemma, with the exponential decay rate determined by the minimum of the characteristic decay rate and the decay rate of this term, remains valid.

math.OC