arXiv ScienceSearch

arXiv · 2602.03939

C-IDS: Solving Contextual POMDP via Information-Directed Objective

Abstract

We study the policy synthesis problem in contextual partially observable Markov decision processes (CPOMDPs), where the environment is governed by an unknown latent context that induces distinct POMDP dynamics. Our goal is to design a policy that simultaneously maximizes cumulative return and actively reduces uncertainty about the underlying context. We introduce an information-directed objective that augments reward maximization with mutual information between the latent context and the agent's observations. We develop the C-IDS algorithm to synthesize policies that maximize the information-directed objective. We show that the objective can be interpreted as a Lagrangian relaxation of the linear information ratio and prove that the temperature parameter is an upper bound on the information ratio. Based on this characterization, we establish a sublinear Bayesian regret bound over K episodes. We evaluate our approach on a continuous Light-Dark environment and show that it consistently outperforms standard POMDP solvers that treat the unknown context as a latent state variable, achieving faster context identification and higher returns.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chongyang Shi, Michael Dorothy, Jie Fu. 2026-02-03. C-IDS: Solving Contextual POMDP via Information-Directed Objective. https://arxiv.org/abs/2602.03939

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Joint Spectrum and Airspace Resource Optimization for Low-altitude Wireless Network

Low-altitude wireless networks have emerged as a promising platform for enabling the safe and efficient operation of unmanned aerial vehicles (UAVs). However, due to the limited spectrum and airspace resources, it is challenging to efficiently accomplish UAV flight tasks without collisions. In this paper, we propose a sequential framework with two coupled stages that coordinates spectrum allocation and airspace planning to construct efficient low-altitude air corridors. Specifically, the low-altitude airspace is discretized into a set of digital grids, where obstacles are modeled as impermeable units. Then, we formulate an optimization problem to minimize the total traversal cost of air corridors, which is challenging to solve due to the tight coupling between spectrum allocation and path planning.Therefore, we first design a constrained Vickrey-Clarke-Groves (VCG) ascending auction mechanism to allocate the spectrum resources. Then, we propose a joint spectrum and airspace resource allocation algorithm to minimize the total traversal cost of air corridors. Finally, simulation results show that the proposed algorithms achieve lower total costs than the baseline algorithms.

eess.SY

IMU-Centric Moving Horizon Estimation for Lateral Dynamics Estimation Across Vehicles and Grip Conditions

Accurate estimation of lateral vehicle dynamics near the adhesion limit is important for stability control and high-performance driving, but lateral velocity is rarely measured directly because sensors such as optical sensors are costly. This paper presents an inertial measurement unit (IMU)-centric Moving Horizon Estimation framework that reconstructs lateral velocity using standard onboard signals, without relying on exteroceptive odometry or detailed tire-parameter tuning. Experimental validation on human-driven sports cars and an autonomous open-wheel race car across tracks, maneuvers, and conditions demonstrates accurate and robust lateral velocity and lateral acceleration estimates. The proposed framework is available at https://github.com/Aseuffo/IMU-Centric-MHE

eess.SY

Structured Stochastic Representations of Integrated Dynamic Strategies

Dynamic allocation decisions couple present resource use to evolving internal conditions, delayed returns, and future costs. We represent this interaction by four probability localizations linked through regime-indexed, graph-constrained column-stochastic operators. Pre-action state or context selects a locally affine model, while action-dependent changes update subsequent regimes, yielding a causal switched representation of nonlinear evolution. We characterize operator identifiability relative to the graph, the stochastic constraints, and the sampled embedding, separating coefficient recovery from predictive equivalence on the decision domain. Decision making is then formulated through implementable return--cost acceptability regions. Finite-horizon error propagation supplies conservative classification margins, and simultaneous intervals distinguish model-relative near-optimality from certified $ε$-optimality over a declared finite policy class. Regime-indexed stochastic feedback is admitted when it satisfies the same certification test. Reproducible synthetic laboratories for personal preparation, supplier participation, and customer retention illustrate exact, operator-supplied, and noisy feedback cases. Multinomial experiments show improving recovery of the feedback function and fewer unresolved decisions with increasing sample size, while unrestricted off-policy recovery remains limited. The contribution is a structure-preserving representation--identification--decision workflow, not a domain-specific physiological or commercial calibration.

eess.SY