arXiv Science⌕ Search

arXiv · 2609.12486

Adaptive Agent Design

Abstract

We consider an agent acting against a general non-Markovian environment. The agent maintains its agent states, but is free to choose a transition kernel across those states and optimize its state-feedback control policies. We study the bi-level agent design problem that optimizes the transition kernel and the policy it induces, given said kernel with offline data of observations and actions obtained via a behavioral policy. For general environments, we show that a soft $Q$-learning algorithm converges almost surely to the fixed point of a soft Bellman equation defined by the stationary averages that the behavioral policy and the chosen kernel induce, and we delineate what separates the resulting policy from an optimal one. In partially observed Markov decision problems, we analyze convergence properties of parametrized transition kernel design via zero-th order and Bayesian optimization techniques.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Raj Kiriti Velicheti, Subhonmesh Bose, Tamer Başar. 2026-09-11. Adaptive Agent Design. https://arxiv.org/abs/2609.12486

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Inference to Control: Structure-Guided Control of Hypergraph Dynamics

Controllability determines whether a system's state can be guided toward any desired configuration, making it a fundamental prerequisite for designing effective control strategies. In the context of networked systems, controllability is a well-established concept. However, many real-world systems, from biological collectives to engineered infrastructures, exhibit higher-order interactions that cannot be captured by simple graphs. Moreover, the interaction structures might be unknown and difficult to measure directly. Here, we close this gap by combining hypergraph inference with the identification of controllable nodes. Building on the inferred structure, we design a controller that, given a set of controllable nodes, steers the system toward a desired configuration. We formulate analytical controllability guarantees for polynomial systems. For non-polynomial dynamics on hypergraphs, we propose a heuristic method for identifying controllable nodes and validate the proposed approach using Kuramoto oscillators.

eess.SY↗

Input Dexterity and Output Negotiation in Feedback-Linearizable Nonlinear Systems

We introduce a task-relative taxonomy of actuator inputs for nonlinear systems within the input-output feedback-linearization framework. Given a flat output specifying the task, inputs are classified as essential, redundant, or dexterity: essential inputs are required for exact linearization, redundant inputs can be removed without effect, and dexterity inputs can be deactivated while preserving exact linearization of a reduced task. We show that a subset is dexterity if and only if, under a suitable dynamic prolongation, it can appear as additional output channels (flat-input complement) on a common validity set. Whenever a family of systems obtained by (de)activating dexterity inputs admits a common prolongation, the family can be interpreted as a single prolonged system endowed with different output selections. This enables a unified linearizing controller that negotiates between full and reduced tasks without transients on shared outputs under compatibility and dwell-time conditions. Simulations on a fully actuated aerial platform illustrate graceful task downgrades from six-dimensional pose tracking as lateral-force channels are deactivated.

eess.SY↗

Input-to-state stabilization of linear systems under data-rate constraints

We study feedback stabilization of linear systems under data-rate constraints in the presence of completely unknown disturbances. A communication and control strategy is proposed based on sampled and quantized state measurements, where the quantization range is dynamically adjusted using reachable-set approximations and a disturbance estimate derived from quantization parameters. The strategy alternates between stabilizing and searching stages to recapture the state after escapes from the quantization range. Under a data-rate condition, it guarantees input-to-state stability (ISS) with respect to the disturbance. An additional quantization symbol is introduced to establish ISS near the equilibrium. A simulation example illustrates the effectiveness of the proposed approach.

eess.SY↗