arXiv Science⌕ Search

arXiv · 2610.03508

Hierarchical Control via MPC-RL for Multi-Timescale Battery Systems

Abstract

Multi-timescale systems present a fundamental challenge, where fast operational decisions must coexist with long-horizon sustainability targets. In this work, we propose a new hierarchical control framework via Model Predictive Control (MPC) and Reinforcement Learning (RL) to separate decision-making on two distinct timescales. The high-level MPC optimizes long-horizon setpoints at the slow dynamic and on a fast timescale, a low-level pretrained RL agent tracks these setpoints in real time to maximize short-term objectives. RL is introduced to learn nonlinear control policies, without relying on model linearizations or requiring the heavy online computation from solving repeated optimal control problems. The framework is applied to a Battery Energy Storage System (BESS) operating in frequency regulation markets to balance fast profit opportunities (seconds) and slow battery degradation (weeks to months). The design employs a degradation-aware RL agent trained offline to generate safe long-horizon setpoints, and a degradation-unaware agent fine-tuned from it for fast runtime setpoint tracking. Compared to MPC baselines, the proposed approach successfully extends battery lifetime by 84% and increases operational profit by 34%.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rasa Pourjam, Ehecatl Antonio del Río Chanona, Paulina Quintanilla. 2026-10-02. Hierarchical Control via MPC-RL for Multi-Timescale Battery Systems. https://arxiv.org/abs/2610.03508

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Inference to Control: Structure-Guided Control of Hypergraph Dynamics

Controllability determines whether a system's state can be guided toward any desired configuration, making it a fundamental prerequisite for designing effective control strategies. In the context of networked systems, controllability is a well-established concept. However, many real-world systems, from biological collectives to engineered infrastructures, exhibit higher-order interactions that cannot be captured by simple graphs. Moreover, the interaction structures might be unknown and difficult to measure directly. Here, we close this gap by combining hypergraph inference with the identification of controllable nodes. Building on the inferred structure, we design a controller that, given a set of controllable nodes, steers the system toward a desired configuration. We formulate analytical controllability guarantees for polynomial systems. For non-polynomial dynamics on hypergraphs, we propose a heuristic method for identifying controllable nodes and validate the proposed approach using Kuramoto oscillators.

eess.SY↗

Input Dexterity and Output Negotiation in Feedback-Linearizable Nonlinear Systems

We introduce a task-relative taxonomy of actuator inputs for nonlinear systems within the input-output feedback-linearization framework. Given a flat output specifying the task, inputs are classified as essential, redundant, or dexterity: essential inputs are required for exact linearization, redundant inputs can be removed without effect, and dexterity inputs can be deactivated while preserving exact linearization of a reduced task. We show that a subset is dexterity if and only if, under a suitable dynamic prolongation, it can appear as additional output channels (flat-input complement) on a common validity set. Whenever a family of systems obtained by (de)activating dexterity inputs admits a common prolongation, the family can be interpreted as a single prolonged system endowed with different output selections. This enables a unified linearizing controller that negotiates between full and reduced tasks without transients on shared outputs under compatibility and dwell-time conditions. Simulations on a fully actuated aerial platform illustrate graceful task downgrades from six-dimensional pose tracking as lateral-force channels are deactivated.

eess.SY↗

Input-to-state stabilization of linear systems under data-rate constraints

We study feedback stabilization of linear systems under data-rate constraints in the presence of completely unknown disturbances. A communication and control strategy is proposed based on sampled and quantized state measurements, where the quantization range is dynamically adjusted using reachable-set approximations and a disturbance estimate derived from quantization parameters. The strategy alternates between stabilizing and searching stages to recapture the state after escapes from the quantization range. Under a data-rate condition, it guarantees input-to-state stability (ISS) with respect to the disturbance. An additional quantization symbol is introduced to establish ISS near the equilibrium. A simulation example illustrates the effectiveness of the proposed approach.

eess.SY↗