arXiv Science⌕ Search

arXiv · 2609.33795

Deep Behaviour Cloning of Model Predictive Control for Real-Time Operation of a Hydrogen-Diesel Dual-Fuel Engine

Abstract

Hydrogen-diesel dual-fuel (H2DF) combustion engines offer a promising pathway for decarbonising hard-to-electrify transport sectors, yet their highly nonlinear dynamics and coupled process variables demand constraint-aware control strategies. Model Predictive Control (MPC) meets these requirements but requires an online optimisation in every combustion cycle, which limits deployment on low-cost embedded hardware. This paper trains a feedforward deep neural network (DNN) by behaviour cloning (BC) to imitate an MPC expert, using 86,000 engine cycles of demonstration data collected at 1500 min-1 on a modified Cummins QSB 4.5-litre hydrogen dual-fuel engine. Two variants, one with process feedback and one without, are validated experimentally. Both track unseen fast-transient load steps (3-8 bar indicated mean effective pressure, IMEP) with normalised root mean square error (NRMSE) values of 7.80% and 9.03% against the expert's 8.01%, while keeping mean NOx and particulate matter emissions at or below those of the expert. Inference takes 2 ms or less on a Raspberry Pi 400 (ARM Cortex-A72 at 2.2 GHz), including 1 ms communication latency, compared to up to 7 ms for the MPC expert, a 3.5x speedup. Open-loop profiling on a low-cost ESP32 microcontroller at 180 MHz gives 4.3 ms per inference, 4x faster than required for the 18 ms cycle window. Beyond the training range the cloned policy saturates its controls but exceeds the pressure-rise-rate limit. To the authors' knowledge, this is the first experimental BC controller for cycle-to-cycle combustion control of an internal combustion engine (ICE), trained from demonstrations recorded on the engine itself.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alexander Winkler, Neeraj Naduvath Mana, David Gordon, Jakob Andert. 2026-09-27. Deep Behaviour Cloning of Model Predictive Control for Real-Time Operation of a Hydrogen-Diesel Dual-Fuel Engine. https://arxiv.org/abs/2609.33795

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Reinforcement Learning to Initialize Newton-Raphson for AC Power Flow with Quantum Annealing-Based Environment Updates

The Newton-Raphson (NR) method is widely used for solving power flow (PF) equations due to its quadratic convergence. However, its performance deteriorates under poor initialization or extreme operating scenarios, e.g., high levels of renewable energy penetration. We propose the use of reinforcement learning (RL) to optimize the initialization of NR, and introduce a quantum-enhanced RL environment update mechanism that addresses the combinatorially large action space at each RL timestep by formulating the voltage adjustment task as a Quadratic Unconstrained Binary Optimization (QUBO) problem, solved with an Ising machine. RL initialization is benchmarked against flat start and start from the DC (linearized) PF solution on a standard 4-bus system, Iwamoto's ill-conditioned 11-bus system, and the IEEE 118-bus system under normal and stressed loading and reactive power limits, with verified operational solutions. On all systems, a supervised initializer refined by RL requires fewer NR iterations than flat and DC starts and than the same initializer without RL, for all seeds. For example, on the 118-bus system under normal and stressed loading, it reached 2.04 and 2.86 NR iterations, compared with 3.02 and 5.13 from DC start and 2.61 and 3.09 without RL. In wall-clock time, this pays off only for an initializer integrated into the solver and reused for many solves on a fixed topology. On the 4-bus system, a quantum-enhanced RL agent with a quantum-inspired annealer moved challenging initial states that required 29 and 44 NR iterations to initializations that required three NR iterations within one RL timestep.

eess.SY↗

Safety for Weakly-Hard Control Systems via Graph-Based Barrier Functions

Despite significant advancement in technology, communication and computational failures are still prevalent in safety-critical engineering applications. Often, networked control systems experience packet dropouts, leading to open-loop behavior that significantly affects the behavior of the system. Similarly, in real-time control applications, control tasks frequently experience computational overruns and thus occasionally no new actuator command is issued. This article addresses the safety verification and controller synthesis problem for a class of control systems subject to weakly-hard constraints, i.e., a set of window-based constraints where the number of failures are bounded within a given time horizon. The results are based on a new notion of graph-based barrier functions that are specifically tailored to the considered system class, offering a set of constraints whose satisfaction leads to safety guarantees despite such failures. Subsequent reformulations of the safety constraints are proposed to alleviate conservatism and improve computational tractability, and the resulting trade-offs are discussed. Finally, several numerical case studies including linear and polynomial systems demonstrate the effectiveness of the proposed approach.

eess.SY↗

Model-Free Output Feedback Stabilization via Policy Gradient Methods

Stabilizing a dynamical system is a fundamental problem that serves as a cornerstone for many complex tasks in the field of control systems. The problem becomes challenging when the system model is unknown. Among the Reinforcement Learning (RL) algorithms that have been successfully applied to solve problems pertaining to unknown linear dynamical systems, the policy gradient (PG) method stands out due to its ease of implementation and can solve the problem in a model-free manner. However, most of the existing works on PG methods for unknown linear dynamical systems assume full-state feedback. In this paper, we take a step towards model-free learning for partially observed linear dynamical systems with output feedback and focus on the fundamental stabilization problem of the system. We propose an algorithmic framework that stretches the boundary of PG methods to the problem without global convergence guarantees. We show that by leveraging zeroth-order PG update based on system trajectories and its convergence to stationary points, the proposed algorithms return a stabilizing output feedback policy for discrete-time linear dynamical systems. We also explicitly characterize the sample complexity of our algorithm and verify the effectiveness of the algorithm using numerical examples.

eess.SY↗