arXiv · 2609.34486
Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction
Abstract
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Victor Paredes, Ayonga Hereid. 2026-09-28. Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction. https://arxiv.org/abs/2609.34486
Cite the original work for its findings. Save a collection to share your selection of sources.