Learning Energy-Efficient Air--Ground Actuation for Hybrid Robots on Stair-Like Terrain
Hybrid aerial--ground robots can use thrust to cross obstacles that impede wheel-driven motion, but deciding how much thrust to apply during contact remains challenging. We present an energy-aware reinforcement learning framework that jointly commands wheels, tilt servos and propellers through a single continuous policy, without prescribing locomotion modes. Hardware-calibrated power models penalise estimated electrical energy, while a terrain curriculum and a terminal reward for upright, settled arrivals support learning of thrust-assisted climbing. In simulation, continuous thrust allocation improves single-step clearance over fixed-thrust and mode-switching baselines as steps become taller. At the wheel-radius step height, it draws approximately 27 percent less mean power than the best fixed allocation. An energy-weight ablation shows the accompanying trade-off between efficiency and reliability. On a physical DoubleBee prototype, the unchanged network with a deployment interface clears two 6 centimetre steps in 8 of 10 trials. Additional hardware tests show both the potential for transfer to different terrain geometries and the remaining limitations in heading control, contact robustness and recovery.