arXiv Science⌕ Search

arXiv · 2609.37316

Battery-Aware Reinforcement Learning for Aggressive Quadrotor Flight

Abstract

Agile flight tasks such as drone racing and pursuit-evasion require strong acceleration and precise turns, but the available thrust changes as the battery discharges and voltage drops under load. Conservative command limits make this variation easier to tolerate, at the cost of unused performance. We investigate how learned controllers can use that additional thrust while retaining the flight controller's voltage compensation and rate control. Our training simulator couples an identified load-transient battery model to rotor dynamics and firmware saturation. The feedforward policy receives filtered voltage during both training and deployment. Controlled ablations distinguish the benefit of a larger thrust-command range from that of voltage information. On a 38 g Crazyflie Brushless, the resulting policy reduces circle tracking error by 49% relative to stock-authority RL at 3.84 m/s, while preserving easy-task precision. Mean 20-lap race time decreases from 106.22 s to 95.24 s. Compared with a voltage-blind policy with the same increased authority, hardware error and race time are lower by 15.3% and 4.5%, respectively. In simulation, replacing the policy's voltage input with a recording from a different battery condition worsens hard-circle tracking, with a smaller, voltage-dependent effect in racing. Together, these results show where a simple voltage input complements existing actuator compensation in aggressive learned flight.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alejandro Sanchez Roncero, Olov Andersson, Petter Ogren. 2026-09-29. Battery-Aware Reinforcement Learning for Aggressive Quadrotor Flight. https://arxiv.org/abs/2609.37316

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An Real-Sim-Real (RSR) Loop Framework for Generalizable Robotic Policy Transfer

The sim-to-real gap remains a critical challenge in robotics, hindering the deployment of algorithms trained in simulation to real-world systems. We propose a flexible Real-to-Sim-to-Real (RSR) framework whose central contribution is an information-theoretic cost function that explicitly accounts for sim-to-real discrepancies. This cost balances two objectives, completing the task and steering the policy to collect real-world samples that are maximally informative for improving transfer. It can be integrated seamlessly into existing reinforcement learning algorithms (e.g., PPO, SAC) and ensures a balanced exploration of critical regions in the real domain. The framework treats differentiable simulation as optional: when a differentiable simulator is available, the collected informative data can also be used to tune simulator parameters. We implement the framework with the MuJoCo MJX platform and demonstrate its generality by evaluating on both manipulation tasks with a 6-DOF robotic arm and locomotion tasks on a legged robot. Empirical results show that our RSR loop yields more efficient data acquisition and substantially improves task performance in real-world that achieves a smoother sim-to-real transfer.

cs.RO↗

Learning-Based Progressive Barrier Control for Robot Manipulators with Initial Errors Outside Prescribed Tracking Bounds

Robot manipulators may start a new task with a tracking error larger than the prescribed tolerance. Conventional barrier controllers generally require the initial error to lie within this tolerance, which prevents their direct use under such conditions. This paper develops a progressive barrier controller that gradually contracts an initial error bound to the required value within a prescribed time. The robot can therefore start outside the final bound, while the direct joint-position error satisfies it after the transition. The closed-form control law combines progressive barrier feedback with an online adaptive torque term based on an actor--critic structure. A Lyapunov analysis establishes bounded closed-loop signals and gives sufficient conditions for satisfaction of the final tracking bound. Two-link simulations consider large initial errors, actuator saturation, dynamic variations, disturbances, and measurement errors. The adaptive term reduces the median root-mean-square tracking error by 45.1\% compared with the zero-weight progressive barrier controller. The simulations also map the initial errors that can be handled at different transition times under fixed torque limits. Finally, an experiment on a Niryo Ned3 Pro illustrates tracking performance using measured position and motor-current data.

cs.RO↗

BIM Informed Visual SLAM for Construction Environments

Monitoring building construction sites requires comparing the as-planned design with the as-built state, which can be estimated in real time using Simultaneous Localization and Mapping (SLAM) techniques. However, visual SLAM is prone to trajectory drift in construction environments, producing maps that are geometrically inaccurate with the actual environment. To address this limitation, we augment an existing RGB-D SLAM system with structural priors derived from the Building Information Model (BIM). The system associates detected walls with their BIM counterparts and includes these correspondences as geometric constraints in the back-end optimization, reducing drift and enhancing global consistency. The proposed method operates in real time and is validated on multiple real construction sites, achieving an average trajectory error reduction of 25.23% and a 7.14% improvement in map accuracy over state-of-the-art baselines. Robustness analyses further demonstrate resilience to incomplete BIM data and geometric discrepancies between as-planned models and the as-built environment.

cs.RO↗