arXiv ScienceSearch

arXiv · 2609.01764

Curriculum-Guided Reinforcement Learning for Energy-Efficient UAV-ISAC in Post-Disaster Search-and-Rescue Operations

Abstract

Uncrewed aerial vehicles (UAVs) are promising platforms for integrated sensing and communication (ISAC), but their limited onboard energy creates a strong coupling among sensing accuracy, communication quality, and propulsion cost. This paper proposes a curriculum-guided soft actor-critic (CG-SAC) framework with propulsion-aware reward shaping for energy-efficient UAV-ISAC, jointly optimizing the 3D trajectory, communication-sensing power split, and per-user power allocation. A rotary-wing propulsion model is incorporated to derive a closed-form propulsion-economic cruising speed, which is used to construct a propulsion-aware speed-shaping term within a normalized composite reward together with navigation, node-visiting, energy-efficiency, and constraint-penalty terms. A log-linear curriculum progressively tightens the communication, sensing, and proximity requirements during training. Across 2000 randomized scenarios, CG-SAC achieves an average energy efficiency of 0.72 Mbits/J, substantially outperforming the evaluated DRL baselines. Among successfully completed missions, it requires 107.6 steps on average, corresponding to a 66%--82% reduction in flight steps relative to the baselines. Crucially, the learned policy exhibits mission-aware speed adaptation by decelerating near service points and accelerating during transit, while achieving a 99.6% communication-rate satisfaction ratio at service instants. Ablation results further demonstrate the complementary roles of the reward components in balancing mission feasibility and energy efficiency.

Explore related subjects

Keep this discovery

BibTeXRIS

Tai-You Guo, Chuan-Chi Lai. 2026-09-01. Curriculum-Guided Reinforcement Learning for Energy-Efficient UAV-ISAC in Post-Disaster Search-and-Rescue Operations. https://arxiv.org/abs/2609.01764

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Predictive Lightweight MARL for Resilient Coverage in Sparse-Signaling Aerial Networks

This letter proposes the Predictive Lightweight Multi-Agent Reinforcement Learning (PL-MARL) framework to ensure resilient coverage in bandwidth-constrained UAV swarms. To counter coordination collapse caused by sparse signaling and information aging, we introduce a Kinematic-Aware Inference Engine that proactively reconstructs neighbor trajectories via physical priors. This approach enables an efficient computation-for-communication trade-off, decoupling structural resilience from signaling frequency. Simulations confirm that PL-MARL maintains superior coverage and mission continuity under extreme signaling scarcity and node failure. Our results validate proactive inference as a scalable, low-latency solution for robust aerial coordination, effectively minimizing control overhead to preserve spectrum for payload services while ensuring resilience against interference.

cs.NI

Adaptive RIS-aided Communications through ML-based Generation of Phase Masks

Reconfigurable Intelligent Surfaces (RISs) are an attractive technology for Millimeter Wave (mmWave) communications due to their ability to passively reflect incident signals. However, current implementations of RIS rely on performing computationally-intensive algorithms offline to generate phase masks, which are stored as a codebook on the embedded microcontroller on the RIS. The codebook size is restricted by the embedded microcontroller's storage capacity, which limits the ability of the RIS to adapt to evolving channel conditions and deployment scenarios. In this demo, we showcase an Machine Learning (ML)-based solution for dynamically generating new phase masks during runtime. Our approach leverages a ML model deployed on the microcontroller for approximating the output of a phase mask generation algorithm, responding to new inputs while remaining smaller than a codebook.

eess.SY

Low-Complexity Control Under Input Saturation and Performance Constraints: A Bidirectional Modification Scheme

This article addresses the output tracking control problem for a class of high-order uncertain highly-coupled MIMO nonlinear systems subject to input saturation and performance constraints. To resolve the problem, a bidirectional modification mechanism is constructed, which is able to not only relax the constraints when saturation occurs to alleviate potential conflict, but also accelerate the recovery of original constraints after saturation ceases, and further tighten the constraints to enhance control performance if saturation remains inactive at the steadystate phase. Based on the mechanism, a model-, approximationand complexity-explosion-free control scheme is proposed. To bypass the obstacle in Lyapunov analysis, a novel stability analysis framework is developed, which, given that two parameter selection conditions are met, ensures satisfaction of modified constraints and boundedness of all closed-loop signals. Simulation results validate the effectiveness and superiority of the methodology.

eess.SY