arXiv Science⌕ Search

arXiv · 2610.09266

Cooperative Dueling DQN SAC Learning for Energy Efficiency in Dynamic OWC Networks

Abstract

Growing wireless traffic is increasing pressure on the congested radio-frequency spectrum. Optical wireless communication (OWC) provides a complementary solution by using the abundant unlicensed optical spectrum. However, indoor OWC networks are dynamic: users move, enter or leave the network, and subscribe to different services. Poorly coordinated resource allocation can consequently waste subcarriers, require excessive transmission power and cause frequent AP reassignments. Energy efficiency (EE), defined as the total delivered data rate divided by the total network power consumption, therefore requires the serving AP, number of allocated subcarriers, and transmission power to be jointly adapted while maintaining QoS. Optimising these in a dynamic time series, multi-service OWC environment produces a complex sequential EE optimisation problem. To address this problem, this work proposes Dual-Agent Resource Allocation using Deep Reinforcement Learning (DARA-DRL). DARA-DRL combines a branching duelling deep Q-network for association and subcarrier allocation with a conditional soft actor-critic agent for continuous power control. The agents are coupled through a common reward and a cooperative value update that evaluates each discrete allocation together with its corresponding power decision. Simulation results show that DARA-DRL remains within 5\% of the optimal solution and, compared with state-of-the-art benchmarks, it improves EE by 28.8\% and QoS satisfaction by 10.5\%, while reducing online decision time by 14.5\%. Results demonstrate that agent specialisation simplifies mixed-action learning, and cooperation outperforms independently trained agents.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Walter Zibusiso Ncube, Ahmad Adnan Qidan, Taisir El-Gorashi, Jaafar M. H. Elmirghani. 2026-10-07. Cooperative Dueling DQN SAC Learning for Energy Efficiency in Dynamic OWC Networks. https://arxiv.org/abs/2610.09266

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Global Geolocated Realtime Data of Interfleet Urban Transit Bus Idling

Urban transit bus idling is a contributor to ecological stress, economic inefficiency, and medically hazardous health outcomes due to emissions. The global accumulation of this frequent pattern of undesirable driving behavior is enormous. In order to measure its scale, we propose GRD-TRT-BUF-4I (Ground Truth Buffer for Idling) an extensible, realtime detection system that records the geolocation and idling duration of urban transit bus fleets internationally. Using live vehicle locations from General Transit Feed Specification (GTFS) Realtime, the system detects approximately 200,000 idling events per day from over 50 cities across North America, Europe, Oceania, and Asia. This realtime data was created dynamically to serve operational decision-making and fleet management to reduce the frequency and duration of idling events as they occur, as well as to capture its accumulative effects. Civil and Transportation Engineers, Urban Planners, Epidemiologists, Policymakers, and other stakeholders might find this useful for emissions modeling, traffic management, route planning, and other urban sustainability efforts at a variety of geographic and temporal scales.

eess.SY↗

Confidence-Aware Safe and Stable Control of Control-Affine Systems

Designing control inputs that satisfy safety requirements is crucial in safety-critical nonlinear control, and this task becomes particularly challenging when full-state measurements are unavailable. In this work, we address the problem of synthesizing safe and stable control for control-affine systems via output feedback (using an observer) while reducing the estimation error of the observer. To achieve this, we adapt control Lyapunov function (CLF) and control barrier function (CBF) techniques to the output feedback setting. Building upon the existing CLF-CBF-QP (Quadratic Program) and CBF-QP frameworks, we formulate two confidence-aware optimization problems and establish the Lipschitz continuity of the obtained solutions. To validate our approach, we conduct simulation studies on two illustrative examples. The simulation studies indicate both improvements in the observer's estimation accuracy and the fulfillment of safety and control requirements.

eess.SY↗

Collision Avoidance for Convex Primitives via Differentiable Optimization Based High-Order Control Barrier Functions

Ensuring the safety of dynamical systems is crucial, where collision avoidance is a primary concern. Recently, control barrier functions (CBFs) have emerged as an effective method to integrate safety constraints into control synthesis through optimization techniques. However, challenges persist when dealing with convex primitives and tasks requiring torque control, as well as the occurrence of unintended equilibria. This work addresses these challenges by introducing a high-order CBF (HOCBF) framework for collision avoidance among convex primitives. We transform nonconvex safety constraints into linear constraints by differentiable optimization and prove the high-order continuous differentiability. Then, we employ HOCBFs to accommodate torque control, enabling tasks involving forces or high dynamics. Additionally, we analyze the issue of spurious equilibria in high-order cases and propose a circulation mechanism to prevent the undesired equilibria on the boundary of the safe set. Finally, we validate our framework with three experiments on the Franka Research 3 robotic manipulator, demonstrating successful collision avoidance and the efficacy of the circulation mechanism.

eess.SY↗