arXiv Science⌕ Search

arXiv · 2610.06148

Reinforcement Learning-Based Optimization of Workload-Aware Power Delivery Networks

Abstract

Power Delivery Networks (PDNs) are critical components of modern VLSI chips, providing stable voltage levels while satisfying electromigration (EM) and IR-drop constraints. Conventional PDN design methodologies typically rely on worst-case assumptions, often resulting in over-provisioned networks and inefficient use of resources. This paper presents a reinforcement learning-based framework for the optimization of workload-aware PDNs. The proposed methodology first generates workload-aware PDNs using architectural power traces obtained from system-level simulations. These power traces are mapped to spatial power density distributions, enabling adaptive allocation of PDN resources according to local current demand. A reinforcement learning agent then performs wire-width optimization to minimize PDN area while maintaining EM and voltage integrity constraints. Electrical and reliability metrics are obtained using SPICE-based circuit analysis and EM lifetime estimation. Experimental evaluation is performed on a dataset of workload-aware PDNs generated from 4-, 8-, and 16-core multiprocessor floorplans using PARSEC and SPLASH-2 benchmark workloads. Furthermore, the proposed Deep Q-Network (DQN)-based optimizer reduces the average normalized PDN area by 47\% while satisfying all EM and IR-drop constraints. Compared to simulated annealing, the proposed approach achieves comparable optimization quality while providing approximately 26$\times$ faster optimization.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Oran Hayes, Maria Pantazi-Kypraiou, Athanasios Tziouvaras, George Stamoulis, Anuj Pathania, Shreejith Shanker, George Floros. 2026-10-05. Reinforcement Learning-Based Optimization of Workload-Aware Power Delivery Networks. https://arxiv.org/abs/2610.06148

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Self-Calibrating Framework for Analog Circuit Sizing Using LLM-Derived Analytical Equations

We present a design automation framework for analog circuit sizing that produces calibrated, topology-specific analytical equations from raw circuit netlists. A large language model (LLM) derives a complete Python sizing function in which each device dimension is traceable to a specific design rationale - a form of interpretable output absent from existing optimization-based and LLM-based sizing methods. A deterministic calibration loop extracts process-dependent parameters from a single DC operating point simulation, while a prediction-error feedback mechanism compensates for analytical inaccuracies. We validate the framework on circuits ranging from 6 to 30 transistors - spanning single-stage, current-mirror (simple and cascoded), folded-cascode, gain-boosted folded-cascode, two-stage Miller-compensated, nested-Miller-compensated, and complementary class-AB output topologies - across six process nodes from 32 nm to 180 nm. On matched-specification benchmarks, including the class-AB opamp case, the framework converges within a few simulations. Despite large initial prediction errors, convergence depends on the measurement-feedback architecture, not prediction accuracy. The one-shot calibration automatically captures process-dependent variations, enabling cross-node portability without modification, retraining, or per-process characterization.

cs.AR↗

Towards Enabling Distance-Based Memory Addressing

Approximate Nearest-Neighbor Search (ANNS) in high dimensional vector datasets is an application of significant prevalence across different AI applications. However, such an operation is significantly bandwidth limited at large workingset sizes owing to the curse of dimensionality. Traditional indices used to accelerate ANNS rely on search-space pruning as a preprocessing step to alleviate such bandwidth requirement, but such optimization occurs either at the cost of increased bandwidth-inefficiency and/or degradation of search quality. This paper proposes a data-parallel hardware/software mechanism for performing large-scale similarity search in-memory. We propose a novel algorithm to simplify the computation requirement for similarity search across various distance metrics through lightweight primitives to perform a fast and approximate data-parallel brute-force search on the entire vector space. We further build a memory system capable of executing the required operations to generate a distance metric per datapoints, which is then used to enable pruning as a post-processing step. We offer adequate software support for user control over the proposed system. By enabling such search-space pruning as a post-processing step, we achieve near-perfect recall across representative workloads while achieving orders of magnitude performance and energy improvement over state-of-the-art algorithmic approaches on million and billion-scale workloads.

cs.AR↗

Terracotta: Enabling the Adoption of New DRAM Techniques via a Flexible DRAM Interface and Memory Controller

DRAM continues to limit the performance, energy efficiency, and robustness of modern systems. Many prior works propose DRAM techniques that support in-DRAM computation, improve memory access latency and parallelism, and enhance DRAM maintenance and reliability. However, adopting each new technique requires repeated modifications to the rigid DRAM interface and memory controller, hindering its deployment. Our goal is to reduce these repeated modifications. We observe that the DRAM commands and memory controller structures of many DRAM techniques are similar. Our key idea is to use these similarities to compose a set of primitives for implementing diverse DRAM techniques. We propose Terracotta, a new framework with two flexible components: (i) custom command extensions that let DRAM vendors define new commands within a single, standardized interface, and (ii) a programmable memory controller that system designers can program to support new DRAM techniques post-silicon. Together, these enable deployment by configuring the memory controller instead of modifying the interface and controller. We design Terracotta for a DDR5-based system and evaluate its performance, energy, and hardware complexity. For four DRAM techniques from four distinct domains (processing-using-DRAM, low-cost DRAM maintenance, subarray-level parallelism, and latency reduction), Terracotta retains almost all of the performance benefits (>96%) of custom implementations. A Terracotta-based composition of two techniques outperforms the Terracotta-based implementation of each technique alone, demonstrating the benefits of adding techniques without repeated interface and controller modifications. Terracotta incurs low DRAM energy (0.6-3.2%), area (0.03%), and power (0.56%) overheads in a high-end server-grade processor. Terracotta's source code is freely available at https://github.com/CMU-SAFARI/Terracotta.

cs.AR↗