arXiv Science⌕ Search

arXiv · 2610.02401

Design Space Exploration of Backside Clock Meshes for 2 nm GAAFET BSPDN Technology

Abstract

Clock meshes are used in high-performance VLSI designs to minimize skew and tolerate on-chip variation, but they spend scarce routing resources on premium metal layers. Backside power delivery creates a new option: it adds thick, low-resistance metal layers on the back of the wafer, and the power grid does not consume all of them. Flip-flops remain on the frontside; a backside mesh therefore cannot drive them directly, and every connection passes through a through-silicon via. We present the first design-space exploration of backside clock meshes, implemented in OpenROAD on GT2N, a 2 nm nanosheet technology. Four benchmarks (1,938 to 15,311 flip-flops) are explored with multi-objective Bayesian optimization, and every design point is verified by transistor-level SPICE simulation, since the cyclic mesh cannot be evaluated by static timing analysis. Across all four designs, the backside mesh consistently outperforms an identical frontside mesh, with on average 45% lower skew, 25% lower sink slew, 4.5% lower power, and 28% less frontside clock wiring, and its skew spread under 10,000-sample Monte Carlo is a third of the frontside mesh's.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wajid Ali, Muhammad Hadir Khan, Dalton Gaddy, Matthew Guthaus. 2026-10-01. Design Space Exploration of Backside Clock Meshes for 2 nm GAAFET BSPDN Technology. https://arxiv.org/abs/2610.02401

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster

The electric power supply for AI data centers is now the most significant bottleneck in the race toward Artificial General Intelligence, surpassing even the constraint of AI accelerator availability. To our knowledge, this paper is the first to describe the end-to-end power management process for a hyper-scale AI datacenter; from early power planning to accommodate next-generation accelerators 6--12 months before their general availability, to tuning power settings after large scale deployment, and finally to dynamic, runtime power management for evolving workloads. We present detailed power measurements for a 150 MW datacenter hosting a cluster of 83K GB200 GPUs. We share insights from building this state-of-the-art AI cluster. We hope this work encourages practitioners across the industry to share their own experiences as well.

cs.AR↗

PEEK: Heterogeneous Parallelism for Privileged Error Detection in Safety-Critical Processors

Heterogeneous parallel error detection architecture has been widely studied for safeguarding OoO superscalar processors in safetycritical systems, as it achieves significantly lower hardware overhead compared to traditional LockStep, by exploiting the parallelism that exists in a secondary execution. However, previous works do not cover the protection of privileged-mode execution, impeding their effectiveness in real-world deployment. Moreover, naive extension to privileged-mode can cause a litany of issues, from abysmal performance due to high synchronization costs, to full deadlocks. Here, we present PEEK, the first privileged parallel error detection architecture. Based on a deep analysis of privileged execution, we redesign the verification pipeline, addressing all the bottlenecks and bugs identified in privileged-mode protection. Evaluated using various metrics on an RTL-level full system running Linux, PEEK achieves full-privilege protection on Linux with negligible performance slowdown and affordable hardware overhead. PEEK has been taped out using a 28nm process, and its source is available at https://anonymous.4open.science/r/PEEK-3000.

cs.AR↗

Divide and conquer: Scalable performance and energy in MCM GPUs

Multi-chip-module (MCM) GPUs offer a promising path to scale compute capability beyond monolithic designs by integrating multiple chiplets on a common package. However, the impact of disaggregation on performance scalability and energy consumption remains underexplored. The design space grows rapidly across dimensions such as SMs per chiplet, chiplet count, and interconnection network. The inter-chiplet network is particularly critical, as it determines whether additional compute resources translate into performance gains. This limited understanding leaves industry and research without clear guidance on the performance and energy trade-offs of MCM GPU scaling. In this work, we investigate whether distributing compute and memory capability across multiple chiplets offers a more scalable alternative to concentrating resources. We quantify their effects on performance, energy, and efficiency and examine how inter-chiplet topology influences scalability at different system sizes. Our results demonstrate that a 16 chiplet Torus configuration with 256 SMs delivers a remarkable $2.40\times$ performance improvement over a state-of-the-art MCM architecture with the same compute capability, while simultaneously reducing energy consumption by $4.45\times$. These substantial gains provide evidence that disaggregation is a first-order architectural factor and will be critical to unlocking the performance and energy-efficiency potential of next-generation GPUs.

cs.AR↗