arXiv Science⌕ Search

arXiv · 2610.06937

FAPO: Fanout-Aware Post-Mapping Optimization for LUT-Based FPGAs

Abstract

Reducing logic depth during FPGA technology mapping does not necessarily eliminate fanout bottlenecks in the resulting LUT networks.Replicating heavily loaded LUTs can alleviate these bottlenecks, but the added copies consume area and may increase upstream loading. We propose FAPO, a mapper-independent post-mapping optimizer that jointly refines sink assignments and LUT cuts under a final-area constraint.Replication, cut switching, and recovery are evaluated by the resulting live LUT count.We derive a necessary all-critical-path condition for delay reduction and obtain minimum-copy proposals by minimum cut within a restricted same-cut family.These proposals complement greedy refinement, with mappings selected by full-network area and timing.Experiments on EPFL benchmarks with three initial mappers achieve geometric-mean reductions of up to 14.48% in fanout-aware area--timing product (ATP) with a 5% final-area allowance.VPR evaluation shows a 7.04% reduction in routed critical-path delay relative to ABC if.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xiaoyu Hao, Keren Zhu. 2026-10-03. FAPO: Fanout-Aware Post-Mapping Optimization for LUT-Based FPGAs. https://arxiv.org/abs/2610.06937

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Predictive Software Scheduling as an Early-Warning Hint Layer for Optical Engine Thermal Drift in Heterogeneous SoIC Packaging

As semiconductor scaling approaches the A16 / 2 nm node, the integration of co-packaged optics (CPO) through TSMC's Compact Universal Photonic Engine (COUPE) architecture introduces critical thermal-optical coupling challenges. Micro-ring resonators embedded in the Photonic Integrated Circuit (PIC) layer are highly sensitive to temperature, with a wafer-level center wavelength deviation of merely +/-1.7 nm across a 300 mm wafer representing the platform's manufacturing control limit. To address this, we propose XRM-SSD V24, a physics-aware scheduling layer that models inference-load density 20-50 ms before execution and issues early-warning hints to the COUPE bias-control firmware, enabling pre-emptive thermal compensation. Simulation-based validation on a software emulation platform (physical characterization pending TSMC tape-out) over 90,000 inference steps yields a simulator-internal thermal-load correlation of R^2 = 0.9911 across a workload density range of pv24 in [0.9, 2.7] (a 3x span), with wavelength drift below 0.354 nm - equivalent to 21% of the +/-1.7 nm wafer-level wavelength control budget and 71% of the tighter +/-0.5 nm per-channel spectral specification. A full Thermal Resistance Fingerprint characterization further confirms Rth = 0.45 deg C/W, a thermal time constant tau = 80 ms, and a thermo-optic coefficient of 0.0852 nm/deg C across five discrete load states (Idle to Peak). Memory stability is reported as zero leakage in the current simulation run; long-duration soak testing to confirm sustained stability remains future work. We establish a formal domain separation between deterministic software scheduling and continuous physical thermal dynamics, ensuring physics-consistent claims suitable for peer review.

cs.AR↗

An Interleaved Parallel Dependent Quantization Hardware Architecture for H.266/VVC

While dependent quantization in H.266/VVC delivers a high compression ratio, its strong serial nature and high complexity result in poor real-time performance, making it difficult to deploy in practical scenarios. To improve the real-time performance of dependent quantization with minimal degradation to its compression performance, we propose a interleaved parallel dependent quantization hardware architecture with low BDBR loss, which achieves four-channel parallel dependent quantization by time-division multiplexing most combinational logic. This architecture adopts the proposed intra-CG context simplification scheme and the encoding scheme that skips decAbslevel during rate estimation. The proposed design incurs a BDBR loss of only 0.42% under the All Intra configuration and 0.38% under the Random Access configuration, respectively. Implemented in Verilog HDL, the dependent quantization hardware architecture achieves quantization speeds of 4K@33.2, 91.4, 331.2, and 454.1 fps at QP = 22, 27, 32, and 37, respectively, when implemented on the Xilinx XCZU19 FPGA. When implemented on ASIC using the TSMC 28nm process standard cell library, the corresponding quantization speeds reach 4K@84.0, 231.0, 837.0, and 1147.7 fps.

cs.AR↗

BenchmarkAnything: Agent-Driven Construction of Simulator-Ready Microarchitecture Benchmarks

The selection of benchmark workloads is of paramount importance in computer architecture, as it establishes the yardstick against which architectural innovations are measured and guided. Yet for decades, the SPEC benchmark suites, comprising merely tens of workloads, have been the de facto standard in academic architectural research, where they are frequently treated as a principal evaluation and optimization target. When a suite this small is relied upon so heavily, it risks architectural overfitting; as our research and prior studies demonstrate, an overly narrow focus can mislead design decisions by overvaluing certain innovations, producing cores that excel on SPEC benchmarks yet underperform on broader, realistic workloads. To mitigate this overfitting, adopting a large, comprehensive benchmark suite is the natural solution. However, the immense engineering effort required to strip software into the clean, interference-free binary executables demanded by simulators often makes this highly impractical. In this work, we demonstrate that AI agents provide an elegant solution to this challenge. Rather than manually curating yet another static benchmark suite, we introduce an agent-driven workflow capable of autonomously transforming arbitrary open-source repositories into simulator-ready executables. This automated approach makes workload collection highly scalable, allowing us to rapidly harvest hundreds of diverse applications from public repositories into our benchmark suite. Through a comparative analysis of our agent-generated suite against SPEC, we show that it not only achieves higher-fidelity performance assessments but also uncovers novel architectural insights that traditional, static suites fail to expose.

cs.AR↗