arXiv · 2404.01908
Optimizing Offload Performance in Heterogeneous MPSoCs
Abstract
Heterogeneous multi-core architectures combine a few "host" cores, optimized for single-thread performance, with many small energy-efficient "accelerator" cores for data-parallel processing, on a single chip. Offloading a computation to the many-core acceleration fabric introduces a communication and synchronization cost which reduces the speedup attainable on the accelerator, particularly for small and fine-grained parallel tasks. We demonstrate that by co-designing the hardware and offload routines, we can increase the speedup of an offloaded DAXPY kernel by as much as 47.9%. Furthermore, we show that it is possible to accurately model the runtime of an offloaded application, accounting for the offload overheads, with as low as 1% MAPE error, enabling optimal offload decisions under offload execution time constraints.
Explore related subjects
Keep this discovery
Luca Colagrande, Luca Benini. 2024-04-02. Optimizing Offload Performance in Heterogeneous MPSoCs. https://doi.org/10.23919/date58400.2024.10546670
Cite the original work for its findings. Save a collection to share your selection of sources.