arXiv Science⌕ Search

arXiv · 2609.34657

Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch

Abstract

Modern deep learning (DL) workloads are limited by data movement, and processing-in-memory (PIM) targets this bottleneck by placing compute units near the memory. However, PyTorch and other DL frameworks lack compiler support for making this decision on the code they lower: existing offloading frameworks target hand-written C/C++ programs, while those that address DL fix the candidate set to a list of operator types before lowering. We present Torch-PIM, a compiler framework that uses profile-guided optimization (PGO) to decide host-versus-PIM placement over the loop nests that progressive lowering materializes. Every parallel loop nest the pipeline emits enters the candidate space, and each is assessed in two stages: the amount of work it carries, and its memory boundedness. Every quantity the assessment consumes is profiled on the host or obtained from the multi-level intermediate representation (MLIR) of the code. Across PIM configurations of 32 to 128 cores, Torch-PIM's offloading decisions yield speedups of up to 8.6x on tensor operators, 2.9x on MLP, 4.4x on Attention, 5.1x on GPT-J-6B, and 3.6x on LLaMA-7B over CPU-only execution.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Heeeon Lee, Hyunwoo Nam, Junyong Heo, Hyunmo Sung, Jay Hwan Lee, Yeonsoo Kim, Seongho Jeong, Shinhyung Yang, Bernd Burgstaller. 2026-09-28. Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch. https://arxiv.org/abs/2609.34657

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Hardware-Attributed Operator Profiling for PyTorch

Framework profilers expose operator timing without hardware counters; GPU profilers expose hardware counters without operator attribution. Bridging this gap manually is error-prone and does not scale. We present Operator Profiler, a hardware attribution pipeline that automatically links hardware metrics to PyTorch operators via three complementary attribution paths: PyTorch profiler CUPTI correlation, NVTX temporal enclosure with per-stream interval trees, and Inductor fusion- map enrichment from debug artifacts. NVIDIA Nsight Compute (ncu) hardware counters are matched to NVIDIA Nsight Systems (nsys) kernel records via invocation-order matching, avoiding timestamp joins across incompatible clock domains. A curated 20-counter metric set with duration-weighted aggregation covers all hardware bottleneck axes, layer deduplication reduces ncu replay time by a factor of N/K for models with N layers across K unique structural classes, and GPU clock locking controls the kernel-duration aggregates used for operator-level comparison. On an NVIDIA RTX PRO 6000 Blackwell, Operator Profiler attributes 95-100% of kernel runtime for compiled workloads (GPT-2, SDPA Attention); black-box library backends such as cuDNN RNN are correctly surfaced as greater than 85% unattributed rather than silently dropped. Applied to profile-guided FX graph optimization, attributed profiles yield 1.76x-2.24x profiled-kernel-time speedups on the two compiled optimization case studies; a third LSTM diagnostic case identifies cuDNN re-dispatch as a structural fix rather than an FX graph rewrite.

cs.AR↗

Lossless Compression of Lookup Tables for Hardware Applications

Large lookup tables are widely used in hardware to store constant-valued arrays for applications ranging from elementary mathematical operations, such as constant-coefficient multiplication and nonlinear function evaluation, to emerging machine learning models, including table-based neural networks (NNs) and Kolmogorov-Arnold networks (KANs). However, storing extensive tables of constant values can lead to excessive hardware costs in resource-constrained edge devices such as FPGAs. In this paper, we propose CompressedLUT, a lossless compression scheme and its decoder hardware architecture for the efficient storage and retrieval of arbitrary data in hardware. Our method combines decomposition, self-similarities, higher-bit compression, and multilevel compression techniques to maximize table size savings without accuracy loss. Its hardware decoder primarily uses addition, arithmetic right shift, and several small lookup tables, ensuring low area and high throughput. We evaluated CompressedLUT on FPGAs by implementing multiple nonlinear functions, constant-coefficient multipliers (CCMs), and KANs at 12-bit resolution. CompressedLUT is available as an open-source tool.

cs.AR↗

Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis

Advancements in AI have greatly enhanced the medical imaging process, making it quicker to diagnose patients. However, very few have investigated the optimization of a multi-model system with hardware acceleration. As specialized edge devices emerge, the efficient use of their accelerators is becoming increasingly crucial. This paper proposes a hardware-accelerated method for simultaneous reconstruction and diagnosis of \ac{MRI} from \ac{CT} images. Real-time performance of achieving a throughput of nearly 150 frames per second was achieved by leveraging hardware engines available in modern NVIDIA edge GPU, along with scheduling techniques. This includes the GPU and the \ac{DLA} available in both Jetson AGX Xavier and Jetson AGX Orin, which were considered in this paper. The hardware allocation of different layers of the multiple AI models was done in such a way that the ideal time between the hardware engines is reduced. In addition, the AI models corresponding to the \ac{GAN} model were fine-tuned in such a way that no fallback execution into the GPU engine is required without compromising accuracy. Indeed, the accuracy corresponding to the fine-tuned edge GPU-aware AI models exhibited an accuracy enhancement of 5\%. A further hardware allocation of two fine-tuned GPU-aware GAN models proves they can double the performance over the original model, leveraging adequate partitioning on the NVIDIA Jetson AGX Xavier and Orin devices. The results prove the effectiveness of employing hardware-aware models in parallel for medical image analysis and diagnosis.

cs.AR↗