arXiv · 2610.09156
Compiling Semi-Ring Dynamic Programming to Tier-Aware 3D-DRAM Processing-in-Memory
Abstract
Processing-in-Memory (PIM) on monolithic 3D (M3D) DRAM is a promising answer to the memory wall for data-intensive dynamic programming (DP), yet extracting its performance today demands hand-written kernels: the programmer must pick a tile size, place data across non-uniform-latency memory tiers, partition work between heterogeneous processing units, and insert the right broadcasts and queues by hand. We present GenMLIR, an MLIR compiler that automates this mapping for tier-aware 3D PIM. GenMLIR encodes the semi-ring generalized grid update, the algebraic form shared by all-pairs shortest path (APSP) and sequence alignment, as a first-class IR abstraction, and lowers it through four GenDRAM-aware pass groups that derive blocked tiling, tier-aware placement, tile-to-PU assignment, and explicit communication. We also characterize precisely which DP recurrences the abstraction admits and which it does not. On the GenDRAM architecture, GenMLIR-generated code runs up to 5.8 (APSP) and 16.7 (alignment) faster than a lowering without compiler support, and 1.2-1.4 faster than a standard affine-tiling PIM compiler we also implement, reaching 90-100% of an achievable-performance bound. It expresses these workloads in 4-9 lines instead of 300-500 and compiles in a negligible fraction of runtime.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mahbod Afarin, Tsung-Han Lu, Tajana Rosing. 2026-10-06. Compiling Semi-Ring Dynamic Programming to Tier-Aware 3D-DRAM Processing-in-Memory. https://arxiv.org/abs/2610.09156
Cite the original work for its findings. Save a collection to share your selection of sources.