arXiv · 2610.02658
Annotation-Driven Migration of CUDA Programs to Tenstorrent Blackhole
Abstract
Tenstorrent Blackhole combines distributed local memories, explicit inter-core communication, and Tensix cores decoupling data movement from computation. CUDA offers a substantial HPC software base but leaves physical data placement and scheduling largely implicit. Migrating CUDA HPC kernels to Blackhole requires spatial mapping across cores and per-core coordination of compute and data-movement kernels. We present an MLIR-based compiler deriving data placement, inter-core communication, and tile computation from a statically shaped affine CUDA subset. Declarative annotations express choices not fixed by the source, including compute/DM operation placement and streaming granularity. We evaluate it on Gaussian elimination, five-point stencil, and symmetric rank-k update. Relative to our compiler's default realizations, the best measured configurations achieve a 4.2x speedup for BF16 Gaussian and 2.1-2.2x for FP32 Gaussian and Jacobi. A choice's performance impact can reverse with surrounding policies, precision, and loop schedule, motivating comparison of alternative per-core realizations rather than independent policy selection.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ayumi Ohno, Shinya Takamaeda-Yamazaki. 2026-10-02. Annotation-Driven Migration of CUDA Programs to Tenstorrent Blackhole. https://arxiv.org/abs/2610.02658
Cite the original work for its findings. Save a collection to share your selection of sources.