arXiv · 2608.10803
Adaptive Matrix Multiplication for Dynamic Shapes on Ascend NPUs
Abstract
Matrix Multiplication (MatMul) faces a "generalization crisis" driven by highly dynamic tensor shapes. This crisis is particularly acute on Ascend NPUs, where explicitly controlled architectures and strict physical constraints render existing GPU-centric optimizations ineffective. To resolve this, we propose AdaptCore, an adaptive framework for universally high-performance MatMul on Ascend NPUs. AdaptCore systematically decouples operator optimization into spatial tiling and instruction orchestration. It first maps dynamic shapes into a hardware-aware 2D tiling taxonomy to balance on-chip capacity limits and multi-core parallelism. Furthermore, it integrates a composable optimization library with a deterministic analytical performance model. By mathematically evaluating hardware state mutations, AdaptCore proactively selects and caches optimal implementations, enabling O(1) overhead runtime dispatching. Evaluations demonstrate that AdaptCore delivers a remarkable 1.85x mean speedup across 80,000 input shapes, and achieves up to a 1.48x acceleration in representative end-to-end models over the highly-tuned native vendor library (ACLNN).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuhang Zhou, Jiang Peng, Qianyu Jiang, Zhibin Wang, Xinghui Tian, Jianwei Zhou, Songxiang Zhu, Jingyi Zhang, Junsong Wang, Chen Tian. 2026-08-11. Adaptive Matrix Multiplication for Dynamic Shapes on Ascend NPUs. https://arxiv.org/abs/2608.10803
Cite the original work for its findings. Save a collection to share your selection of sources.