arXiv · 2610.02288
AdaptViT: Runtime-Adaptive Vision Transformer Deployment on Custom RISC-V
Abstract
Deploying Vision Transformers (ViTs) on low-power edge devices is challenging due to high computational demands. Conventional pruning frameworks require a separate compiled binary for each sparsity level, increasing storage overhead and limiting runtime adaptability. This paper presents an end-to-end deployment pipeline that transforms pretrained ViTs into a single runtime-configurable binary, enabling dynamic compute-budget switching on embedded CPUs. This is achieved by restructuring generated C kernels with modified loop bounds and binary-mask control logic, allowing execution to switch across discrete sparsity levels via compact external configuration files. Compared to multi-binary deployment, the proposed runtime-adaptive approach reduces on-device storage by up to 4.86x, requiring only 163 MB for ViT-Base instead of nearly 800 MB. To maximize pruning efficiency, we introduce a hardware-aligned block pruning strategy for Multi-Layer Perceptron (MLP) layers. In addition, a custom ISA extension is proposed to exploit input-reuse patterns in linear projection kernels. On a Synopsys TRV32P3FX RISC-V processor, the full system achieves up to 2.8x speedup at 65% MLP and 50% attention-head pruning for ViT-Base. The ISA extension alone provides a 1.56x speedup and 33% lower inference energy, with a 24.7% area overhead in a TSMC 28 nm implementation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Vishnu PS, Ajay Kumar M, Yike Li, Robert Bogdan Staszewski, Deepu John. 2026-10-01. AdaptViT: Runtime-Adaptive Vision Transformer Deployment on Custom RISC-V. https://arxiv.org/abs/2610.02288
Cite the original work for its findings. Save a collection to share your selection of sources.