arXiv · 2610.05123
HiNa-MoE: High-Performance, Non-Intrusive MoE Inference on CPUs with Matrix Engines
Abstract
Mixture-of-Experts (MoE) inference is increasingly deployed in local and on-premise environments, where expert parameters often exceed GPU memory capacity. In latency-sensitive, low-concurrency settings, repeatedly staging routed-expert weights from CPU memory to the GPU can be prohibitive, leaving routed-expert feed-forward networks (FFNs) on the critical path of multi-socket CPUs. Existing CPU accelerations often rely on intrusive, hardware- or topology-specific requirements, such as AMX-specific weight layouts or manual NUMA-aware placement. These requirements reduce portability and complicate integration with standard CPU-GPU offloading pipelines. We present HiNa-MoE, a high-performance, non-intrusive operator library for MoE inference on CPUs with Intel AMX. HiNa-MoE (1) exploits AMX with an optimized micro-kernel that keeps expert weights in standard layouts and instead fuses lightweight layout transforms into token gathering and stores; (2) applies NUMA-aware task partitioning under a simple page-interleaved policy without modifying the framework allocator; and (3) converts decode-phase memory matrix-vector operations into small matrix-matrix execution to utilize AMX. Across multiple MoE models, HiNa-MoE achieves up to 3.37x speedup for FFN kernels and up to 2.09x end-to-end inference speedup over state-of-the-art baselines, while remaining plug-and-play with existing frameworks and deployment workflows.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Weiling Yang, Junwen Zhang, Dezun Dong, Jianbin Fang, Enda Yu, Zhe Bai, Xiaopeng Deng. 2026-10-04. HiNa-MoE: High-Performance, Non-Intrusive MoE Inference on CPUs with Matrix Engines. https://arxiv.org/abs/2610.05123
Cite the original work for its findings. Save a collection to share your selection of sources.