AirMoE: Realizing Over-the-Air Distributed Mixture-of-Experts Inference at the Wireless Edge
Mixture-of-experts (MoE) architectures enable efficient large language model (LLM) inference at the wireless edge through sparse activation. The wireless distributed MoE (WIDE) architecture addresses edge-resource constraints by distributing experts across devices coordinated by an edge server. However, WIDE suffers from repeated uplink transmissions of high-dimensional expert outputs via orthogonal multiple access. To overcome this bottleneck, we propose AirMoE, an over-the-air computing (AirComp)-enabled framework for simultaneous expert-output aggregation via wireless waveform superposition. Integrating AirComp into MoE inference introduces three challenges: fast-varying aggregation weights, layer-dependent error sensitivity, and channel-aware expert placement. To address these challenges, we construct an inference-aware AirMoE error metric to quantify aggregation distortion effects on end-to-end (E2E) inference accuracy via perturbation-based layer-sensitivity calibration. We then formulate a joint optimization problem to minimize this error and decompose it, without loss of optimality, into a two-timescale framework. At the fast timescale, we derive a globally optimal threshold-based power-control policy that separates devices into coefficient-aligned and full-power groups. At the slow timescale, we develop an activation- and channel-aware expert placement strategy that assigns more important experts to devices with lower channel-power cost. Extensive experiments demonstrate that AirMoE outperforms representative baselines in E2E inference accuracy, especially under strong device heterogeneity.