arXiv Science⌕ Search

arXiv · 2609.36763

SCORAS-MoE: Joint Compression and Resource-Adaptive Deployment of MoE-VLMs in LEO Satellite Networks

Abstract

Deploying large vision-language models (VLMs) onboard satellites enables onboard data processing and reduces raw data downlink. However, onboard inference faces two resource challenges. Limited onboard memory and energy require model compression and distributed deployment. Dynamic resource availability requires fast deployment decisions as illumination, battery levels, and communication conditions change. We present SCORAS-MoE, a joint compression and deployment framework for mixture-of-experts (MoE) VLMs in low Earth orbit (LEO) satellite networks. To address limited resources, SCORAS-MoE measures the perturbation of the routed MoE output caused by low-rank approximation, assigns higher ranks to more sensitive experts, and distributes compressed model shards across satellites for cooperative inference. The compressed models yield profiles of measured accuracy and inference energy. To adapt to dynamic resources, the online scheduler selects profile compositions and shard placements in each slot. For each candidate composition, it reduces placement to a minimum-cost assignment problem solved by the Hungarian algorithm, while enumerating the compositions yields the optimal deployment for the current-slot objective. Experiments on Qwen3-VL-30B-A3B-Instruct show that allocating ranks based on output perturbation is particularly effective under aggressive compression, with an absolute gain of $3.7\%$ in mean accuracy over uniform rank allocation when expert projections retain $30\%$ of their original parameters. The fixed-profile scheduler achieves higher throughput with fewer service switches and lower battery impact than the evaluated proximal policy optimization (PPO) and evolutionary baselines, with respective speedups of $8.7\times$ and $183.5\times$. Adaptive profile selection further improves the balance between service quality and energy use.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tong Quan, Yuanlong Wan, Huasen He, Yunpeng Hou, Shuangwu Chen, Xiaofeng Jiang, Jian Yang. 2026-09-29. SCORAS-MoE: Joint Compression and Resource-Adaptive Deployment of MoE-VLMs in LEO Satellite Networks. https://arxiv.org/abs/2609.36763

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

GATE: GPU-Accelerated Traffic Engineering for the WAN

Traffic engineering (TE) has become a crucial tool for enforcing routing policy and maintaining operational efficiency in large networks. Existing TE solutions pick an objective function to optimize, aiming to balance (i) allocating traffic optimally with (ii) reacting quickly to demand changes and disruption events. However, as the scale of networks grows, the runtime of the existing optimal solution becomes infeasibly large. The alternative - approximate solvers - result in costly inefficiencies. We present GPU-Accelerated Traffic Engineering (GATE), which achieves the best of both worlds: enabling fast TE runtimes through a highly-parallelizable GPU-compatible decomposition, while iteratively converging to the provably optimal solution. GATE unlocks a unique set of desirable properties: it becomes increasingly parallelizable with network size, supports a wide spectrum of fairness objectives, and offers theoretically guaranteed convergence to the optimal solution and near-optimal convergence within a bounded time. We evaluate GATE on production traces from two large cloud WANs, and show that GATE achieves near-optimal solutions 4-10x faster than state-of-the-art.

cs.NI↗

Throughput-Optimized Networks at Scale

Data center network design plays a critical role in AI training by supporting scaling to thousands of accelerators. An open problem, designing a near-optimal throughput-oriented network-topology, routing, and collectives-has not been achieved at scale and with broad applicability to physical or implementation constraints. We address this problem with a compelling use-case, Google's TPU v4-8t supercomputer where the topology may be reconfigured to achieve higher All-to-All throughput, supporting large, parallelized AI training. We show that the existing TPU networks leave terabytes per second of throughput on the table and we fill that gap. This paper presents Throughput-Optimized Networks at Scale (TONS), an automated network synthesis framework that meets the high-throughput demands of modern computing. TONS formulates topology synthesis as a linear optimization problem that maximizes a throughput-centric proxy metric, using theory and heuristics to scale to thousands of nodes while producing state-of-the-art network topology performance metrics. We further introduce a state-of-the-art deadlock-free routing scheme compatible with limited virtual channels and optical switch faults, enabling the synthesized topologies to realize their predicted throughput gains in simulation. Evaluating uniform random and All-to-All traffic, TONS networks have a geometric-mean speedups of 2.1x and 1.6x over the best TPU v4-8t torus variants.

cs.NI↗

Cell-Free Massive MIMO Under Mobility: A Fairness-Differentiated Handover Scheme

While cell-free massive MIMO (CF-mMIMO) offers high and uniform network-wide throughput in static networks, its performance in mobile networks is not yet fully addressed. In this paper, we evaluate the throughput performance of urban mobile CF-mMIMO networks under a comprehensive throughput model and show that it suffers from large performance degradation due to the combined effect of channel aging and handover overheads. To restore the uniformly good performance of CF-mMIMO under mobility, we formulate a novel optimization problem to maximize the nett throughput that considers both channel aging and handover cost. We derive a near-optimal solution nearOpt for our transformed and relaxed optimization problem with Newton's method. We then design a heuristic handover algorithm, FairDiff, to differentiate prioritized and optional handovers using a policy threshold based on Jain's fairness index, in order to achieve uniform throughput over the network. Our extensive evaluation of the mobile throughput performance of our handover schemes in realistic urban mobile networks shows that, unlike the existing literature benchmarks that obtain very low throughput under mobility, our FairDiff scheme consistently achieves the near-optimal throughput comparable to nearOpt and highest network-wide throughput with the lowest computational complexity among all considered schemes. We thus for the first time propose a handover scheme that delivers the promise of uniformly good throughput for mobile CF-mMIMO, making it a feasible architecture for practical mobile networks.

cs.NI↗