arXiv Science⌕ Search

arXiv · 2610.03415

RailWave: Adaptive Spatial and Temporal Scheduling for Expert-Parallel Communication

Abstract

Irregular All-to-All communication is a major bottleneck in expert-parallel Mixture-of-Experts (MoE) models. Even with fixed expert routing and placement, uneven utilization of parallel network Rails and incast can limit communication performance. We present RailWave, a phase-adaptive communication layer built on DeepEP that addresses these bottlenecks below the routing layer through spatial and temporal traffic shaping. RailBalance redistributes source traffic across eligible Rails using source-local information, while a reusable, topology-derived permutation schedule limits concurrent senders per receiver without rebuilding demand-dependent schedules for each communication phase. A lightweight calibrated selector chooses an execution path according to each phase's traffic characteristics and offline profiling results. On training-derived communication workloads from the 106B GLM-4.5-Air model, RailWave delivers up to 5.84x speedup on H800 and 4.36x on H20 over Native. Code is available at https://github.com/CyberSecurityErial/RailWave-EP.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chutian Wang, Wenhao He, Jingmin Zhu, Qingyu Yin, Heng Xu, Xiuyu Li. 2026-10-02. RailWave: Adaptive Spatial and Temporal Scheduling for Expert-Parallel Communication. https://arxiv.org/abs/2610.03415

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Grassroots Bonds: Financing by the People, for the People

Grassroots currencies turn mutual trust into liquidity: a grassroots coin is a unit of its issuer's debt, backed by the issuer's goods and services, which the issuer must redeem, 1-for-1, against any coin they hold; liquidity arises from mutual credit lines, formed by the voluntary exchange of coins among persons who trust each other. As coins are redeemable 1-for-1, the exchange must be 1-for-1 as well, lest prompt redemption after it leave one party with undue profit. Thus, grassroots coins are incongruent with interest-bearing credit. Here, we extend grassroots currencies to include also grassroots bonds, units of their issuer's debt due at a later date. Upon maturity, the bearer of the bond may redeem it against a coin of the issuer, so liquid coins can be lent against interest-bearing bonds. We specify bonds by adding Date and Escrow clauses to the social contract of grassroots currencies: the contract enforces the redemption of a bond when it is mature according to the date stated by the issuer, who undertakes to keep it current. We show that the voluntary swap of coins and bonds can realise the basic financial instruments: loans, sale of debt, and forward contracts, and, with an escrow agent, loans with payment schedules, options, collateral, guarantees, insurance, credit default swaps, letters of credit, and credit lines. We extend the liquidity ratios of corporate finance to bonds, with maturity as asset class, and prove that a community clears its debts by redemptions exactly when no member owes more than they hold, and that it does so without coordination once dates advance and persons act on their rights. Grassroots currencies that include coins and bonds are implemented in GLP, a concurrent logic programming language running on Dart, as a program derived from the contract and demonstrated by a village market of six agents and an escrow agent.

cs.DC↗

Characterization-Guided GPU Fault Resilience in NVIDIA MPS

NVIDIA Multi-Process Service (MPS) enables fine-grained GPU sharing by allowing multiple processes to execute concurrently on the same GPU, making it an important mechanism for improving GPU utilization. However, MPS has weak fault resilience: a fault in one process can terminate all co-running processes, limiting its adoption in resilience-critical settings such as multi-tenant GPU clusters. In this work, we design fault-resilient MPS to solve this problem. Our design is guided by insights from a systematic characterization of GPU faults and a deep analysis of their end-to-end processing pipeline. Based on these insights, we design two complementary mechanisms. First, we design a fault isolation mechanism for the dominant memory-related faults that can be fully isolated while preserving process-level fail-stop semantics by software intervention in the open GPU driver kernel module. For other faults whose process is within proprietary software, we design a fast-recovery substrate that combines virtual-memory-based GPU-resident state sharing with pre-initialized standbys. Our evaluation across GPUs and workloads demonstrates effective fault isolation and fast recovery with minimal overhead: in an end-to-end case study, isolation incurs no visible outage, while recovery restores pre-fault throughput in 355\,ms.

cs.DC↗

ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum

Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.

cs.DC↗