arXiv Science⌕ Search

arXiv · 2609.37916

RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust

Abstract

Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runtime roles around one primitive-level, three-level intermediate representation (IR), plus a transparent dispatch contract that resolves each operator to native, common-IR, or rewritten lowering and fails compilation when legalization is not possible. The same IR targets fourteen runtime devices (cpu, metal, mlx, ane, cuda, rocm, oneapi, tpu, hexagon, gpu, vulkan, opengl, directx, webgpu) and two specialty codegen paths (Cortex-M INT8 and FPGA), ingests safetensors, GGUF, ONNX, and rten formats, supports F16/BF16/F64/C64 and quantized INT4/INT8 flows with AMP/PTQ/QAT, and scales via tensor-/pipeline-parallel collectives over TCP and RDMA transports. Beyond neural workloads, RLX also extends to scientific/physics-style domains through sparse and dense linear algebra extensions (e.g., CSR LU/CG/matvec and LAPACK- backed factorizations) and 3D Gaussian splatting operators. We evaluate RLX against PyTorch, TensorFlow, JAX, candle, burn, tch, rten, MLX, CoreML, IREE, Glow, TensorRT, and tinygrad under identical input generation and p50 measurement methodology on one host. On all-MiniLM-L6-v2, RLX-Metal is fastest at every batch (e.g., 16.6 ms at batch 32 vs. PyTorch-MPS 26.7 ms). In the MNIST training table, RLX also has the top-throughput entry (graph-fused MLP: 946,487 img/s), above NumPy+BLAS (787,349 img/s), while retaining 100% top-1 parity on reference checks (e.g., Qwen3).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Eugene Hauptmann, Nataliya Kosmyna. 2026-09-29. RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust. https://arxiv.org/abs/2609.37916

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Deterministic Self-Stabilizing BFS Construction in Constant Space

In this paper, we resolve a long-standing question in self-stabilization by demonstrating that it is indeed possible to construct a spanning tree in a semi-uniform network using constant memory per node. We introduce a self-stabilizing synchronous algorithm that builds a breadth-first search (BFS) spanning tree with only $O(1)$ bits of memory per node, converging in $2^ε$ time units, where $ε$ denotes the eccentricity of the distinguish node. Crucially, our approach operates without any prior knowledge of global network parameters such as maximum degree, diameter, or total node count. In contrast to traditional self-stabilizing methods, such as pointer-to-neighbor communication or distance-to-root computation, that are unsuitable under strict memory constraints, our solution employs an innovative constant-space token dissemination mechanism. This mechanism effectively eliminates cycles and rectifies deviations in the BFS structure, ensuring both correctness and memory efficiency. The proposed algorithm not only meets the stringent requirements of memory-constrained distributed systems but also opens new avenues for research in self-stabilizing protocols under severe resource limitations.

cs.DC↗

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

General Matrix Multiplication (GEMM) is the cornerstone of HPC workloads and Deep Learning. State-of-the-art (SOTA) vendor libraries tune tensor layouts, parallelization schemes and cache blocking to minimize data movement across the memory hierarchy and maximize throughput. However, optimal settings for these parameters depend on the target platform and matrix shapes, making exhaustive tuning infeasible. In this work, we address this cumbersome scheduling search and tuning using space-filling curves (SFC). We partition the matrix multiplication using advancements in SFC, and obtain platform-oblivious and shape-oblivious matrix multiplication schemes with a high degree of data locality. We extend the SFC-based work partitioning to implement Communication-Avoiding (CA) algorithms with replication techniques in a seamless fashion. The resulting SFC-CA GEMM achieves provable asymptotic communication optimality for both square and rectangular matrix regimes. Across four x86 and Arm platforms, SFC-CA GEMM outperforms vendor libraries by up to 5.5$\times$ per shape and 1.8$\times$ in weighted harmonic mean (WHM) throughput. Last, we show the impact of our work on two real-world applications by leveraging our SFC-CA GEMM as a compute backend: i) prefill of LLM inference with speedups up to 1.85$\times$ over SOTA inference runtimes, and ii) distributed-memory matrix multiplication with speedups up to 2.3$\times$ over the SOTA distributed-memory GEMM framework with vendor-optimized compute backend.

cs.DC↗

MoEless: Efficient MoE LLM Serving with Serverless Experts

Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.

cs.DC↗