RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of backend inference and deployment reliability arising from the separation of compilation and execution in machine learning stacks. To this end, it proposes the first Rust-based unified architecture that deeply integrates the compiler and runtime. Methodologically, a three-level intermediate representation (IR) is designed to enable transparent multi-backend scheduling and distributed scaling, supporting 14 device types and customized code generation paths alongside strict fault-handling mechanisms. The system further incorporates TCP/RDMA communication, mixed-precision training, and sparse linear algebra techniques to enhance overall capabilities. Experimental results demonstrate that the proposed system maintains 100% accuracy on MiniLM and MNIST benchmarks while surpassing mainstream frameworks such as PyTorch in throughput and significantly reducing latency.
📝 Abstract
Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runtime roles around one primitive-level, three-level intermediate representation (IR), plus a transparent dispatch contract that resolves each operator to native, common-IR, or rewritten lowering and fails compilation when legalization is not possible. The same IR targets fourteen runtime devices (cpu, metal, mlx, ane, cuda, rocm, oneapi, tpu, hexagon, gpu, vulkan, opengl, directx, webgpu) and two specialty codegen paths (Cortex-M INT8 and FPGA), ingests safetensors, GGUF, ONNX, and rten formats, supports F16/BF16/F64/C64 and quantized INT4/INT8 flows with AMP/PTQ/QAT, and scales via tensor-/pipeline-parallel collectives over TCP and RDMA transports. Beyond neural workloads, RLX also extends to scientific/physics-style domains through sparse and dense linear algebra extensions (e.g., CSR LU/CG/matvec and LAPACK- backed factorizations) and 3D Gaussian splatting operators. We evaluate RLX against PyTorch, TensorFlow, JAX, candle, burn, tch, rten, MLX, CoreML, IREE, Glow, TensorRT, and tinygrad under identical input generation and p50 measurement methodology on one host. On all-MiniLM-L6-v2, RLX-Metal is fastest at every batch (e.g., 16.6 ms at batch 32 vs. PyTorch-MPS 26.7 ms). In the MNIST training table, RLX also has the top-throughput entry (graph-fused MLP: 946,487 img/s), above NumPy+BLAS (787,349 img/s), while retaining 100% top-1 parity on reference checks (e.g., Qwen3).
Problem

Research questions and friction points this paper is trying to address.

Tensor Compiler
Distributed Runtime
Multi-Backend
Graph Compilation
Machine Learning Infrastructure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tensor Compiler
Intermediate Representation
Multi-Backend Runtime
Distributed Execution
Rust
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Eugene Hauptmann
Massachusetts Institute of Technology
N
Nataliya Kosmyna
Massachusetts Institute of Technology