Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the pipeline bottleneck caused by straggler nodes when serving large language models across heterogeneous devices within trusted private clusters. We propose a NAT-traversal-based ring serving protocol that leverages heterogeneous hardware performance modeling and bandwidth capacity measurement to design a statically optimal partitioning strategy for a constrained member set. By jointly optimizing embedding layers, output heads, and model blocks, our approach overcomes the limitations of conventional uniform or proportional partitioning, enabling efficient inference scheduling under a fixed execution order. Experimental results demonstrate that the proposed method reduces the latency of the slowest stage by 5.2×, achieves 75%–87% hardware utilization, and improves LAN throughput by 1.56×, thereby significantly enhancing multi-user concurrent serving performance.
📝 Abstract
Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding membership removes what they depend on: a swarm holds each part of the model on several peers and routes around a slow one. A bounded session must use every device it admits. Its pipeline advances at the pace of whichever device received a share it cannot serve quickly, so the division has to be right before serving begins. We propose Kafila, whose protocol assembles a ring from behind NATs, preferring direct paths and relaying where traversal fails, while its planner measures each device's memory bandwidth, capacity and reachability, divides the model exactly for a fixed ring order, and places the head, which holds the embedding and output projection, together with that division rather than beforehand. On machines with different capabilities across three fleets, from a shared LAN to five devices spanning two continents, Kafila shortens the slowest pipeline stage by up to $5.2\times$ against the even split of pipeline parallelism, as in GPipe, and up to $3\times$ against the memory-proportional split of personal-device inference, as in exo, keeps 75 to 87 per cent of the committed hardware doing work where those divisions fall below half, and serves a model no uniform split can place on the fleet at all. What that is worth to a user depends on how much of a token is computation rather than network. Where the members share a network the same division returns $1.56\times$ the throughput of a uniform split and $1.25\times$ of a memory-proportional one, and under four concurrent users that lead compounds to $3.2\times$ rather than fading, each user served at almost the rate of one.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model Serving
Heterogeneous Commodity Machines
Trusted Device Cluster
Pipeline Parallelism
Distributed Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous Computing
Pipeline Parallelism
LLM Serving
NAT Traversal
Model Partitioning