arXiv Science⌕ Search

arXiv · 2609.34045

Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines

Abstract

Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding membership removes what they depend on: a swarm holds each part of the model on several peers and routes around a slow one. A bounded session must use every device it admits. Its pipeline advances at the pace of whichever device received a share it cannot serve quickly, so the division has to be right before serving begins. We propose Kafila, whose protocol assembles a ring from behind NATs, preferring direct paths and relaying where traversal fails, while its planner measures each device's memory bandwidth, capacity and reachability, divides the model exactly for a fixed ring order, and places the head, which holds the embedding and output projection, together with that division rather than beforehand. On machines with different capabilities across three fleets, from a shared LAN to five devices spanning two continents, Kafila shortens the slowest pipeline stage by up to $5.2\times$ against the even split of pipeline parallelism, as in GPipe, and up to $3\times$ against the memory-proportional split of personal-device inference, as in exo, keeps 75 to 87 per cent of the committed hardware doing work where those divisions fall below half, and serves a model no uniform split can place on the fleet at all. What that is worth to a user depends on how much of a token is computation rather than network. Where the members share a network the same division returns $1.56\times$ the throughput of a uniform split and $1.25\times$ of a memory-proportional one, and under four concurrent users that lead compounds to $3.2\times$ rather than fading, each user served at almost the rate of one.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Murtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya. 2026-09-28. Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines. https://arxiv.org/abs/2609.34045

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Recolorable Graph Exploration by an Oblivious Agent with Fewer Colors

Recently, Böckenhauer, Frei, Unger, and Wehner (SIROCCO 2023) introduced a novel variant of the graph exploration problem in which a single memoryless agent must visit all nodes of an unknown, undirected, and connected graph before returning to its starting node. Unlike the standard model for mobile agents, edges are not labeled with port numbers. Instead, the agent can color its current node and observe the color of each neighboring node. To move, it specifies a target color and then moves to an adversarially chosen neighbor of that color. Böckenhauer~et al.~analyzed the minimum number of colors required for successful exploration and proposed an elegant algorithm that enables the agent to explore an arbitrary graph using only eight colors. In this paper, we present a novel graph exploration algorithm that requires only six colors. Furthermore, we prove that five colors are sufficient if we consider only a restricted class of graphs, which we call the $φ$-free graphs, a class that includes every graph with maximum degree at most three and every cactus.

cs.DC↗

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines

Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic workflows onto a GPU cluster. Scepsy exploits the insight that, while the end-to-end latency of an agentic workflow is unpredictable, each LLM's fraction of execution time is comparatively stable across requests. Scepsy profiles each LLM under different parallelism degrees and combines the profiles with these fractions into an Aggregate LLM Pipeline, a lightweight throughput and latency predictor for allocations. To minimize latency at a target throughput, Scepsy uses the Aggregate LLM Pipeline to search over fractional GPU shares, tensor parallelism degrees, and replica counts. A hierarchical heuristic then places the chosen allocation onto the cluster, minimizing fragmentation and respecting network topology. On realistic agentic workflows, Scepsy achieves up to 2.5x higher throughput before saturation and 1.0-3.3x lower latency than systems that optimize LLMs independently or rely on user-specified allocations.

cs.DC↗

The World's Fastest Matching Engine Algorithm

We drove 247 matching engines through one C ABI harness on one identical workload: every open-source FIFO implementation we could find, deduplicated, and our own, on the same gate. The workload doubles as a byte-identical correctness oracle, 1,000,000,000+ order messages per engine, replayed against an independent-engine consensus. Only 47 are correct as shipped; we filed 181 GitHub issues upstream, 28 already fixed by their maintainers, none declined. Our engine leads the 160 that reproduce the consensus by ~95 M/s (12.6x the second best) on worst-case throughput. One core sustains 103.4 million order messages per second (122.09 million on AMD EPYC processors) worst-case, and in the engine's production configuration its wire pipeline keeps single-book P99 host-path latency, OUCH parsing and OUCH/ITCH encoding included, under a microsecond through 82 M msgs/s; a 96-core server (~$1,630/month, 3-year reserved) reaches ~1.3 billion/s across over 10,000 symbols, for scale, over 45x the CTA consolidated quote feed's provisioned capacity. The lead is structural: the 73 engines written inside the trading industry sit under the same 8.19 M/s ceiling as the rest of the field, and it is one of them that sets it.

cs.DC↗