Efficient Clustering with Provable Guardrails for LLM Inference at Scale

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high cost and latency of large-scale LLM inference, where existing clustering methods struggle to simultaneously ensure intra-cluster similarity, exact matching of categorical attributes, and scalability on datasets with tens of millions of samples. The authors propose a two-stage clustering algorithm: first generating initial clusters via Mini-batch K-Means, then greedily selecting representative points within α-balls in the embedding space—a procedure equivalent to applying the Johnson–Chvátal set cover heuristic—which rigorously enforces per-sample similarity and attribute constraints. This approach is the first to achieve linear scalability, guaranteed minimum intra-cluster similarity, and exact categorical attribute matching in large-scale settings, with provable quality bounds. Experiments demonstrate 10–1000× speedups over state-of-the-art methods, and successful deployment in a recommendation system serving 38 million users, reducing downstream cost and latency by up to 50×.
📝 Abstract
Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of modern foundation models. A natural fix is to cluster the inputs and call the LLM only on cluster representatives, letting other members inherit the output -- but this is only safe if each member is measurably close to its representative. Existing clustering methods do not offer such per-sample quality control at scale: none jointly guarantee a minimal within-cluster similarity, exact matching of categorical attributes, and scalability to tens of millions of samples. We propose a two-stage algorithm that generates initial clusters with Mini-batch K-Means, then greedily selects representatives within each initial cluster -- a step equivalent to the Johnson-Chvatal heuristic for Set Cover over alpha-balls in embedding space. The algorithm enforces the similarity and attribute guardrails exactly by construction, and runs in $O(nd + n^2 d/K)$ time and $O(nd + n^2/K^2)$ memory for $n$ samples, feature dimension $d$, and $K$ initial clusters -- linear in $n$ when $K$ grows proportionally with $n$. We provide benchmarks against common clustering methods on internal and public datasets: our method not only delivers per-sample guardrails but also runs 10-1000x faster and scales to data sizes where most standard methods become intractable. Deployed on 38 million customers for a persona-based recommender, the clustering method cut downstream cost and latency by 50-fold while preserving personalization and unblocked the production launch.
Problem

Research questions and friction points this paper is trying to address.

LLM inference
clustering
scalability
similarity guarantee
categorical constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

provably safe clustering
LLM inference optimization
per-sample guardrails
scalable clustering
representative selection
L
Longshaokan Wang
Amazon
W
Wai Tsang Keung
Amazon
P
Punit Ghodasara
Amazon
R
Roman Wang
Amazon
A
Ali Dashti
Amazon
Francesc Moreno-Noguer
Francesc Moreno-Noguer
Amazon Science
Computer VisionDeep Learning