RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决大规模使用LLM进行聚类的问题,提出RAILS方法,通过检索增强和增量处理提升效率与准确性。
📝 Abstract
Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real workloads require. We present RAILS, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple loop over a growing label pool and scales through document batching with bounded concurrency. On six public benchmarks RAILS exceeds the strongest prior LLM-clustering method on average, lifting accuracy from 51.2% to 59.3%, NMI from 67.2% to 74.8%, and ARI from 45.4% to 54.7%. We further report production-deployment evidence from a SaaS ticket-topic-discovery pipeline, where RAILS has replaced a traditional HDBSCAN stage with higher clustering quality, transparent prompt-driven control, and stateful incremental operation.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model (LLM)
clustering
production scale
throughput
label space
Innovation

Methods, ideas, or system contributions that make the work stand out.

retrieval-augmented
incremental LLM clustering
document batching
bounded concurrency
production-scale
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Armin Oliya
Zendesk
A
Aleksandra Sawczuk
Zendesk
R
Radosław Białobrzeski
Zendesk