🤖 AI Summary
This work addresses the inefficiency in traditional fixed-probe IVF approaches, which ignore application-level similarity thresholds during training data deduplication, leading to redundant computation. The paper proposes SieveIVF, the first dynamic search mechanism that integrates similarity thresholds directly into the IVF execution process, enabling early termination without modifying the index structure or top-k interface—specifically halting when no qualified candidates are found in W consecutive probes. By employing a dynamic query scheduling strategy based on continuous batching, look-ahead scheduling, and partition prioritization, SieveIVF significantly enhances performance while preserving main-batch efficiency. Implemented atop Lance, it achieves 4.1–8.4× speedups over baselines across four 10M-scale and two 100M-scale datasets, with recall degradation limited to 0.03–2.29 percentage points.
📝 Abstract
Embedding-based training data deduplication retrieves candidate duplicate edges above an application similarity threshold, but fixed-probe inverted-file (IVF) search ignores this predicate when giving every query the same partition budget. Across four Hunyuan workloads, qualifying neighbors appear early despite sharply varying search depths. We present SieveIVF, a threshold-aware IVF executor that stops after $W$ consecutive searches find no qualifying candidate. The systems challenge is to preserve partition-major batching when each query's remaining work depends on prior results. Continuous batching groups ready queries by partition. A lookahead scheduler layers on top, exposing only work committed by the stopping rule to increase concurrency without changing stopping decisions or returned results. We implement SieveIVF in Lance. At $W=8$, SieveIVF is $4.1$--$7.6\times$ faster than fixed-probe IVF on four 10M Hunyuan workloads and $6.1$--$8.4\times$ faster on two public 100M workloads under the same index and search parameters, with pooled filtered top-10 recall losses of $0.03$--$1.13$ percentage points on Hunyuan and $1.43$--$2.29$ percentage points on the public workloads. These results show how an application predicate can guide IVF work allocation without changing the index or bounded top-$k$ interface.