🤖 AI Summary
This study addresses the challenge of distinguishing nascent topic signals from ordinary noise during the early stages of document publication. To this end, it proposes a forward-looking prediction method grounded in the geometric properties of embedding spaces. For the first time, this approach introduces multi-model consensus to annotate anomalous trajectories, relying exclusively on instantaneous information available at the time of publication. By integrating text embeddings with cross-validation techniques, the method successfully identifies critical early documents without requiring retrospective data. Experimental results demonstrate that the proposed model achieves an F1 score exceeding 0.90 on high-consensus subsets and maintains a robust performance of 0.76–0.80 under rigorous temporal evaluation settings, significantly outperforming existing baselines.
📝 Abstract
Some documents that embedding-based topic models initially classify as noise later become founding members of emerging topics. At publication time, however, they appear as scattered points in embedding space and are difficult to distinguish from ordinary noise without the benefit of hindsight. We study whether such anticipatory outliers can be predicted prospectively, using only information available when a document first appears. We derive labels from the subsequent trajectories of outlier documents, distinguishing those that anticipate new topics from those that reinforce existing topics or remain isolated, and estimate label confidence through agreement across multiple embedding models. On two French news corpora, anticipatory outliers prove predictable at publication time. Under cross-validation, $F_1$ rises from about 0.77 over the full eligible population to above 0.90 on high-consensus subsets, and remains at 0.76-0.80 under a strictly chronological evaluation. Predictive performance is driven mainly by geometric features capturing each outlier's position in embedding space.