latent event clustering

Designs and implements unsupervised, bottom-up clustering methods that group textual event mentions or their learned representations into latent event clusters or groups. These methods produce compact, de-duplicated event representations, mitigate bias from co-occurring or repeated events, and improve robustness of downstream models that consume event information.

latenteventclustering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Unsupervised Episode Detection for Large-Scale News Events

Aug 09, 2024
PK
Priyanka Kargupta
🏛️ UIUC

Existing automated event detection methods suffer from limited interpretability and poor adaptability to large-scale critical events, failing to capture “episodes”—semantically coherent subunits composed of core entities, actions, and spatiotemporal contexts. This paper introduces, for the first time, the unsupervised episode detection task, aiming to automatically identify such natural narrative fragments from news corpora. To address key challenges—including the absence of explicit spatiotemporal markers and the inadequacy of semantic similarity–based clustering—we propose EpiMine, a two-stage framework: (1) discriminative term migration to detect episode boundaries, and (2) large language model–driven reasoning and clustering over candidate fragments. Evaluated on three real-world, human-annotated datasets, EpiMine achieves an average 59.2% performance gain over all baselines. Our approach establishes a novel, interpretable, and scalable paradigm for critical event modeling.

Detecting cohesive episodes in news for key eventsImproving interpretability and adaptability in event detectionOvercoming lack of explicit markers in episode detection

EC-GCD (Event-Centric Generalized Category Discovery) faces two key challenges: (1) inconsistency between clustering and classification groupings due to subjective labeling, and (2) unfair representation of minority classes. To address these, we propose PaMA—a novel framework featuring: (i) an event-pattern extraction-and-refinement mechanism that aligns clustering and classification semantics; and (ii) a rank-filter-mine pipeline with dynamic prototype-balanced sampling to ensure equitable representation of minority classes in long-text, highly imbalanced settings. PaMA integrates large language model–driven event pattern mining, contrastive latent-space alignment, and H-score–based optimization. Evaluated on our newly constructed EC-GCD benchmark (e.g., Scam Report), PaMA achieves a 12.58% absolute improvement in H-score. Moreover, it demonstrates strong generalization on standard GCD benchmarks, confirming its robustness across diverse category-discovery scenarios.

Address imbalanced class distributions in event-centric contextsClassify known and novel categories with partial labelsImprove cluster-class alignment using LLM-extracted event patterns

Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka Embeddings

May 30, 2025
HW
Hans W. A. Hanley
🏛️ Stanford University

Existing multilingual topic modeling and clustering approaches suffer from poor scalability, opaque similarity metrics, and insufficient semantic granularity. To address these limitations, we propose the Multilingual Matryoshka Embedding framework, which encodes hierarchical semantic structures—from event-level to topic-level—into a single vector via nested, multi-granular representations. Our method integrates multilingual pretraining with a dimensionality-adaptive pruning mechanism, and introduces a hierarchy-aware similarity metric alongside a lightweight hierarchical clustering algorithm. Evaluated on SemEval 2022 Task 8, it achieves a Pearson correlation coefficient of 0.816—establishing a new state-of-the-art. The framework significantly advances cross-lingual story discovery and thematic abstraction, and, for the first time, enables multilingual hierarchical clustering that is unified (single embedding), multi-granular, interpretable, and strongly generalizable.

Addresses poor scaling and opaque metrics in topic modelingEnables hierarchical story and theme identification in news dataImproves multilingual news clustering scalability and interpretability

Existing event extraction datasets commonly suffer from limited event type coverage, domain closure, and a lack of large-scale human validation. To address these limitations, this work introduces EVENT5Ws—a large-scale, human-annotated, and statistically validated open-domain event extraction dataset. Through a systematic annotation pipeline and rigorous quality control mechanisms, EVENT5Ws achieves, for the first time, broad cross-regional coverage of diverse event types. The study also establishes strong baselines leveraging pretrained large language models. Experimental results demonstrate that EVENT5Ws substantially enhances model generalization across varied geographical contexts, offering a reliable resource and practical guidance for advancing open-domain event extraction research.

dataset limitationsevent extractionlarge-scale dataset

Existing approaches to media narrative analysis often struggle to balance fine-grained detail with scalability, either sacrificing nuance through coarse-grained modeling or relying on domain-specific annotations that limit generalizability. This work proposes an unsupervised method that jointly models events and characters and incorporates structured clustering to automatically induce interpretable narrative schemata from large-scale news corpora. By avoiding manual annotation, the approach yields narrative frameworks that align with theoretical expectations while demonstrating strong generalization capabilities. The resulting schemata preserve fine-grained semantic structure and enable efficient, interpretable discovery of media narratives across diverse datasets.

framing theorymedia narrativesnarrative structures

Latest Papers

What's happening recently
View more

This work addresses the limitation of traditional embedding models, which capture only semantic topics and fail to reveal structured functional patterns in text—such as narrative modes or registers. To overcome this, the authors propose a contrastive learning approach grounded in temporal co-occurrence relations, mapping pretrained embeddings into an associative space where recurrent cross-text transitional structures can be discovered under compression constraints. The study extends the Predictive Associative Memory framework from episodic memory to unsupervised concept formation, enabling an abstract shift from “what a text says” to “what a text does.” Evaluating on a corpus of 9,766 Project Gutenberg books, the method constructs a multi-resolution conceptual map and achieves a zero-shot assignment accuracy of 42.75%, substantially outperforming purely semantic clustering baselines.

corpus-scale analysispredictive associative memorytextual function

This work addresses a key limitation of mainstream centroid-based clustering methods in unsupervised term discovery: their inductive bias impedes the recovery of vocabulary that follows the Zipfian distribution observed in natural language. To overcome this, the authors propose a bottom-up graph clustering approach that constructs a similarity graph from pairwise embeddings of speech segments and partitions it using the Leiden algorithm to recover word- or syllable-level lexicons better aligned with Zipf’s law. This study provides the first systematic demonstration of graph clustering’s marked advantage in modeling Zipfian distributions, consistently outperforming baseline methods—including K-means, Gaussian Mixture Models, and BIRCH—across three languages. Although average-linkage agglomerative clustering yields comparable performance, it suffers from lower computational efficiency. By challenging the dominance of centroid-based paradigms, this work offers a novel pathway for unsupervised term discovery that more faithfully captures the statistical properties of language.

clustering biaslexical distributionspeech segmentation

This work addresses the common limitations of unsupervised text clustering—namely, semantically incoherent, redundant, and poorly interpretable clusters, coupled with a lack of effective validation mechanisms. The authors propose a three-stage reasoning framework leveraging large language models (LLMs) to semantically validate and reconstruct any given clustering result without requiring labeled data. The framework sequentially performs coherence checking, redundancy adjudication, and unsupervised label generation. Innovatively treating the LLM as a semantic adjudicator rather than an embedding generator, the approach decouples representation learning from structural validation. Experiments on two real-world social media datasets demonstrate that the method substantially improves cluster coherence and human alignment of generated labels, with manual evaluations strongly endorsing label quality and cross-platform robustness.

cluster coherencelabel interpretabilitysemantic redundancy

This study investigates effective strategies for selecting representative disaster event samples from news texts to support socio-environmental research, with a focus on landslides. It systematically compares two sampling paradigms: a “top-down” approach that retrieves news articles based on existing disaster inventories, and a “bottom-up” method that employs natural language processing (NLP) to perform spatiotemporal clustering of news reports. The findings demonstrate that the choice of sampling strategy significantly influences sample composition, thereby introducing biases in assessments of media reporting equity, disaster monitoring coverage, and inventory completeness. By highlighting the critical role of sampling design in shaping disaster information extraction, this work provides a methodological foundation for mitigating systematic biases inherent in media-derived datasets.

disaster coveragemedia representationnews data sampling

Hot Scholars

LH

Lei Hou

RMIT University
Building Information Modeling (BIM) - Project Management - Construction IT - Productivity Research - Lean Construction
BG

Banglei Guan

National University of Defense Technology
PhotomechanicsVideometrics
ME

Muhammad Ejaz Ahmed

CSIRO's Data61
Securitydigital forensicsmalware detectionthreat hunting
DW

David Windridge

Professor of Data Science & Machine Learning, Head of AI/ML Group, Middlesex University, London
Machine Learning/AIQuantum Machine LearningAstrophysics
LG

Li Gao

Search Science, Baidu Inc.
Information RetrievalRecommender System