π€ AI Summary
This study investigates effective strategies for selecting representative disaster event samples from news texts to support socio-environmental research, with a focus on landslides. It systematically compares two sampling paradigms: a βtop-downβ approach that retrieves news articles based on existing disaster inventories, and a βbottom-upβ method that employs natural language processing (NLP) to perform spatiotemporal clustering of news reports. The findings demonstrate that the choice of sampling strategy significantly influences sample composition, thereby introducing biases in assessments of media reporting equity, disaster monitoring coverage, and inventory completeness. By highlighting the critical role of sampling design in shaping disaster information extraction, this work provides a methodological foundation for mitigating systematic biases inherent in media-derived datasets.
π Abstract
News articles are an important source of information on disaster impacts and adaptation. A key methodological challenge in socio-environmental studies is how to select a representative data sample. Two approaches are common: querying news databases top-down with the aid of an existing disaster inventory or using NLP methods to cluster news texts bottom-up based on temporal and spatial features. Using a dataset of German news about landslides worldwide, we compare these approaches and discuss variations in event coverage. Such research design decision can influence the resulting news sample, affecting its use in studies of inequality in media coverage, disaster monitoring and inventory enrichment.